VoxAI 隐私政策
最后更新:2026 年 7 月 15 日 · 适用版本:v2.0
VoxAI 是什么
VoxAI 是 macOS 上的 AI 副驾。它听你的会议、通话和面谈,实时转成文字并区分说话人;你可以把自己的 AI 助手(Claude Code、Cursor 等)接进来,让它读到这些内容并给你建议。VoxAI 也保留了最初的语音输入能力——说话转文字,自动复制到剪贴板。
VoxAI 不做什么
- VoxAI 没有自己的服务器。你的音频和转录文字从不上传给作者——技术上无处可传。
- 不收集任何个人信息。不要求注册、不需要登录、没有用户账号系统。
- 不追踪你,不集成任何第三方分析或广告 SDK。
- 不向任何服务器发送使用统计、崩溃日志或遥测。
- 不读取你的文件、通讯录、日历或浏览器数据。
你的数据存在哪里
全部在你这台 Mac 上,在 App 的沙盒容器内:
~/Library/Containers/com.ethanys.voxai/Data/Library/Application Support/VoxAI/
录音文件
会议、通话、面谈场景的录音会写入磁盘,保存为未压缩的 WAV 文件,存在上面那个目录的 recordings/ 下。
转录文字
所有会议的逐句转录存在同目录的 meetings.json 里,是明文 JSON:每句包含时间戳、说话人标签和原文。
说话人识别(声纹)
会议和面谈场景会实时区分说话人。这项识别完全在你的 Mac 上本地完成,用的是本地 CoreML 模型(FluidAudio)。
- 声纹特征只存在于内存中,录制结束即释放——从不写入磁盘,从不上传到任何地方。
- 留在本机的只有识别结果标签(如「说话人 1」「说话人 2」),不含任何声学特征。
- VoxAI 不建立跨会议的声纹库——每场录制都从零开始,上一场的说话人身份不会带到下一场。
- 模型本身会在首次使用时从 HuggingFace 下载(约 13 MB,一次性)——详见下面「VoxAI 什么时候联网」。
语音对话模式
你在语音对话模式里说的话,最近 100 条会存在本机的 dialog_input.json,供你接入的 AI 读取。这个文件在本机,不会上传。
偏好设置与试用次数
VoxAI 在本机的 UserDefaults 里存 App 自己的偏好(识别语言、外观、是否自动复制到剪贴板等)。副驾功能的试用次数也只记在本机(单场录制满 60 秒才计一次)——这个计数不发送给任何人,包括作者和 Apple。
系统剪贴板
语音输入停止后,VoxAI 会把转录文字写入系统剪贴板(默认开启,可在设置里关闭)。剪贴板的内容完全由你控制。
通话场景:关于录到对方的声音
技术上是这样实现的:通话场景需要 macOS 的屏幕录制权限,VoxAI 用它捕获你 Mac 播放出来的声音,也就是对方在通话里说的话。
对方的声音会和你的声音分别录成两个 WAV 文件、分别转成文字,对方的部分标记为「对方」。两者都和其他录音一样存在你本机的容器里,永久保留直到你删除,不会上传给 VoxAI——VoxAI 没有服务器可传。
VoxAI 什么时候联网
VoxAI v2.0 的 App 沙盒权限有三条:
com.apple.security.app-sandbox(App Store 强制要求)com.apple.security.device.audio-input(麦克风访问)com.apple.security.network.client(发起对外网络连接)
第三条是 v2.0 新增的。以下是它的全部用途,没有别的:
| 什么时候 | 连向谁 | 发出去什么 | 默认 |
|---|---|---|---|
| 会议 / 面谈首次录制 | HuggingFace | 只下载说话人识别模型(约 13 MB,一次性)。不发送任何音频或文字。对方能看到的只有你的 IP 地址和这次下载请求。 | 开启(说话人识别的必需品) |
| 语音转文字 | Apple | 在 macOS 26 及以上,VoxAI 优先使用 Apple 的本地识别引擎;在更早的系统上使用 SFSpeechRecognizer,按 Apple 的隐私政策,音频可能被发送到 Apple 的服务器处理。两者都是 Apple 的系统框架,VoxAI 不经手。 |
系统行为 |
| App 启动、购买、恢复购买 | Apple App Store | StoreKit 查询商品价格与你的购买状态。Apple 会知道这台设备 / 这个 Apple ID 是否购买过。不涉及你的录音或转录。 | 开启 |
| AI 朗读(仅当你选择 Azure 语音) | 微软 Azure | 要朗读的那段文字,用你自己的密钥和额度发给微软。VoxAI 不中转、不留副本。适用微软的隐私政策。 | 关闭——默认用 macOS 自带的本地朗读,完全不联网 |
| 会后云端精处理(仅当你主动对某场录制点「精处理」) | AssemblyAI | 那场录制的完整音频文件,用你自己的密钥和额度上传,换回质量更好的转录与说话人分离。VoxAI 不中转、不留副本,适用 AssemblyAI 的隐私政策。这是唯一会让你的录音离开这台 Mac 的功能。 | 关闭 |
VoxAI 从不把你的录音或转录文字上传给作者——我没有服务器可以收。上表里唯一会送出你会议内容的是会后云端精处理,而它默认关闭、需要你逐场主动触发、用你自己的账号直接发给 AssemblyAI,不经过我这里。
会后云端精处理(默认关闭)
实时转录难免有误。你可以选择在一场录制结束后,把它送到 AssemblyAI 做一次更高质量的转录和说话人分离。这需要你自己在 AssemblyAI 注册、拿到密钥、填进设置——用的是你自己的账号和额度,VoxAI 不代收费用也不代为中转。
- 要上传的是那场录制的完整音频。点下去之前请确认:这场录音里如果有别人的声音,你已经取得了他们的同意。
- 只有你逐场主动点击才会发生。VoxAI 不会自动上传任何东西,也不会批量处理你的历史录音。
- 不填密钥,这个功能就完全不存在——不填就没有任何东西会被上传。
- 原始转录会被保留,你可以对比、也可以撤销精处理的结果。
- 音频到了 AssemblyAI 之后如何被处理和保留,适用他们的隐私政策,不是这一份。
你的密钥存在哪
如果你使用云端精处理或 Azure 语音朗读,需要填入你自己的密钥。两者的存放方式不同,我们照实说明:
- 云端精处理(AssemblyAI)的密钥:存在 macOS 钥匙串(Keychain)里,由系统加密保管。
- Azure 语音朗读的密钥:以明文存放在 VoxAI 应用容器内的配置文件里,没有加密。之所以没放进钥匙串:朗读是由容器外的伴生 server 执行的,而它读不到 App 的钥匙串。macOS 会在其他程序首次访问 VoxAI 容器时向你请求授权,但这不等于加密保护。如果你不使用 Azure 朗读(默认就不使用),就不必填这个密钥。我们打算在后续版本里解决它。
两者相同的是:只留在你这台 Mac 上,不会发送给作者,VoxAI 也不会把它们显示给你接入的 AI。
你接入的 AI 能读到什么
这是 VoxAI 的核心功能,也是最需要你了解的一条数据流。
你可以选择安装一个伴生的 MCP server,把 VoxAI 接进你的 AI 客户端(Claude Code、Cursor、Codex 等)。接入之后,你的 AI 可以读取:会议转录、实时转录、说话人标签、你在语音对话模式说的话。
关于这条链路的几点澄清:
- 这条数据流不经过 VoxAI 的任何服务器——是你的 Mac 直接把内容交给你自己安装的 AI 客户端。
- MCP server 是一个独立进程,在 App Store 之外单独分发(GitHub Release)。VoxAI App 本体不下载它、不执行它——安装与运行都由你和你的 AI 客户端完成。
- 它只读写 VoxAI 容器内的文件,从不上传你的会议内容。
- 你填的密钥不会被它回显给 AI。
伴生 server 自己会联网做什么
如果你接入了 AI 并让它开口说话,伴生 server 有两处会联网——都不涉及你的会议内容:
- Azure 语音朗读(默认关闭):见上表。开启后把要朗读的文字发给微软,用你自己的密钥。
- 本地语音组件(默认未安装):VoxAI 提供一套可离线运行的中文语音。你的 AI 可以在你同意下触发安装,届时伴生 server 会从
voxai.ethanflow.com(作者的域名)下载语音组件,并从 HuggingFace 下载语音模型(较大,约 2 GB)。这是本政策里唯一一个连向作者服务器的请求——作者由此能看到的只有你的 IP 地址和这次下载请求,没有你的任何音频、文字或身份信息。组件装好后完全离线运行,朗读不再联网。
这两件事都发生在伴生 server 里,VoxAI App 本体全程不参与下载、也不执行任何下载来的代码。不安装伴生 server 的话,这一整节都与你无关。
儿童使用
VoxAI 不专门面向 13 岁以下儿童,也不会向作者收集任何用户数据(包括儿童的数据)。
政策变更
如果未来版本改变了本政策描述的任何数据行为——新增联网用途、改变存储方式、引入新的第三方服务——本政策会在该版本发布时同步更新,页首的「最后更新」日期会随之变化。
联系
如有疑问,请发邮件至 redlilyholmes@gmail.com 联系作者。
VoxAI Privacy Policy
Last updated: July 15, 2026 · Applies to: v2.0
What VoxAI is
VoxAI is an AI copilot for macOS. It listens to your meetings, calls, and in-person conversations, transcribes them in real time, and identifies who's speaking. You can connect your own AI assistant (Claude Code, Cursor, etc.) so it can read along and advise you. VoxAI also keeps its original dictation feature — speak, get text, auto-copied to your clipboard.
What VoxAI does NOT do
- VoxAI has no servers of its own. Your audio and transcripts are never uploaded to us — there is technically nowhere to upload them to.
- Does not collect any personal information. No registration, no login, no user accounts.
- Does not track you. No third-party analytics, no advertising SDKs.
- Does not send usage stats, crash logs, or telemetry to any server.
- Does not read your files, contacts, calendar, or browser data.
Where your data lives
All of it on this Mac, inside the app's sandbox container:
~/Library/Containers/com.ethanys.voxai/Data/Library/Application Support/VoxAI/
Audio recordings
Meeting, call, and in-person sessions are written to disk as uncompressed WAV files, under recordings/ in the directory above.
Transcripts
Every meeting's transcript is stored in meetings.json in the same directory, as plain JSON: each line has a timestamp, a speaker label, and the text.
Speaker identification (voice embeddings)
Meeting and in-person modes tell speakers apart in real time. This runs entirely locally on your Mac using an on-device CoreML model (FluidAudio).
- Voice embeddings exist in memory only and are released when recording ends — never written to disk, never uploaded anywhere.
- All that persists is the resulting label ("Speaker 1", "Speaker 2"), which contains no acoustic features.
- VoxAI does not build a voiceprint library across meetings — every session starts fresh; speaker identities do not carry over.
- The model itself is downloaded from HuggingFace on first use (~13 MB, one time) — see "When VoxAI uses the network" below.
Voice dialog mode
What you say in voice dialog mode is kept locally — the most recent 100 entries in dialog_input.json — so your connected AI can read it. This file stays on your Mac.
Preferences and trial count
VoxAI stores its own preferences in local UserDefaults (recognition language, appearance, auto-copy toggle, and similar). The copilot trial count is also kept only on this Mac (a session counts only once it passes 60 seconds) — that count is not sent to anyone, including us and Apple.
System clipboard
When dictation stops, VoxAI writes the transcript to the system clipboard (on by default, can be turned off in Settings). The clipboard's contents are entirely under your control.
Call mode: about recording the other party
How it works technically: Call mode requires macOS Screen Recording permission, which VoxAI uses to capture the audio your Mac plays back — that is, the other party's side of the call.
The other party's audio and yours are recorded as two separate WAV files and transcribed separately, with their side labeled "partner". Both are stored in your local container like any other recording, kept until you delete them, and never uploaded to VoxAI — there is no VoxAI server to upload to.
When VoxAI uses the network
VoxAI v2.0 declares three sandbox entitlements:
com.apple.security.app-sandbox(required by the App Store)com.apple.security.device.audio-input(microphone access)com.apple.security.network.client(outbound network connections)
The third is new in v2.0. Here is every use of it — there are no others:
| When | To whom | What is sent | Default |
|---|---|---|---|
| First meeting / in-person recording | HuggingFace | Download only — the speaker identification model (~13 MB, one time). No audio or text is sent. All they can see is your IP address and the download request. | On (required for speaker ID) |
| Speech to text | Apple | On macOS 26 and later, VoxAI prefers Apple's on-device engine. On earlier systems it uses SFSpeechRecognizer, and per Apple's privacy policy audio may be sent to Apple's servers for processing. Both are Apple system frameworks; VoxAI doesn't handle the transmission. |
System behavior |
| App launch, purchase, restore | Apple App Store | StoreKit queries product pricing and your purchase status. Apple learns whether this device / Apple ID has purchased. None of your recordings or transcripts are involved. | On |
| AI speech output (only if you choose Azure) | Microsoft Azure | The text to be spoken, sent to Microsoft using your own key and quota. VoxAI neither proxies it nor keeps a copy. Microsoft's privacy policy applies. | Off — the default is macOS's built-in local speech, which uses no network at all |
| Cloud refine (only when you explicitly refine a session) | AssemblyAI | That session's complete audio file, uploaded using your own key and quota, in exchange for a higher-quality transcript with speaker separation. VoxAI neither proxies it nor keeps a copy; AssemblyAI's privacy policy applies. This is the only feature that takes your recordings off this Mac. | Off |
VoxAI never uploads your recordings or transcripts to us — we have no server to receive them. The only row above that sends your meeting content anywhere is cloud refine, which is off by default, must be triggered by you per session, and goes directly to AssemblyAI under your own account, never through us.
Cloud refine (off by default)
Live transcription is never perfect. After a session ends, you can choose to send it to AssemblyAI for a higher-quality pass with speaker separation. This requires you to sign up with AssemblyAI, get your own key, and enter it in Settings — it runs on your own account and quota; we neither charge for it nor proxy it.
- What gets uploaded is that session's complete audio. Before you click it, make sure that if the recording contains other people's voices, you have their consent.
- It only ever happens when you click it, session by session. VoxAI never uploads anything automatically and never batch-processes your history.
- No key, no feature — with no key entered, nothing can be uploaded at all.
- Your original transcript is kept, so you can compare the two or undo the refine.
- Once the audio reaches AssemblyAI, how it is handled and retained is governed by their privacy policy, not this one.
Where your keys are stored
If you use cloud refine or Azure speech, you supply your own keys. The two are stored differently, and we'd rather say so plainly:
- The cloud refine (AssemblyAI) key is stored in the macOS Keychain, encrypted by the system.
- The Azure speech key is stored in plain text in a config file inside VoxAI's app container — it is not encrypted. Why it isn't in the Keychain: speech output is performed by the companion server, which lives outside the container and cannot read the app's Keychain items. macOS will ask for your permission the first time another program accesses VoxAI's container, but that is not the same as encryption. If you don't use Azure speech (the default is not to), you never need to enter this key. We intend to fix this in a future release.
What's true of both: they stay on this Mac, are never sent to us, and VoxAI never shows them to your connected AI.
What your connected AI can read
This is VoxAI's core feature, and the data flow you most need to understand.
You can install a companion MCP server to connect VoxAI to your AI client (Claude Code, Cursor, Codex, etc.). Once connected, your AI can read: meeting transcripts, the live transcript, speaker labels, and what you say in voice dialog mode.
A few clarifications about this path:
- This data flow does not pass through any VoxAI server — your Mac hands the content directly to the AI client you installed.
- The MCP server is a separate process, distributed outside the App Store (via GitHub Releases). The VoxAI app itself neither downloads nor executes it — installing and running it is done by you and your AI client.
- It only reads and writes files inside VoxAI's container; it never uploads your meeting content.
- The keys you enter are not echoed back to the AI.
What the companion server itself connects to
If you connect an AI and let it speak, the companion server has two network uses of its own — neither involves your meeting content:
- Azure speech (off by default): see the table above. Once enabled, the text to be spoken goes to Microsoft using your own key.
- Local voice component (not installed by default): VoxAI offers a set of Chinese voices that run fully offline. Your AI can trigger the install with your approval, at which point the companion server downloads the voice component from
voxai.ethanflow.com(the author's domain) and the voice models from HuggingFace (large — around 2 GB). This is the only request in this policy that reaches the author's server — all we can see from it is your IP address and the download request, none of your audio, text, or identity. Once installed, the component runs entirely offline; speaking never touches the network again.
Both happen inside the companion server. The VoxAI app itself never downloads or executes any of it. If you don't install the companion server, this whole section doesn't apply to you.
Children
VoxAI is not directed at children under 13, and collects no user data of any kind for us (children's or otherwise).
Policy changes
If a future version changes any data behavior described here — a new network use, a change in storage, a new third-party service — this policy will be updated when that version ships, and the "Last updated" date at the top will change with it.
Contact
For questions, contact the author at redlilyholmes@gmail.com.