- Does it work with Zoom, Teams, and Google Meet?
-
Yes. It taps the audio an app plays rather than integrating with any
one service, so a Zoom window, a Slack huddle and a call in a browser
tab all caption the same way. Browsers and Electron apps included:
it captures the helper processes those spawn, which is where their
audio actually comes from.
- Which languages does the app caption?
-
Sixteen: English, Spanish, French, Italian, Portuguese, German,
Dutch, Turkish, Russian, Arabic, Hindi, Japanese, Korean,
Vietnamese, Ukrainian and Mandarin. The app detects which one is
being spoken, or you can name the language in the menu to stop the
guessing. The list is the same on every supported macOS version,
with no country or region condition, and from macOS 15 any of them
translates into any other.
- Can it translate what it hears?
-
Yes, since 1.3. Translate To in the
menu takes any of the sixteen languages into any of the other
fifteen, through Apple's on-device translator, so a translated
caption never leaves the machine, just like an untranslated one
never does. It is off until you pick a target, and holding
⌃ shows the original language underneath while it runs.
You can also show both the translation and the source in the
overlay at the same time, and Speak Translation, the row below
Translate To, reads the translation aloud, with the original turned
down underneath. It needs macOS 15; captioning alone still runs on
14.2.
- Does it work with VoiceOver?
-
Yes. VoiceOver reads the caption box and presses its button, and
Speak Translation, on by default whenever VoiceOver is running,
reads the translation aloud, so a call or a film in another
language can be followed by ear. What the app asks you, such as
whether to save a transcript, is spoken as it appears, and
VoiceOver's own speech is never captioned or translated.
- Doesn't an overlay get in the way?
-
It takes no clicks at all, so you work straight through it, and it
fades away under the pointer so you can read and interact with what
is beneath it. How much it fades, how far that reaches, how solid
the box is, and how many lines it holds are all settings. Hold
⇧ to keep it solid and drag it somewhere else; wherever
you put it on your screen is remembered.
- Does it caption my own voice?
-
Not unless you choose the microphone in Listen To. By default, it
hears what your Mac plays, so on a call you read the other people,
not yourself. With the microphone chosen, it captions the room
instead, macOS asks for the permission for it the first time, and
with Show Both Languages on under Translate To, two people can
talk in two languages, with each reading their own.
- Does it need an internet connection?
-
Once, to fetch the model on first run, about 633 MB. After that
it runs entirely offline. The only other requests it can make are a
daily check for a newer version, which it asks permission for first
and which you can decline, and a license check from time to time.
- Can I save or export a transcript?
- Yes, when you ask. After a minute or more of speech, the caption box offers to save the session when the audio stops, or before a long enough silence clears the rewind history, and Save Transcript… in the menu saves it at any time, as plain text, Markdown or SRT subtitles; nothing is written to disk otherwise. While the app is running, holding ⌥ rewinds through the finished boxes and ⌥F searches them, for finding the line you looked away from. If you need notes or a summary of a conversation, this is the wrong tool; it was designed as complementary to note-taking apps.
- Does IT have to approve it?
-
There is no vendor here to approve: no account, no server, no
transcript leaving the Mac, and nothing joining the call as a
participant, so nobody in the meeting is being recorded by anything.
It is one signed, notarized app on one machine, and after the model
downloads once, the only requests it can make are a version check you
can switch off, and a license check. Whether that clears your own
policy is your call, but there is no third party in it to review.
- Does it work with headphones or AirPods?
-
Yes. It captures the audio an app produces, not the sound coming out
of a speaker, so what you listen on makes no difference.
- Will it slow my Mac down?
-
Transcription runs on the Neural Engine rather than the CPU, at roughly
a 0.15 real-time factor, or six seconds of speech per second of
compute. The same model on the CPU measured about 100× too slow to keep
up.
- Which Macs does it run on?
-
Apple silicon, macOS 14.2 or later. The version floor is Core Audio
process taps, which is the API the whole thing is built on.
Translations run from macOS 15 upward.
- Is it real-time?
-
Yes. The model is a streaming ASR model, so a caption appears while
the sentence is still being said rather than after it ends. It reads
the system audio an app plays, with no loopback driver to install,
and the English options allow you to trade latency against accuracy
from 160 ms upward. If you searched for speech to text or
real-time transcription for the Mac, this is that, drawn on screen
as it plays, rather than saved to a file.
- Is it a subscription?
-
No. The download is a free seven-day trial. A license key is $19 once,
every future update included, no account to create and no minutes to
buy; it covers every Mac you use, and the app checks it with Gumroad
about once a month, sending the key and nothing about you. If it is not
for you, there is a 14-day refund.
- Can I see the source?
-
Yes. It is published under FSL-1.1-ALv2, which becomes Apache 2.0
two years after each release. Read it, or build it yourself, at
github.com/daformat/subtitles.