Subtitles · Captions for deaf and hard-of-hearing users

Live captions on your Mac,
without asking anyone to turn them on.

Most people who need captions still hear something, and the job is usually filling the words that went missing rather than replacing the ones that arrived. Subtitles reads the audio your Mac is playing, so it captions every app on the machine without asking anyone's permission, and the overlay stays on top while you work in front of it.

14-day refund, no questions asked

One-time $9, every future update included · macOS 14.2 or later · Apple Silicon
Unlimited captions, for as long as you like. No minutes to buy, no account, no subscription.

The parts that are actually hard

Captions existing is not the same as captions being usable. These are the places where the ones you are given tend to fall down, and what this does about each.

  • The line that is already gone. Google Meet keeps no caption history at all, and Zoom holds about three minutes. Look down to type and the sentence you half-read is not recoverable. Hold ⌥ here and the last captions come back, scrollable, so missing one line does not mean losing the thread of the meeting.
  • Two people talking at once. Overlapping speech is where every caption system struggles and where you need it most. A new box on speaker change is optional and worth switching on: in a conversation you read rather than hear, working out who said what is most of the work.
  • Video nobody captioned. Closed captions and SDH exist only where somebody made them, which rules out most training videos on an internal portal, most conference recordings, most screen recordings a colleague sent you and most podcasts. This reads the audio itself, so whether the publisher bothered stops being the question.
  • Calls that are not meetings. FaceTime, a WhatsApp or Signal call on the desktop, a number dialled through the Mac, a voicemail played back. None of these has a captions button to press and none of them has a host to ask, which is most of why phone calls stay the part of the week people dread.
  • Watching something with people who can hear it. Turning subtitles on for a film is a negotiation when nobody else in the room needs them, and plenty of things have no subtitle track to turn on anyway. This draws on your Mac only, sized how you like it, so it stops being everyone's decision.
  • Hearing aids that already stream from the Mac. Made for iPhone hearing aids and sound processors pair straight to an M1 Mac or later, and audio arriving cleanly is not the same as fast, overlapping dialogue arriving intelligibly. Captions over streamed audio is an ordinary setup rather than a redundant one. It reads the app's own audio stream rather than your output device, so what you are listening on is not part of the picture.
  • The tiredness at four in the afternoon. Concentrating hard enough to follow speech all day is exhausting in a way that shows up on nobody else's calendar. Reading a word you only half-heard costs less than reconstructing it from context, and it costs a great deal less than asking a room to repeat itself for the third time.

What it hears, and what it does not

This part belongs at the top rather than in the small print, because it decides whether the app is any use to you at all. Subtitles reads one thing: the audio a Mac application is playing. That is a narrow tap, and the narrowness is deliberate, but it rules some things out completely.

  • It does not caption the room. The microphone is never opened, so a person speaking in front of you produces nothing on screen. If what you need is the conversation at the table, this is the wrong tool and no setting will change that. macOS Live Captions can take microphone audio, and Live Listen on an iPhone is built for the room; either is a better answer than this one.
  • It does not caption you. Only what the Mac plays is read, so on a call you get the other people and never yourself.
  • It is speech recognition, so it is wrong sometimes. Names, jargon, crosstalk and unfamiliar accents are where it slips. For an appointment, a legal conversation or a class you are assessed on, a human captioner or an interpreter is the accommodation, and asking for one is your right rather than a favour. This is for the rest of the day, where nobody was going to provide one anyway.
  • It writes speech, not sound. This is the real gap against SDH, which also marks the doorbell, the laughter and the music cue. A voice detector runs ahead of the recogniser here so that a backing track never becomes lyrics, and the cost of that is the rest of it: you get what was said, and nothing about what else was audible.
  • It suits some hearing more than others. If you hear much of a conversation and lose words in noise, at speed or over a bad microphone, this is aimed squarely at you. If you use ASL and want an interpreter, or you read text far faster than any recogniser can produce it, then it is a smaller addition and worth trying before buying.
  • It keeps nothing. Captions are drawn on screen and then they are gone: nothing is written to disk, so there is no transcript to read back afterwards. If you need a record, this is not where to get one.

Try the one already on your Mac first

macOS has Live Captions built in, free, on every Apple silicon Mac: System Settings ▸ Accessibility ▸ Live Captions. It is genuinely good, it costs nothing, and if it covers what you need then you are done and should not spend $9 to find that out. Spending a week with it is also the fastest way to learn what you actually want from captions.

What sends people looking for something else is usually one of three things: the language they need is not in Apple's list, the caption window sits behind whatever they are working in, or they want to choose a slower and more accurate model for a lecture and a faster one for a call. Those are the differences, laid out row by row, on the comparison page.

14-day refund, no questions asked.

When you are the only one who needs it

The tiring part of captions is rarely the reading. It is that every source of them belongs to somebody else, so getting them means asking, and asking means explaining, and explaining is a thing you do again on the next call with the next set of people.

Nobody to ask

Zoom's captions need the host to switch them on. Teams' depend on what the tenant allows. Meet's vary by Workspace tier. A tap on your own machine answers to none of that, so the captions are simply there on a call with a stranger, a recruiter, or a company that has never once thought about it.

Nothing announces you

No bot joins, nothing appears in the participant list, no recording badge lights up. Telling people you are deaf or hard of hearing stays a decision you make in your own time, rather than one the software makes for you at the exact moment you need to follow along.

No repeat performances

Asking a room to say it again has a cost that goes up each time you spend it, and most people stop paying it long before they stop missing things. Reading the line you half-heard costs nothing and interrupts nobody, which is the difference between following a meeting and nodding through one.

On for as long as you need it

The model runs on the Neural Engine in your Mac and costs nothing per hour, so there is no meter and nothing to ration. Captioning services priced by the minute quietly ask you to decide which parts of your day are worth hearing; this one does not have an opinion about that.

Built to be read for hours

Captions you lean on all day are read differently from captions you glance at once. Everything below is a setting rather than a decision somebody made for you, because the right answer depends on how much you are hearing and how much you are reading.

  • It does not go behind your work. The overlay sits above every window and the app switcher, so taking notes, checking a document or looking something up does not cost you the sentence being said while you do it. Following a meeting and writing anything down at the same time is the whole difficulty, and a caption window you have to keep in front does not solve it.
  • You can put it where your eyes already are. It takes no clicks, so you work straight through it, and it fades under the pointer so it never hides what is beneath. Size, opacity, how far that fade reaches and where the box lives are all yours; hold Shift to keep it solid and drag it. Sitting it just under the speaker's face is what most people end up doing, and it is why lip-reading and captions can run together instead of competing.
  • Lines wait for you. A caption holds for about a second and a half plus a tenth per word, so a long one does not vanish mid-sentence, and holding ⌥ brings the last few back to scroll through. Reading speed is personal and the default is a guess; the history is there for when the guess is wrong.
  • Names and jargon are where it will let you down. A drug name, a surname, an acronym your team invented: these are exactly the words you cannot infer from context, and exactly the ones a recogniser gets wrong. Picking a slower, more accurate model helps and does not fix it. It is worth knowing which words to double-check rather than finding out later.

On-device is not a preference here

For most people local transcription is a nice principle. If captions are how you access speech, it is a different calculation: a cloud captioning service would receive a transcript of your appointments, your work, your family and everything else you needed help hearing. That is not a by-product of the accommodation, it is a record of your day held by a company you did not choose to tell.

Nothing here is uploaded, nothing is recorded and nothing is written to disk. The audio is read and discarded as fast as it arrives, and after the model is fetched on first run the app makes no network calls at all. There is no account, no telemetry and nothing to opt out of, because there is nothing collected to opt out of. The source is public if you would rather check than take that on trust.

Questions

Does it caption people talking in the room?
No. It only ever reads what your Mac plays, and it never opens the microphone. For a conversation happening in front of you it will show you nothing at all. Apple's own Live Captions can caption microphone audio on a Mac, and Live Listen on iPhone is built for exactly that job, so for the room those are the right tools and this is not.
Is this a replacement for CART or a human captioner?
No. It is speech recognition, so it makes mistakes, particularly on names, jargon, crosstalk and heavy accents. For a medical appointment, a legal meeting, a class you are graded on or anything else where being wrong has a cost, a professional captioner or interpreter is the accommodation and you are entitled to ask for it. This is for the rest of the day, where nobody was ever going to provide one.
Do I have to ask the host, or tell anyone I am using it?
No. It captions the audio your Mac plays rather than integrating with any service, so nothing joins the meeting, nothing appears in a participant list, and no host, admin or lecturer has to enable anything. Nobody is told. Whether you disclose is your decision rather than the software's.
I already stream my Mac's audio to my hearing aids. Does this add anything?
Often, yes. Streaming fixes the path the sound takes, not how hard the speech is: fast talkers, crosstalk and unfamiliar accents are still fast talkers, crosstalk and unfamiliar accents once they arrive in your ear. Captions over streamed audio is a common setup rather than a redundant one. It reads the app's own audio stream rather than your output device, so what you listen on is not part of the picture and nothing about your pairing changes.
What if a video has no subtitles or SDH at all?
That is the case it is for. Closed captions and SDH only exist where somebody made them, which leaves out most internal training videos, most conference recordings and most things a colleague records and sends you. This captions the audio itself, so it does not matter whether the file came with a subtitle track.
Do the captions lag behind, like the ones in my meetings?
Less, but not never. It captions continuously as words arrive rather than waiting for a finished sentence, and the newest word is drawn dimmer until the model is sure of it, so you are reading along rather than reading afterwards. It is still speech recognition: there is some delay between a word being said and that word being readable, and somebody speaking quickly will always be a little ahead of any caption.
Can I make the captions bigger and easier to read?
Yes. Box size, opacity, how many lines it holds and where it sits are all settings, and the settings window previews them on the real overlay rather than in a mockup. Hold Shift to keep it solid and drag it where you want it; where you put it is remembered.
Does it caption FaceTime and phone calls?
Anything that plays through an app on the Mac, which is FaceTime, a WhatsApp or Signal desktop call, a number dialled through your Mac and a voicemail played back. You read the other person and never yourself, because your microphone is never opened. These are the calls with no captions button and no host to ask, so they are often the first thing people try it on.
Will it tell me who is speaking, or mark sounds like SDH does?
It can start a new caption box when the speaker changes, which is optional and worth turning on for calls, but it does not put a name to anyone. It does not mark non-speech sound at all: a voice detector runs ahead of the recogniser so that music never turns into invented lyrics, and the price of that is you get the speech and nothing about what else was audible. If sound cues matter to you, this is the honest gap against a proper SDH track.
What happens when I miss a line?
Hold ⌥ and the last captions come back, and you can scroll through them while they are up. It is deliberately short-term: nothing is written to disk, so it is there to catch the sentence you just lost rather than to keep a record of the day.
Is there a per-minute or monthly cost?
No. It is $9 once and the captions are unlimited, because the model runs on your Mac and costs nothing to run. There are no minutes to buy and no subscription, which matters more when captions are how you follow everything rather than something you switch on occasionally.
Does it need an internet connection?
Once, to fetch the model on first run. After that it runs entirely offline and makes no network calls at all, so nothing you listen to is ever sent anywhere.

Captions on everything, without asking.

14-day refund, no questions asked

One-time $9, every future update included · macOS 14.2 or later · Apple Silicon
Unlimited live captions, for as long as you own the Mac. No minutes to buy, no account, no subscription.