Key Takeaways
Accessible captions are not just captions that exist. They carry the dialogue plus the non-speech sound that matters, they mark every speaker change, and they appear and disappear in step with the speech. Subtitles are a different thing again, usually dialogue only and often translated. A transcript is the text without timing, and it serves people who cannot use the player at all, so it is a separate deliverable rather than a substitute. Automatic captions with no correction pass are the usual failure.
Someone has told you the captions have to be accessible. You have captions. You are now trying to work out what the gap is between those two sentences, and nobody has been very specific about it.
The gap is almost never that a video has no caption track. It is that the caption track is unusable in a way that does not show up until a deaf or hard-of-hearing viewer tries to follow the video with it. Below is what that actually means, in the order you will run into it.
Three words, three deliverables, three audiences
Captions, subtitles and transcripts get used interchangeably in ordinary conversation, and the requirement you have been handed almost certainly does not use them interchangeably. If it names two of them, it wants two things.
Captions carry the dialogue plus the non-speech sound that matters. They are written on the assumption that the viewer is not hearing anything, so anything the soundtrack is doing that the picture does not show has to be in the text.
Subtitles usually carry dialogue only, and are often a translation. They assume you can hear the video perfectly well, you just do not follow the language being spoken. A subtitle track will not tell you the alarm went off, because a hearing viewer in the target language does not need telling.
A transcript is the text on its own, without timing. No cues, no sync, no player involved. It serves people using a screen reader, people who cannot operate the video player, and people who would simply rather read the thing than sit through it.
That last one is the distinction people most often collapse. A transcript posted next to the video does not discharge a caption requirement, and a caption file does not discharge a transcript requirement, because they reach different people by different routes. If you produce one and call it both, you have produced one.
Worth saying plainly: accessibility obligations differ by country, by sector and by organisation. There is no single global standard you can check yourself against and be done. The rule that matters is the one your organisation is actually bound by, and someone in your organisation knows which one that is even if they have not told you. Ask before you build a process around a guess.
Captions carry the sound, not only the speech
This is the part that surprises people who have only ever thought of captions as speech-to-text. A door closing. An alarm. Laughter from off screen. A phone ringing that makes the person on screen stop mid-sentence. None of that is dialogue, and none of it is visible.
If a viewer who cannot hear the soundtrack watches your video and something happens that they cannot account for, the captions have failed at exactly the thing they exist for. The person on screen reacting to a noise you never mentioned is now behaving inexplicably.
The test is not “was there a sound” but “does the sound carry meaning”. Room tone does not need a caption. Background music that nobody reacts to usually does not either. The alarm that makes everyone leave the room does. So does laughter, when it tells you how a line landed. The judgement call is yours, and it is a real judgement call rather than a rule you can automate, which is one reason a human pass over the file matters. If you want the longer version of how this is handled in practice, the guide to closed captions covers the mechanics.
Do not overdo it either. A caption track that annotates every rustle competes with the dialogue for reading time and pushes the viewer behind the video. Include what changes the meaning of the scene. Leave the rest.
Showing who is speaking
A caption track that runs three people’s dialogue together with no indication of where one stops and the next starts is technically present and functionally unusable. This is one of the most common defects in a file that has passed every box-ticking check, because the check asks whether captions exist, not whether they can be read.
Think about what a hearing viewer gets for free. Voices sound different. You know the answer came from someone else because it came from a different throat, and often you can hear which side of the room it came from. Strip the audio away and all of that disappears. On screen, a reaction shot or a tight framing can leave the speaker off camera entirely for the length of a sentence.
The conventions for marking a change vary between houses and platforms, and any of the common ones work as long as you use them consistently within a file. What does not work is nothing at all. An interview, a panel, a meeting recording, anything with more than one voice, needs the handover marked every single time it happens, including when the same two people are trading short lines quickly. Especially then, in fact, because that is where a reader loses the thread fastest.
Timing that tracks the speech
Captions have to appear when the words are spoken and leave when they stop. A caption that lands a beat late means the viewer reads the line after watching the reaction to it. A caption that lingers over the next shot means they are reading one scene while looking at another.
Neither of those is fatal on a single cue. Across ten minutes they are exhausting, and the video becomes harder to follow with the captions than a viewer might have managed without them. Drift is the usual culprit: a file that starts in sync and slips further behind as it goes, often because it was generated against one version of the video and applied to another that has an extra second of top-and-tail.
Check the end of the file, not the beginning. Everyone checks the beginning. If the last cue in a long video still lands on the words, the middle is probably fine. If it has slid, the whole file needs re-timing rather than patching. A structural check with an SRT file validator will catch overlapping or malformed cues before you upload, though it cannot tell you whether the words are the right words.
Automatic captions and the pass that makes them usable
Automatic captions with no correction pass are the single most common way an organisation ends up with captions that do not meet its own requirement. They typically arrive with punctuation decisions you would not trust, no speaker changes at all, and no non-speech sound, because the model was asked to transcribe speech and that is what it did.
None of that makes automatic captions useless. They are a good first draft and they save real time on the part of the job that is pure typing. The mistake is treating the output as the deliverable rather than as the raw material. What turns one into the other is a person reading the file against the video and fixing four things: the words that came out wrong, the sentence boundaries, the speaker changes, and the sounds that carry meaning.
For the first draft itself, SozAI transcribes audio and video files, recordings and YouTube links into text with speaker labels across 99+ languages in its iOS, Android and macOS app, and the site also carries free subtitle and counting tools that run entirely in your browser with nothing uploaded. Speaker labels give you a head start on the handovers; the sound cues and the final timing are still yours to add.
Budget for the pass. It is shorter than transcribing from scratch and longer than nothing, and it is the whole difference between a file that exists and a file that works.
What still goes wrong after you have done all this
Plenty, and it is better to know where.
- Names, jargon and product terms come back wrong from automatic transcription and stay wrong, because a reviewer who is not close to the subject does not notice that a term is slightly off.
- The caption file gets attached to the wrong cut of the video after an edit, and the sync goes with it. Nobody re-checks after a trim.
- The transcript quietly never gets produced, because it is the deliverable with no obvious slot in the publishing workflow and no error message when it is missing.
- A translated subtitle track gets uploaded as if it were a caption track, which leaves deaf viewers in the original language with dialogue only and no sound cues.
None of these show up in a spot check of the first thirty seconds, which is the check most videos actually get. If you are building a process, build the check at the end of the file and after the final edit, not before.
The other thing to expect is that no amount of care in the caption file settles what your obligation is. That comes from the rule your organisation is bound by, and the practical detail of how those requirements tend to be written up is covered in the piece on transcription and accessibility compliance. Get that answer first, then apply the craft above to whatever it turns out to require.

