Key Takeaways
Subtitling into a language you don’t speak is two jobs, not one. Get a transcript in the original language first, read it, fix it, and only then translate. A translation built on a flawed transcript carries the flaw forward into every language you produce. Cue timings belong to the speech, so they mostly survive the switch; what changes is how much text has to fit inside each cue. Check names by hand. Have a native reader look at the result if you possibly can.
You have a video and an audience that reads a language you don’t. The subtitles have to be right, and you have no way of telling whether they are. That’s an uncomfortable position, and most of the advice about it skips the part that actually decides the outcome.
The decision is about order. Almost everything that goes wrong with a translated subtitle track was already wrong before the translation started, and translation is very good at making an error look deliberate. What follows is the sequence that keeps the damage down, and an honest account of what you’re still exposed to at the end of it.
Two steps, and the order is not optional
Subtitling in another language is a transcript in the original language, followed by a translation of that transcript. That’s it. Every tool that offers to go from video to foreign-language subtitles in one move is doing both steps, just without showing you the middle one.
The problem with hiding the middle step is what happens to mistakes. If the transcription mishears a word, the translation step doesn’t know it was a mishearing. It translates it faithfully, as though it were correct. The output is grammatical, fluent, and wrong, and because it’s fluent there’s nothing in the text to catch a reader’s eye. A garbled transcript at least looks garbled. A clean translation of a garbled transcript looks fine.
So treat the transcript as a document you own. Produce it, open it, read it against the audio, fix it, and only then send it to be translated. If you’re publishing into three languages, this matters three times over, because all three are made from the same source file. The same applies to audio-only material, where translating a recording still means transcribing it first, whether or not you ever see that intermediate text.
One tool that keeps the middle step visible: SozAI transcribes video and audio files and YouTube links on iOS, Android and macOS, and will translate across 99+ languages once you’ve read and corrected the original. There are plenty of ways to get a transcript; the requirement is that you can see and edit it before anything is translated, not that it comes from any particular place.
Correcting the original is where the work pays off
Every minute you spend fixing the source transcript is a minute paid back in each language you output. Every error you leave in gets reproduced in each one. This is the only stage of the process where you have complete competence: it’s your language, your recording, your subject matter. Once you cross into translation, your ability to judge the output drops to roughly nothing.
Read the transcript with the audio playing rather than reading it cold. Cold reading catches things that look wrong. Reading against the audio catches things that look right and aren’t, which is the larger category. Technical terms, proper nouns and anything said quickly or over background noise are where the errors cluster.
Pay particular attention to sentences that make sense but sound slightly off for you. A transcript that says something you would never say is usually a transcript that got a word wrong. Fix it in the source. If you fix it after translation, you have to fix it once per language, and you have to explain the fix to someone who can read the language.
This is also the moment to decide how the subtitles should read as text rather than as speech. False starts, filler, repeated words: strip them here, in the language you understand, rather than letting a translator render them literally into a language where they read as incoherence.
The timings survive. The text inside them doesn’t
Cue boundaries are attached to the speech, not to the words. A cue starts when someone starts talking and ends at a pause, a sentence end or a speaker change, and translating the words doesn’t move any of those events. So the timing structure you build for the original mostly carries over intact, which is the one genuinely easy part of this.
What doesn’t carry over is length. Languages differ substantially in how much room the same meaning takes, and a sentence that sat comfortably on two lines in the source can come back needing three. When a translation overruns its cue, you have two options: tighten the text, or re-cut the cue to give it more time.
Tighten, in almost every case. Re-cutting a cue means it no longer lines up with the speech, and once one cue moves the ones around it usually have to move too. Do that in four languages and you have four different timing files to maintain instead of one. Tightening keeps a single timing structure across every version, which is worth a small loss of nuance in the wording.
Practically, that means asking whoever or whatever does the translation to fit a length, not just to translate. If you need the finished file in a different format for the platform you’re publishing to, a subtitle format converter that runs in your browser handles that without touching the timings. Building the cue file in the first place is a separate job, and generating a subtitle file from a transcript is where the cue boundaries get set.
What to check when you can’t read a word of it
You can’t check the translation. You can check specific items inside it, and there’s a category worth going through one by one: proper nouns. Names, places and product names are the things a translation is most likely to alter, and an altered name is wrong in a way viewers spot instantly, even viewers who can’t otherwise fault the subtitles.
Before you translate, write down the ones that appear in the video. After you translate, find each one in the output and confirm it survived. You don’t need to read the surrounding sentence to do this. You’re pattern-matching.
- People’s names, including anyone mentioned but not on screen
- Your company name and the names of any products
- Place names, especially ones that have a conventional local form
- Anything you spell out loud in the video, like a URL or a code
- Numbers and units, which sometimes get reformatted rather than translated
Where a name legitimately has a different form in the target language, you have a judgement call you’re not equipped to make, and that’s a question to put to a person rather than to a tool. Scripts that don’t share an alphabet with your source make the visual check harder but not impossible; you’re still looking for a recognisable token in a predictable position.
A native reader is the only real check
Someone who reads the target language natively, watching the video with the subtitles on, is the only quality control that actually works. Not a back-translation, which tends to launder errors into plausibility. Not a second machine pass. A person, watching, telling you whether it reads like something a human would say.
They don’t need to be a professional translator, and they don’t need to spend long. Most of what’s wrong with a machine-translated subtitle track is obvious to a native reader in the first minute, because it’s tonal rather than lexical: too formal, too literal, oddly stiff. Ask them for that impression rather than for a line-by-line audit and you’ll get useful information out of a short conversation.
When you can’t get one, the remaining move is upstream. Plainer source language produces a safer translation. Short sentences, ordinary vocabulary, no wordplay, no cultural references that only work in one place. If you already know a video will be subtitled into languages you can’t read, that knowledge should change how you speak on camera, not just how you edit afterwards. It’s a real constraint on the writing, and it costs you some personality. It also removes most of the ways the translation can embarrass you.
What you’re still exposed to
Idiom and humour are where machine translation is least reliable and where a wrong subtitle does the most damage. A mistranslated technical term is a small error that a viewer works around. A joke rendered literally changes how the whole video reads: it makes the speaker sound strange rather than funny, and the viewer has no way of knowing the strangeness came from the subtitle rather than from you.
Tone is the other loss, and it’s harder to see coming. Machine translation defaults to a register, and it may not be your register. A conversational video can come out sounding like a manual. Nothing in it is wrong. It just isn’t you, and the audience reading those subtitles will never meet the version of you that exists in the original audio.
None of this is fixable by trying harder at the translation step. It’s fixable by correcting the source, keeping the source plain, checking the names, and getting a native reader when you can. Do those things and you’ll ship a subtitle track that’s accurate and a little flat, which is a reasonable outcome for something you cannot read. Skip the source correction and you’ll ship errors you’ll never find, in as many languages as you published.

