Key Takeaways
macOS will not transcribe a video file on its own. What it will do, with no downloads, is export the audio track out of the video using QuickTime Player, which leaves you a much smaller file to work with. Built-in dictation types what the microphone hears as you speak, live, so it is not the tool for a recording you already have. Transcription is the separate step, and it is the same step whatever computer you are on.
You have a video sitting on your Mac. A screen recording of a meeting, a lecture someone sent you, a clip you saved. You want the words. You would rather not create an account with a company you have never heard of to get them.
The job splits cleanly in two. First, get a clean audio file out of the video. Second, turn that audio into text. The first half your Mac can do already, right now, with software that is on the machine. The second half it cannot, and no amount of poking through System Settings will change that. Knowing where the line sits saves you an afternoon of looking for a menu item that does not exist.
What macOS already does
QuickTime Player ships with every Mac. It plays video, it records the screen, and it can export the audio from a video file. That last one is the useful part here, and it is often the only step you need to take before handing the file to something else.
macOS also has dictation, which turns speech into text as you talk. Voice Memos is there too, and it syncs with the same app on an iPhone over iCloud. Between the three of them you can capture sound, store it, and type with your voice.
None of them will read an existing video file and give you a transcript. That is worth stating flatly, because the tools sit close enough to each other that it feels like the feature ought to be hiding somewhere. It is not. The Mac is good at producing recordings and bad at converting them into text.
Pull the audio out first
A video file is a container. It carries picture and it carries an audio track, and transcription only ever touches the audio. The video is dead weight for this purpose.
So export the audio before you do anything else. QuickTime Player can do it, which means the first step costs you nothing and installs nothing. What you get back is a much smaller file than the video you started with — sometimes dramatically smaller, depending on what the video was.
That size difference stops being an abstraction the moment the recording is long. A two-hour lecture as video is an awkward thing to move around, upload, or keep several copies of while you work out which tool you are using. The same two hours as audio is a modest file you can handle without thinking about it. If you have a folder of recordings to work through, extracting first is the difference between an evening and a weekend.
The extraction also settles a question that trips people up: whether they need something that specifically handles video files turned into text, or whether plain audio transcription is enough. Once the audio is a separate file, the distinction stops mattering. Any transcription tool that takes audio will take yours.
Dictation is a different job
Dictation and transcription get used as if they were the same word. They are not, and the difference is exactly the thing standing between you and the transcript you want.
Dictation listens to a microphone and writes down what it hears, in real time, as it happens. You talk, text appears. It is a typing method. Transcription takes a recording that already exists — something that finished happening an hour or a year ago — and produces text from it. One is an input device, the other is a conversion.
You can try to bridge the gap by playing the video out loud and letting dictation listen to your speakers. People do this. It works badly. You are asking a microphone to re-record sound that has already been recorded once, in a room, at whatever the room sounds like, and you have to sit through the entire runtime while it happens. An hour of video costs you an hour. For anything longer than a short clip it is not a serious option. There is more to say about speech to text on a Mac and where the built-in path runs out, but the short version is that dictation was never built for files.
When you do not have the file yet
Everything above assumes the video is already on your disk. Sometimes it is not. The thing you want the words from is a call in progress, or a video playing in a browser tab, and there is nothing to extract audio from because nothing has been saved.
This is where the Mac puts up its real obstacle. Capturing your Mac’s own system audio — the sound coming out of the machine, rather than into the microphone — is not something macOS exposes by default. Screen recording gets you the picture. Getting the sound that went with it is a separate problem, and it is the single step that most often stops people cold. If you have ever ended up with a silent recording of a meeting and could not work out what you did wrong, this is what you ran into. You did not do anything wrong.
There are ways around it, and they are worth understanding before you need them rather than five minutes into a call you cannot repeat. The mechanics of recording the screen on a Mac with its audio are their own topic. What matters here is that it is a capture problem, not a transcription problem, and solving it gets you back to the ordinary case: a file on disk.
One route sidesteps it entirely. Voice Memos on a Mac shares recordings with the same app on an iPhone through iCloud, so anything you record on the phone shows up on the Mac without you emailing yourself a file or hunting for a cable. For an in-person conversation, or a talk you are sitting in, recording on the phone and picking the file up on the desktop is less fuss than fighting the system audio question. It does nothing for sound the Mac itself is producing, obviously.
Turning the audio into text
Once the audio exists as a file, none of the previous decisions matter any more. Transcription does not care whether the file came out of a video, a phone, a screen recording, or something a colleague sent you. It does not care that you are on a Mac. From here it is the same job it would be on any machine, which means your choice is about the transcription tool and nothing else.
This is also the point where the no-signup preference collides with reality, so let me be straight about it. Anything that runs entirely inside your browser can promise that your file never leaves the machine, and some tools genuinely work that way. But browser tools of that kind handle things like subtitle files and counting — text that is already text. Actual speech recognition on a two-hour recording is heavier work, and tools that do it generally want either an installed app or your file on their servers. Whichever you pick, that is the question to ask, and the answer should be easy to find.
SozAI is one option: a transcription app for macOS that handles audio and video files, recordings and YouTube links, with speaker labels, summaries and translation across 99+ languages. The same site has free browser-based subtitle and counting tools that run entirely in your browser with nothing uploaded and no account — those do not transcribe audio, but they are useful once you have a transcript to work on.
What the transcript will not do for you
Ask for timestamps if the point is to go back to a moment. A transcript without them is a wall of text you can search — you will find the sentence — but finding the sentence is not the same as finding the place in the video where it was said. If you are transcribing a lecture so you can jump to the ten minutes that mattered, timestamps are the whole feature. If you are transcribing so you can read it once and never open the video again, they are clutter. Decide which one you are doing before you start, because adding them afterwards means doing the work twice.
Expect the text to reflect the recording it came from. Speech recognition is working with whatever was actually captured, and a recording made across a room with people talking over each other gives it less to work with than a clean one. Extracting the audio from the video does not improve the audio; it just isolates what was already there. If the source sound is poor, you will be editing the transcript, and no step earlier in this process fixes that.
And accept that some of the video does not survive the trip. Slides, diagrams, anything written on a whiteboard, whoever was gesturing at what — all of it drops out, because you deliberately threw the picture away to get a file you could work with. For a conversation that loses nothing. For a demo or a presentation, the transcript is half the record, and you will want to keep the original video next to it.

