Skip to content

How to Transcribe an Interview: Verbatim Levels, Speaker Labels and Notation

•9 min read• 22 views •Last updated: Oct 2, 2026
Two chairs facing each other at a small table by a window, a phone recording between them next to a notebook

Key Takeaways

Decide the verbatim level before you type a word: full verbatim keeps every um and false start, clean verbatim keeps the meaning without the fillers, and an edited transcript tidies the grammar for reading. Pick a notation for pauses, overlaps and inaudible stretches, write it as a key at the top of the document, and label speakers consistently. By hand, the Texas Tech University Libraries oral history guide puts the ratio at six hours of typing per hour of recording; a Cornell University course guide says four to six. An automatic transcript gives you a first draft in minutes; the human work shifts from typing to checking it against the recording.

You’ve got the recording. Now someone wants a transcript they can quote, code or hand in, and the first question isn’t which app to use – it’s how much of the speech you’re going to keep.

Fillers and false starts matter in conversation analysis and some legal contexts; they’re noise everywhere else. This article goes through the three verbatim levels, the notation that marks pauses and overlaps, how long the job takes by hand, and what an automatic transcript gets right and wrong on a real recording.

What is the easiest way to transcribe an interview?

Run the recording through a speech-to-text engine first, then fix names, spot-check against the audio and add whatever notation your project needs – that’s faster than typing from a blank page, and it’s what most people actually mean by easiest. Typing a full hour of speech from scratch, word for word, is the slow way in, not the easy one.

For a short clip, our browser tool at transcribe an interview online does that first pass with no account: upload a recording and it transcribes the first 5 minutes, labels each turn Speaker A, Speaker B and so on with a timestamp, and detects the language automatically across 99 languages. Files up to 25 MB go through as they are; longer recordings get cut to 5 minutes in the browser before anything is sent. The audio and the transcript are deleted at the provider once the result comes back, and you get three previews per device per day.

That’s enough to see whether a tool is worth using on a given recording. For a whole interview in one pass, with turns labelled all the way through and every transcript kept in a searchable library afterwards, the SozAI app does that – the first 30 minutes are free. See how interview transcription works for the full picture, or download the app.

What is the difference between verbatim and clean verbatim?

Full verbatim keeps every word as spoken: the ums, the false starts, the repeated words, the stutters and the audible laughs. Clean verbatim keeps the speaker’s words and meaning but drops the fillers, the false starts and the accidental repetitions. An edited transcript goes a step further and tidies grammar and sentence order so it reads well – which is exactly why it stops being usable as evidence of what was actually said.

  • Full verbatim: conversation analysis, some legal and evidential work.
  • Clean verbatim: most qualitative research, journalism, user research.
  • Edited: published question and answer articles.
Three verbatim levels and what each keeps
LevelWhat it keepsTypical use
Full verbatimEvery word as spoken: fillers, false starts, repetitions, stutters, audible soundsConversation analysis, some legal and evidential work
Clean verbatimThe speaker’s words and meaning, fillers and false starts droppedMost qualitative research, journalism, user research
EditedTidied grammar and sentence order for readingPublished question and answer articles

If the transcript is going into a research project rather than a published piece, decide this before the first interview, not after the fifth – switching levels partway through means redoing earlier transcripts to match. The research-specific side, consent, anonymisation and storage, is covered in transcribing interviews for research.

How do you write a transcript of an interview?

Start with a key: the verbatim level, the symbols you’re using, and how speakers are labelled, written at the top of the document before the first line of speech. Anyone who codes or quotes the transcript later needs that key, and so will you, weeks on, when you’ve forgotten what (.) was supposed to mean.

The best-known system for detailed transcripts comes from Gail Jefferson, set out in her Glossary of transcript symbols with an introduction (2004). As summarised by Clift, Kendrick, Raymond and Robinson (2024), a left square bracket marks the point where overlapping talk begins, an equals sign joins talk with no gap between it, a number in parentheses such as (0.5) is a timed silence in seconds, and (.) marks a pause too short to time. Sources disagree on exactly how short a micropause is, so if your key uses (.) it should say what you mean by it.

Most interview projects don’t need that level of detail and are easier to read without it. A bracketed convention covers what actually comes up: [inaudible 00:12:31] for a passage you can’t make out, with a timestamp so you can find it again; [crosstalk] where two people talk over each other; [laughs] and other non-speech events in square brackets; and a speaker label at the start of every turn.

Transcript notation: Jeffersonian symbols and a lighter bracketed convention
SymbolMeaningSystem
[Point where overlapping talk beginsJeffersonian
=Talk joined with no gap (latching)Jeffersonian
(0.5)Timed silence, in secondsJeffersonian
(.)Pause too short to timeJeffersonian
[inaudible 00:12:31]Passage that can’t be made out, with a timestampBracketed
[crosstalk]Two people talking at onceBracketed
[laughs]Non-speech soundBracketed

Beyond notation, the UK Data Service guidance on transcription says a transcript should carry a unique identifier, use a uniform layout throughout the project, use speaker tags for turn-taking, and keep line breaks – small things, but they’re what let you find and compare transcripts later instead of re-reading all of them.

Speaker labels are names the transcription software can’t know – expect Speaker A and Speaker B, which you rename by hand once you know who’s who. Labelling speakers consistently goes through how to do that across a whole project without the names drifting between files.

Anonymise as you go, not afterwards: replace names and identifying details with consistent placeholders such as [Participant 2] or [city] right in the transcript, and keep the key that maps placeholders to real names in a separate file. Recording a person needs their consent almost everywhere, and a research ethics board will usually want that consent in writing; transcribing the recording doesn’t change that requirement.

How long does it take to transcribe an interview?

By hand, plan for hours, not minutes. The Texas Tech University Libraries oral history guide says even experienced transcribers typically work at six hours of typing for every hour of audio. A Cornell University course guide puts it lower, at four to six hours per hour of recording – the gap is mostly down to how clean the audio is and how much notation gets added along the way.

An automatic transcript changes where the time goes rather than making it disappear. The draft comes back in minutes; what’s left is listening through it against the recording, fixing names and technical terms, correcting who said what where the labels crossed over, and adding whatever notation the project needs. For a one-hour interview that’s still real work, just a different kind of it – how long transcription actually takes breaks that down further.

Can ChatGPT transcribe an interview?

Whether a chatbot can turn an audio file into text depends on the product, but an interview transcript needs more than text: timestamps you can go back to and speaker turns you can trust, which is what a transcription engine is built to produce. Where a chatbot earns its place is after that step, tidying, summarising or reformatting a transcript you already have.

Either way, automatic text has to be checked against the recording before you quote it, and that check is the part no tool removes.

How much checking a given engine needs varies quite a bit. How accurate AI transcription is covers what to expect and where the errors tend to cluster.

What a real automatic transcript gets wrong

An automatic transcript is a draft, and the clearest way to see what that means is to run one and read it against the recording. On 2 October 2026 we did exactly that, on the first 5 minutes of a public recording that anyone can check against this article.

We ran the engine behind SozAI’s interview tool on the first 5 minutes (300 seconds) of NASA Johnson Space Center’s podcast Houston We Have a Podcast, episode 180, ‘Artemis Mission Management’, host Gary Jordan and guest Mike Sarafin – a NASA work, so not protected by copyright. It came back with 778 words, detected English, and found 2 speakers. The opening lines, as the engine rendered them:

[00:00] Speaker A: Houston, we have a podcast. Welcome to the official podcast of the NASA Johnson Space Center, episode 180, Artemis Mission Management. I'm Gary Jordan, and I'll be your host today.
[01:19] Speaker B: T-minus five seconds and counting. Mark.
[01:22] Speaker A: Launch commit lights are correct. There she goes. Houston. We have a podcast. Mike Serafin, thanks for coming on Houston, We Have a Podcast today.
[01:36] Speaker B: Thank you. It's a pleasure to be here to talk about these exciting Artemis missions we have ahead of us. Yeah, right?
[01:43] Speaker A: Returning to the Moon. Not a bad thing to be a part of.

Three things in just those five turns needed a human pass. The launch countdown at 01:19 and ‘Launch commit lights are correct. There she goes.’ at 01:22 are archive launch audio spliced into the podcast, not the host or the guest speaking – the engine labelled them Speaker B and Speaker A anyway. ‘Yeah, right?’ at the end of the 01:36 turn is actually the host picking the conversation back up; it belongs at the start of the next turn, not the end of the guest’s. And the guest’s surname comes out as Serafin here, though the episode’s own credits spell it Sarafin.

A later turn shows what clean verbatim is for. The engine’s output at 02:51 reads: ‘Yes, so you’re, you’re right in that I lived in the operations realm for 22 of my 27 years in, in my NASA career.’ A clean verbatim pass turns that into: ‘You’re right in that I lived in the operations realm for 22 of my 27 years in my NASA career.’ Full verbatim would keep the repetitions exactly as they came out; the engine, if anything, leans toward keeping them – it held onto most of the repeats and false starts across the file (‘you’re, you’re right’, ‘in, in my NASA career’, ‘That’s, that’s quite a number of’), which puts its output closer to verbatim than to clean verbatim. If a project needs clean verbatim, budget for an editing pass no matter which engine produces the draft.

Answers

Frequently Asked Questions

What is the easiest way to transcribe an interview?

Run the recording through a speech-to-text engine for a first draft, then check it against the audio and fix speaker names and unclear words - typing it from scratch by hand is slower, not easier. A browser tool like SozAI's interview transcriber can produce a labelled, timestamped draft of the first 5 minutes with no account, which is enough to judge whether it's worth running the whole file.

What should a finished interview transcript look like?

It opens with a short key stating the verbatim level, the notation used and how speakers are labelled, followed by labelled turns such as Speaker A, with timestamps. The UK Data Service guidance on transcription adds that it should carry a unique identifier, a uniform layout, speaker tags and line breaks, so it can be found and compared later rather than re-read in full.

Can ChatGPT transcribe an interview?

It depends on the product, and it is the wrong tool for the first step. An interview transcript needs timestamps and speaker turns you can rely on, which is what a transcription engine produces. A chatbot is useful afterwards, for tidying, summarising or reformatting a transcript you already have, and any automatic text still has to be checked against the recording.

Should I use verbatim or clean verbatim?

Use full verbatim for conversation analysis or legal and evidential work, where every um, false start and repetition matters; use clean verbatim for most qualitative research, journalism and user research, where the meaning matters more than the disfluencies. Decide before the first interview, since switching partway through means redoing earlier transcripts.

How do you mark inaudible parts in a transcript?

Write [inaudible] with a timestamp, like [inaudible 00:12:31], so you or anyone else can find that spot in the recording again later. For overlapping speech use [crosstalk], and for non-speech sounds like laughing use [laughs] - all in square brackets, explained once in the key at the top of the transcript.

How long does it take to transcribe a one-hour interview?

By hand, the Texas Tech University Libraries oral history guide puts it at around six hours of typing per hour of audio; a Cornell University course guide estimates four to six hours. An automatic transcript returns a draft in minutes, shifting the remaining time to checking it against the recording rather than typing it out.

Merey Tleugazin

Founder of SozAI. Building tools that turn speech into text for professionals worldwide.

SozAI
SozAI — Free DownloadTranscribe audio & video instantly
Get App