Key Takeaways
Decide the level of detail before you transcribe anything. If your analysis is about what people said, a clean readable transcript is enough. If it is about how they said it, you need pauses, overlaps and hesitations preserved, and automatic transcription strips exactly those. Check what your consent form and your ethics approval say about where recordings may go before you upload anything anywhere. Anonymise in the transcript rather than the recording, keep any mapping separate, and use timestamps so quotations stay traceable.
Nobody tells you the standard because there isn’t one standard. There are several, and which one applies to you depends on what you plan to do with the transcripts once you have them. That decision is methodological, and it belongs before the tooling decision, not after it.
Get it the wrong way round and the cost is real. You transcribe six interviews cleanly, get into your analysis, discover you need to know where the participant hesitated, and now you’re going back through six recordings with different ears. So spend an hour on the question first.
Decide what the transcript is for before you transcribe a word
How much detail a transcript needs is a methodological decision, not a matter of care or thoroughness. It follows from your analytic approach.
If your analysis is concerned with what was said, a clean transcript is fine. You want the words, accurately, attributed to the right speaker. Filler words, false starts and short pauses can be tidied out without damaging anything, because none of them will ever appear in your coding. Most thematic and content-oriented work sits here, and a clean transcript is quicker to produce, quicker to read and quicker to code.
If your analysis is concerned with how something was said, the situation reverses. Pauses, overlaps, hesitations, the moment where two people talk over each other, the three seconds before an answer arrives, the repair mid-sentence: those are the data. A transcript that removes them has removed the thing you’re studying. This matters for automatic transcription specifically, because smoothing out exactly those features is what it does well. It gives you fluent, readable prose, and fluent readable prose is a lossy rendering of a hesitant, overlapping conversation.
There’s a temptation to hedge by transcribing everything to the finest level of detail, on the grounds that you can always ignore what you don’t need. Usually that’s the wrong move. Detailed transcription is slow, and slow work on twenty interviews eats months you don’t have. It also makes the transcripts harder to read, which makes coding slower. If you genuinely don’t know yet whether interaction matters to your analysis, that’s a conversation to have with your supervisor now, not a problem to solve by brute force. Whatever you decide, the recordings stay, and the recording is the actual data. The transcript is a representation of it.
Write your convention down before the second interview
Once you’ve decided the level of detail, fix the conventions and write them into a short document you keep beside you while you work. Changing them later means going back through every recording you’ve already done, and that’s the kind of task that quietly consumes a fortnight.
The things worth settling in advance:
- Whether filler words and false starts are kept, tidied, or kept only where they seem meaningful, which is the option most likely to produce inconsistency across interviews.
- How pauses are marked, and whether short and long pauses are distinguished from each other.
- How overlapping speech is shown, and what happens when you can’t tell who spoke first.
- How you mark passages you can’t make out, and whether you note the reason.
- How speakers are labelled, and how that labelling relates to your anonymisation scheme.
- Where timestamps go: at a regular interval, at each speaker change, or only against passages you expect to quote.
Timestamps deserve a moment of their own. They let a quotation in your thesis be traced back to its position in the recording, which is what makes the claim checkable by anyone who has access to your data, including an examiner. Adding them as you go costs almost nothing. Adding them afterwards, to a finished transcript, means listening to the whole recording again with a stopwatch. Whichever tool you use, get them in on the first pass.
Consent and ethics approval decide where the audio can go
Your consent form is not paperwork that happened before the research started. It is a set of conditions on what you may now do with the recordings, and it usually covers three things: how the recording is stored, who is allowed to hear it, and when it gets destroyed.
Sending audio to a third-party transcription service is a disclosure. Someone or something outside your project now has a copy of a participant’s voice. That has to be compatible with what the participant agreed to. Sometimes it plainly is, because the form anticipated it. Sometimes it plainly isn’t, because the participant was told only the research team would hear the recording. Often the form is ambiguous, and that ambiguity is worth resolving with your supervisor or ethics committee before you upload rather than after.
Ethics approval frequently goes further and specifies where recordings may be stored, on what kind of device or institutional system, and for how long. The transcript inherits those conditions. So does anything derived from it: your coded files, your quotation extracts, the working document where you paste passages while drafting. Those all count as data, and they all belong inside whatever storage arrangement you were approved for. It’s easy to lose track of this once you’re deep in analysis and copies start multiplying across a laptop, a cloud drive and an email to yourself.
Practically, this narrows your options for tooling before you’ve compared any of them on quality or price. A service that processes files on your own device is a different proposition from one that uploads them, and if you’re weighing that up, it’s worth reading how a given tool handles your files and data rather than assuming. If the answer isn’t clearly stated, treat that as an answer.
Anonymisation happens in the transcript
The recording stays as it is. You can’t anonymise a voice by editing text, and you shouldn’t be trying to alter your primary data anyway. Anonymisation is something you do to the transcript.
That means replacing names, place names, employer names, job titles that identify one person, and the small circumstantial details that make someone recognisable to anyone who knows the setting. Do it consistently. If a colleague is [Colleague A] in one interview, she is [Colleague A] everywhere she appears, and the same convention applies across your whole set. Inconsistent pseudonymisation is worse than none, because it looks anonymised while still being traceable by someone who reads carefully.
If you keep a mapping between real identities and replacements, keep it separately from the transcripts, under the storage conditions your approval sets out, and be clear with yourself about why you’re keeping it and when it gets destroyed. In some projects there’s a good reason to hold it: you may need to re-contact participants, or link interviews across waves. In others there isn’t, and the safest mapping is the one that doesn’t exist. Deciding deliberately is the point.
One caution about doing this at speed. Find-and-replace catches the obvious names and misses the rest: the nickname used twice, the misspelling, the description of a building that only one place in the city matches. The anonymisation pass wants human eyes on the whole document.
Where automatic transcription actually helps
Automatic transcription produces a first draft. That’s a genuine saving, and refusing it on principle mostly means typing out several hours of speech by hand for no methodological gain. What it doesn’t produce is a finished transcript.
For interview data, the correction pass against the audio is part of the method, not an optional tidy-up at the end. You listen through with the draft in front of you, fix what’s wrong, restore whatever your convention requires that the machine dropped, insert or check the timestamps, and apply your anonymisation. It’s slower than reading and faster than typing. It’s also where you become familiar with the material, which is not a side benefit. Researchers who transcribe their own interviews arrive at analysis already knowing what’s in them, and the correction pass buys back a decent share of that familiarity.
If you want a draft without sending files off to a server, the SozAI app for iOS, Android and macOS transcribes recordings and files with speaker labels, which handles the mechanical part of an interview transcript and leaves you the pass that matters; more on how that applies to transcribing interviews specifically. It’s one option among several, and the storage question above should filter your shortlist before quality does.
What the first draft won’t do for you
Expect the draft to be weakest exactly where interviews are hardest. Crosstalk confuses speaker separation, and two people talking at once tends to come out as one person saying something slightly incoherent. Strong accents, quiet participants, background noise from a cafe or a corridor, and field-specific vocabulary all reduce accuracy. Speaker labels are usually close but not reliable at the boundaries, particularly around short interjections.
And it will not give you prosody. No automatic transcript preserves the length of a pause or marks the point where the participant’s voice changed. If your analysis needs that, the machine draft is a scaffold for the words and nothing more, and you should budget your time accordingly. A colleague working on the same corpus with a content-oriented question will finish in a fraction of the time, and that’s not because they’re working less carefully. Their transcript is answering a different question.
Budget more time than the arithmetic suggests, too. The correction pass is not one pass through the audio at normal speed; it’s stopping, rewinding, re-listening to the six seconds you can’t quite make out. Doing this on a full set of interviews is substantial work, and it is normal for it to be. If you want a sense of how the pieces fit together across a whole project, there’s a walkthrough of a thesis researcher’s workflow that follows one set of recordings from file to quotable transcript.
The last thing to hold onto is that the transcript is a representation, and every convention you chose is a decision about what to keep and what to lose. Write those decisions into your methods chapter. An examiner who knows why your transcripts look the way they do will read them as considered. One who doesn’t will wonder.

