Skip to content

Free preview

Interview Transcription Online

Drop an interview recording and get the first 5 minutes as text with the interviewer and participant labelled. No account, automatic language detection, file deleted after processing.

  • Free preview: first 5 minutes
  • Deleted after processing
  • No signup

Drop an interview recording here

or

Supports MP3, M4A, WAV, OGG/OPUS, AAC, FLAC and WEBM

How it works

Interview Transcription Online in three steps

01

Drop the interview audio

Phone recording, Zoom audio export or a dictaphone file. The browser keeps the first 5 minutes.

02

Speakers are labelled

Interviewer and participant get separate labels with timestamps. Language detected automatically.

03

Copy into your analysis

Download a .txt or copy the text. The full interview continues in the app.

From recording to quotable text

Whether it is a research interview for a thesis, a source for an article or a user interview for a product, the recording is only useful once it is text: coded, quoted, searched. Transcribing by hand takes four to six hours per hour of audio. This page shows what the automatic version looks like on your own recording, using the first 5 minutes, before you commit the rest to the app.

Who said what

Interviews are two-voice recordings, and the transcript reflects that: each turn is labelled Speaker A or Speaker B with a timestamp, so the interviewer's questions and the participant's answers stay apart. Three-person panels work the same way. The labels are consistent across the transcript — Speaker B at minute one is Speaker B at minute four — which is what makes coding by respondent possible.

What the preview does

Your browser reads the file and keeps the first 5 minutes, cutting longer recordings locally and re-encoding that part as 16 kHz mono WAV so the rest is never uploaded. The audio goes over an encrypted connection to the transcription engine behind the SozAI app (AssemblyAI). The language is detected automatically across 99 languages, punctuation is added, and filler words are lightly cleaned — a readable transcript rather than a strict verbatim one. When the text is delivered, the audio and transcript are deleted at the provider.

Formats and limits

  • M4A, MP3, WAV, AAC, FLAC, OGG/OPUS, WEBM, up to 25 MB as-is; longer recordings are trimmed to 5 minutes in the browser first (up to 30 MB on a computer, 12 MB on a phone).
  • Three previews per device per day, no account.
  • Phone calls recorded through a speaker and interviews over a café table both transcribe, with more errors on the quieter voice.

Consent and confidentiality

Recording a person requires their consent almost everywhere, and research ethics boards usually require it in writing. Transcribing does not change that. If your protocol restricts where audio may be processed, note that the preview is processed by a US-based provider unless the site has been set to its EU endpoint, and use the app's project settings for anything covered by a data agreement.

The whole interview

An hour-long interview is twelve previews' worth of audio. The SozAI app transcribes it in one go, keeps the speaker labels, produces a summary of themes, and stores every interview in a library you can search across. Import from Voice Memos, Files or a shared recording. The first 30 minutes are free.

Answers

Interview transcription — frequently asked questions

Does it identify who is speaking?

Yes. Each turn is labelled Speaker A, Speaker B and so on with a timestamp, and the labels stay consistent through the transcript. Names are not guessed; you rename the labels in your document or in the app.

Is the transcript verbatim?

It is a clean-read transcript: punctuation is added and obvious fillers are trimmed so the text reads naturally. If your method requires every 'um' and false start, note that in your protocol and expect to restore some of them by ear.

Which languages work?

The language is detected automatically across 99 languages. An interview conducted in one language with occasional words from another is transcribed in the dominant language.

How long a recording can I preview?

Any length, but you get the first 5 minutes. Files up to 25 MB are sent whole; longer recordings are cut to 5 minutes in your browser first (up to 30 MB on a computer, 12 MB on a phone).

Where is the audio processed and is it kept?

It is sent over an encrypted connection to the transcription provider (AssemblyAI, US region unless the EU endpoint is enabled), transcribed, and deleted there along with the transcript once delivered. This site stores neither.

Is this acceptable for research ethics purposes?

That depends on your board's rules about third-party processing. The preview is meant for a quick check of your own recording. For interviews covered by a data agreement, review the SozAI security page and use the app, where the full recording is handled under one account.

Can I transcribe a recorded phone or Zoom call?

Yes, as long as you have an audio file: Zoom and Teams can save an audio-only recording (.m4a), and call recorders produce MP3 or M4A. Two voices sharing one channel are still separated.

How accurate is it?

On a quiet, close-miked recording the transcript is close to verbatim. A phone on a café table between two people typically loses a few words from the quieter speaker. Speaker labels are most reliable when the voices differ clearly.

The whole recording, not just the first minutes

Record or import any length in the app: speaker labels, timestamps, a summary and translation into 100+ languages.

Get the App — Free

Free on iOS and Android. 30 minutes of transcription included, no card.