Transcribe a focus group discussion
Group discussions transcribed on your own machine, with participants’ words never leaving it — which is usually exactly what the consent form promised.
Focus groups are the hardest recordings in qualitative research, and this page is direct about what to expect and how to make it workable.
Language and translation need the Multilingual model.
Drop an audio or video file
or , paste with Ctrl + V, or
- MP3
- MP4
- WAV
- M4A
- MOV
- WEBM
- OGG
- FLAC
- AAC
The file is read by this page and never uploaded — no account, no size limit but your own RAM.
Preset: balanced English model, timestamped output.
Download
Click a line to jump the audio there. Click the text to fix a word — your edits go into every download.
How it works
- Open this page — the tool is already set up for “Focus groups”.
- Drop the file in, browse for it, or paste it with Ctrl+V.
- The model runs on your own device; the transcript appears as it goes.
- Correct anything misheard, then download text, subtitles or notes.
Several voices at once is the hard case
A focus group breaks most of the assumptions speech recognition relies on: six or eight people at varying distances from one microphone, frequent overlap, side conversations, and moments where three people agree at once. Where speech overlaps, no model produces a correct transcript — it emits one of the voices, or a blend of both, and gives no indication that it has done so.
The realistic expectation is a good transcript of the clearly-spoken parts and an unreliable one wherever people talked over each other. Those overlap points are often where the interesting disagreement happened, so they are worth listening to directly rather than trusting the text.
What actually improves the result
Recording matters far more than model choice. Two or three microphones placed among the participants, or a table array, dramatically outperform a single phone in the middle. A moderator who names people as they come in ("Go ahead, Priya") gives you attribution anchors that survive into the transcript and make manual labelling far quicker afterwards.
If you have separate microphone tracks, transcribe each one on its own. That is the only way to get genuinely reliable speaker attribution in a group, and it is better than any automatic diarization — including the paid kind.
Coding group data
The CSV export gives one row per segment with start and end times, speaker and text, escaped correctly for Excel. That is the natural unit for coding turns in a group discussion, and it imports into NVivo, MAXQDA, Dedoose and Taguette.
Label speakers before exporting. In a group transcript, unattributed turns lose most of their analytic value — you cannot see who converged with whom, which is often the whole point of running a group rather than individual interviews.
Things that save a re-run
- Use more than one microphone if you possibly can — it helps more than anything else.
- Ask the moderator to name participants when inviting them to speak; it makes labelling much faster.
- Listen to every overlap point directly; the transcript is least reliable exactly where the discussion got interesting.
- Separate per-participant tracks give real speaker separation — transcribe each individually.
Read the transcript against the audio before you rely on it. Every speech model — this one and the paid cloud ones — mishears names, numbers and crosstalk, and it does so confidently. The player and editor here exist so that check takes minutes rather than an afternoon.
Frequently asked questions
Can it tell six participants apart?
No. Whisper has no speaker diarization, and automatic diarization of a six-person group with overlap is unreliable even in expensive tools. Label manually, and use separate microphone tracks where you can — that is the only genuinely reliable approach.
How much of a group discussion will be wrong?
It depends almost entirely on the recording. Clean turn-taking with good microphones transcribes well. Anywhere two people speak at once, expect the transcript to be wrong without flagging it — check those moments against the audio.
Is participant audio uploaded anywhere?
No. The recording is read and processed in your browser, and the site has no upload endpoint. Nothing is transmitted, so no third party receives participants’ words.
What is the best export for coding?
CSV — one row per segment with times, speaker and text. It imports into every major qualitative analysis package and into a spreadsheet.