Transcribing a focus group means converting a multi-speaker recording into a text transcript that correctly attributes each line to the right participant, including moments where people talk over each other. It is meaningfully harder than transcribing a one-on-one interview: most transcription tools are built and tested on single-speaker audio, and accuracy drops when speech overlaps. Tools like Skimle handle multi-speaker attribution as part of the same platform used for analysis, so a transcript is ready to code as soon as it is produced.
Why is focus group transcription harder than interview transcription?
A one-on-one interview transcript has one obvious complication: making sure the words are right. A focus group transcript has that same problem plus a second, harder one: making sure every line is attributed to the correct person, in a recording where multiple people frequently speak at once, interrupt each other, laugh together, or trail off mid-sentence when someone else jumps in.
Automatic speech recognition systems, the technology behind every AI transcription tool, are trained and benchmarked overwhelmingly on clean, single-speaker audio. A focus group violates that assumption structurally: it is multiple voices, inconsistent distance from the microphone depending on where someone is sitting, and genuine overlapping speech that a system built for one voice at a time was never optimised to separate.
The size of that gap is measurable. In a 2026 survey of multi-speaker speech recognition research, the best-performing systems reached a word error rate of around 3.4 to 4.9 on controlled two-speaker benchmark audio, but 21.2 on AMI, a corpus of real meeting recordings with three to five speakers. In plain terms: roughly one word in five is wrong on realistic multi-speaker recordings, against roughly one in twenty-five on clean two-speaker audio. A focus group is much closer to the meeting recording than to the benchmark.
That figure is worth carrying into your planning. It does not mean AI transcription is unusable for focus groups, but it does mean that treating an unreviewed automated focus group transcript as clean data is a mistake, in a way that is less true for a well-recorded one-on-one interview.
5 challenges specific to focus group transcription
1. Speaker attribution when participants aren't formally introduced
In a well-run session, the moderator asks each participant to state their name before their first response, which gives a transcription tool (or a human transcriber) a clean reference point. Many real sessions skip this, especially when the group already knows each other or time is tight, which makes attributing later statements to the right person far harder after the fact.
2. Overlapping speech and crosstalk
Two or three people responding to a provocative comment at the same time is exactly the kind of group energy a good moderator wants to capture, and exactly the audio a transcription system struggles with most. Overlapping segments frequently come out garbled, merged into one speaker's line, or dropped entirely.
3. Variable audio quality across a room
A single lapel mic on the moderator, plus a room mic for participants, means some voices are consistently louder and clearer than others depending on where people are sitting. Quieter or more distant participants are more likely to be misattributed or transcribed inaccurately, which introduces a bias into your transcript before analysis even starts: the quietest voice in the room becomes the least reliably captured one.
4. Non-verbal group dynamics don't survive in text
A pause where the room goes quiet after someone makes an uncomfortable admission, a round of laughter that shifts the tone, someone visibly disagreeing without speaking up: none of this appears in a transcript unless the moderator or a note-taker captures it separately. Focus group analysis depends on these dynamics as much as the words themselves, so losing them at the transcription stage is a real cost, not a minor detail.
5. Transcript length and group size scale together
An eight-person, 90-minute focus group produces a transcript that is both longer and harder to read than an equivalent one-on-one interview, because a reader has to track who is speaking throughout rather than following a single voice. This compounds the review burden: checking a focus group transcript for accuracy simply takes more careful reading than checking an IDI transcript of the same length.
What are your transcription options in 2026?
General-purpose AI transcription tools like Otter.ai, Fireflies.ai, and OpenAI Whisper all include some form of speaker diarisation (separating and labelling different speakers), with accuracy that is good on two-person conversations and noticeably weaker as speaker count and overlap increase. Our comparison of AI transcription tools covers accuracy, pricing, and GDPR considerations for each in detail.
Skimle includes built-in transcription as part of the same platform used for analysis, so a focus group recording goes directly into a project and comes out as a transcript ready to code, without exporting from one tool and importing into another. For a market researcher running several focus groups a week, removing that handoff step is often a bigger practical win than a marginal accuracy improvement on any single recording. Credits cover both transcription and analysis, from €20 ($22) per month with a free trial.
How do you prepare a focus group recording for good transcription?
A few moderator habits meaningfully improve transcription quality before a single word is processed:
- Ask every participant to state their name before their first response. This gives any transcription tool, human or automated, a clean anchor for the rest of the session.
- Establish a "one person at a time" norm early, and have the moderator gently enforce it. Groups that self-regulate crosstalk produce measurably better transcripts than groups that don't.
- Use a dedicated microphone setup for the room, not just the moderator's mic, so quieter participants are captured at a usable volume.
- Record video alongside audio where possible. Even if you transcribe from audio, having video available lets you resolve ambiguous attribution or unclear passages by watching who was speaking.
What does focus group transcription cost?
Worth budgeting explicitly, because group sessions are longer than interviews and the review burden is higher.
| Approach | Typical 2026 cost | Realistic for focus groups? |
|---|---|---|
| Human transcription service | From roughly $1.99 (€1.83) per audio minute, so about $180 (€165) for a 90-minute group | Highest accuracy, including on crosstalk; cost adds up fast across a multi-group study |
| Standalone AI transcription | From roughly $0.25 (€0.23) per minute, or a flat monthly subscription | Cheap and fast, but expect meaningful speaker-attribution correction |
| AI transcription bundled with analysis | Included in platform pricing (Skimle from €20 / $22 per month, credits covering transcription and analysis) | Removes the export step between transcription and coding |
A six-group study at 90 minutes each is nine hours of audio. Human transcription at the rate above would be roughly $1,075 (€985) for that study; AI transcription would be a small fraction of it, with the difference reappearing as analyst review time. The comparison that matters is therefore cost plus review hours rather than cost alone, and the answer shifts depending on how much crosstalk your sessions contain and how senior the person doing the review is.
Where transcription and coding happen in two different tools, that review time is compounded by an export step. Skimle transcribes the recording and opens it straight into the analysis environment, so a finished transcript is already a codeable one. See how that fits market research and customer insights teams.
What should you check in a focus group transcript before analysis?
Before coding begins, read through (or spot-check, for a long transcript) with three questions in mind: are speakers attributed correctly, especially in sections with crosstalk; are unclear or inaudible passages clearly marked rather than silently guessed at; and does the transcript preserve enough structure (pauses, overlapping speech, laughter) to support the kind of group-dynamics reading that focus group analysis and how to analyse focus group transcripts both depend on.
Transcription quality genuinely determines analysis quality here. A transcript that silently misattributes a third of a group's statements will produce a coding pass that reflects the transcription errors as much as what participants actually said, and that error is invisible unless someone checks the transcript against the recording directly.
What about consent and recording ethics?
Group recordings raise a consent wrinkle that one-on-one interviews don't: every participant needs to consent not just to being recorded, but to appearing in a recording alongside other participants whose privacy also matters. Best practice is to get explicit written consent before the session starts, state clearly how the recording and transcript will be used and who will see them, and confirm whether participants can be identified by name in any output or only by role or segment.
Once a transcript exists, anonymising or pseudonymising it before it circulates beyond the immediate research team is good practice for any study involving personal opinions, and often a compliance requirement under GDPR when the data includes anything that could identify a participant. This matters more for focus groups than IDIs in one specific way: participants in a group setting often reference each other by name or reveal identifying details about one another that they would not disclose about themselves in a one-on-one interview, so a transcript can contain identifying information about someone who never directly consented to that detail being recorded.
What about focus groups in other languages?
Multi-market studies add a decision most transcription guides skip: whether to transcribe in the source language or transcribe-and-translate in one step.
Transcribing in the source language first, then translating, keeps the original available for anyone who needs to check what a participant actually said, and preserves idiom and hedging that translation tends to flatten. Going straight to a translated English transcript is faster and cheaper but means every subsequent coding decision rests on an interpretation nobody can audit without going back to the audio.
For a study where findings will be scrutinised, keep the source-language transcript as the record of what was said. Skimle analyses documents in multiple languages within one project, so transcripts from different markets can be coded against a single shared category structure without translating everything into one language first, which removes the trade-off in most cases.
Speaker attribution also degrades further in languages with different phonetic structures to the (largely English) data most speech recognition models were trained on, so budget more review time per group for non-English sessions, not less.
Frequently asked questions
How accurate is AI transcription for focus groups?
Meaningfully less accurate than for one-on-one interviews. Published benchmarks put the best multi-speaker systems at roughly 3.4 to 4.9 word error rate on controlled two-speaker audio but around 21.2 on real meeting recordings with three to five speakers, so expect roughly one word in five to need attention on realistic group audio before review.
How long does it take to transcribe a 90-minute focus group?
With AI transcription, a 90-minute recording typically processes in a few minutes, though reviewing and correcting speaker attribution afterwards can take 30 to 60 minutes for a group of six to eight participants, more if the session had significant crosstalk.
Can AI accurately distinguish speakers in a group setting?
AI speaker diarisation has improved substantially but remains meaningfully less accurate on group recordings than on one-on-one interviews, particularly during overlapping speech. Most researchers should expect to spend some time correcting speaker attribution on a focus group transcript, even with a good tool.
Should you clean up filler words and false starts in a focus group transcript?
Keep a verbatim version for analysis, since hesitations, false starts, and filler words can carry meaning (uncertainty, discomfort, backing away from a strong opinion). A cleaned-up version is fine for a client-facing quote in a report, but the working transcript used for coding should stay as close to verbatim as possible.
Do you need video, or is audio enough for focus group transcription?
Audio alone is usually enough for a well-run session where the moderator manages turn-taking and introductions carefully. Video becomes valuable specifically for resolving ambiguous attribution and capturing non-verbal reactions that audio alone misses.
Is it worth transcribing a focus group at all, or can you analyse it from moderator notes?
Moderator notes and a debrief memo capture the moderator's impressions, which is useful but is not the same as the actual data. A full transcript lets you, or anyone else reviewing the study later, check whether a reported finding actually matches what participants said, which matters for any study where a client or stakeholder might ask "show me where this came from."
Ready to transcribe and analyse your focus groups in one platform? Try Skimle for free and see how transcription and coding work together without a separate export step.
Related reading: Focus group analysis: a complete guide, how to analyse focus group transcripts, and the best AI transcription tools for research in 2026. If you work in market research or customer insights, see how Skimle fits your workflow specifically.
About the authors
Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile



