To code interview transcripts: read each transcript closely, attach short descriptive labels (codes) to meaningful segments, collect those codes into a codebook, apply them consistently across every transcript, and cluster related codes into higher-level themes. Keep each code linked to its source quote so findings stay traceable. Tools like Skimle automate the first pass across large sets.
Coding is how a pile of interview transcripts becomes analysable evidence. Done well, it turns scattered quotes into a structure you can compare, count, and defend. Done casually, it produces a tidy-looking set of themes that nobody can trace back to the data. This guide walks through the process from first read to finished code structure, first as a manual method any researcher can follow, then with a note on where AI-assisted tools fit once the transcript count grows.
What does coding an interview transcript mean?
Coding means attaching a short label to a specific segment of a transcript so that its meaning can be retrieved and compared later. The label is the code. When several codes share a deeper meaning, you group them into a theme.
The distinction between the two trips up a lot of beginners, so it is worth being precise.
| Concept | What it is | Example |
|---|---|---|
| Code | A short, descriptive label on a specific excerpt | "waiting for sign-off", "unclear owner" |
| Theme | An interpretive pattern grouping several codes | "Decision-making friction" |
| Codebook | The organised list of codes with definitions | The full scheme you apply to every transcript |
Codes are descriptive and stay close to the data. Themes are interpretive and sit above the codes. If you want the wider methodological grounding, our complete guide to thematic analysis covers how codes become themes across Braun and Clarke's six phases, the most widely used framework in the field.
How do you code an interview transcript step by step?
Step 1: Prepare and read the transcript
Clean transcripts first: consistent speaker labels, timestamps if useful, and no obvious transcription errors. Then read the whole transcript once without coding anything. This first pass, what Braun and Clarke call familiarisation, is where you start to notice what the interview is really about. Resist the urge to label as you read the first time.
Step 2: Do a first pass of open coding
Now go through slowly and attach a short code to each meaningful segment. Stay descriptive rather than interpretive at this stage. If a respondent says "I kept waiting for someone to make a decision and it never came," code it "decision paralysis" or "unclear ownership," not "poor governance," which is an interpretation you may not end up standing behind. Descriptive codes at first pass keep you close to the data and stop you over-interpreting too early.
Step 3: Build and define your codebook
As codes accumulate, organise them into a codebook with a short definition for each. A definition matters more than it sounds: without one, the same code drifts in meaning between interview 3 and interview 30. Decide whether your codes are:
- Inductive, emerging from the data, or
- Deductive, brought from an existing framework or prior literature.
Most projects are abductive, moving between the two. See inductive, deductive and abductive coding for when to use each, and qualitative coding for open, axial and selective coding techniques.
Step 4: Apply codes consistently across every transcript
This is the laborious part and the part that determines whether your analysis holds up. Apply the codebook to every transcript the same way. When you meet something the codebook does not cover, add a new code and, crucially, go back and check whether earlier transcripts also contained it. Consistency is what lets you later say "eleven of twenty participants raised this" with a straight face.
Step 5: Cluster codes into themes
Step back and group codes that share a meaning into candidate themes. A good theme is coherent (its codes belong together) and distinct (it does not overlap heavily with another). Most published analyses settle on between three and seven themes; fewer suggests under-analysis, more usually means themes need consolidating.
Step 6: Check reliability and trace back to source
Before you report anything, trace each theme back to the specific excerpts that support it. If you code as a team, check inter-coder agreement on a subset and reconcile disagreements by refining code definitions. A finding you cannot trace to a quote is not yet a finding. For the full workflow from transcript to synthesis, see how to analyse interview transcripts.
Where does manual coding start to break down?
Manual coding is a sound method and a good way to learn. It also does not scale gracefully. Coding a single hour-long interview thoroughly can take several hours, and the cognitive load of holding a codebook consistently in your head degrades as fatigue sets in over dozens of transcripts. Beyond roughly 15 to 20 interviews, full manual coverage becomes hard to sustain within realistic timelines, which is exactly why so many researchers find manual coding too slow.
There is also a saturation argument. Guest, Bunce and Johnson (2006), in Field Methods, found that most codes in a homogeneous sample appeared within the first twelve interviews. But real projects are rarely that homogeneous, and once you add subgroups you need in order to compare them, the transcript count, and the coding burden, climbs quickly. See how many interviews are enough for the nuance.
How do AI-assisted tools code interview transcripts?
Once the dataset is large, or the deadline is short, AI-assisted coding does the mechanical first pass while you keep the interpretation. The important distinction is between a general chatbot, which summarises text and cannot be trusted to code it consistently, and a purpose-built tool that codes systematically with traceable evidence. We cover that difference in how to code qualitative interviews with AI.
Using Skimle as the example, the workflow mirrors the manual steps above, compressed:
- Add the transcripts. Text files import directly; audio and video are handled by built-in transcription with speaker identification, so recordings become coded transcripts without a separate tool.
- Choose the approach. Automatic thematic analysis or inductive analysis for emergent codes, or predefined categories to apply your existing codebook.
- Let it code every transcript. Each transcript is read systematically, and every insight is tied to a verified verbatim quote, so nothing is asserted that the source does not contain.
- Review and refine. Merge, split, rename, and re-file codes, exactly the judgement work manual coding requires, but without the weeks of labelling.
- Slice by subgroup. Attach participant metadata and compare how themes distribute across roles, segments, or sites.
If you need to finish in a traditional environment, the coded data exports to REFI-QDA for NVivo, MAXQDA, or ATLAS.ti, so AI coding and manual coding are complementary rather than mutually exclusive.
5 common mistakes when coding interview transcripts
- Coding interpretively too early. Jumping straight to "poor governance" instead of "waiting for sign-off" bakes in conclusions before you have the evidence. Stay descriptive on the first pass.
- Letting code definitions drift. Without written definitions, the same code means different things by interview 30. Define each code and hold to it.
- Never revisiting earlier transcripts. When you add a code late, earlier transcripts probably contain it too. Failing to go back understates how common it really is.
- Confusing frequency with importance. A code that appears often is not automatically the most important finding. A rare but vivid account can matter more. Count, but interpret.
- Losing the audit trail. If you cannot trace a theme back to specific quotes, reviewers and stakeholders have no reason to trust it. Keep every code linked to its source.
Avoiding these is easier when the tool enforces traceability for you, which is one reason manual coding at scale is so error-prone.
Frequently asked questions
What is the difference between a code and a theme?
A code is a short descriptive label applied to a specific excerpt of a transcript. A theme is a higher-level, interpretive pattern that groups several codes around a shared meaning. Codes are close to the data; themes are your interpretation of what the codes collectively mean.
How many codes should an interview transcript have?
There is no fixed number. A rich hour-long interview commonly yields dozens of coded segments across a handful of codes. What matters is that codes are applied consistently and each is defined clearly, not that you hit a target count.
Should I code inductively or deductively?
Inductively when you are exploring and want codes to emerge from the data; deductively when you have an existing framework or hypotheses to apply. Many researchers combine both, starting with a few deductive codes and adding inductive ones as they read. This blended approach is abductive.
Can I code interview transcripts in Excel?
You can, and for a handful of transcripts it works. The limits appear quickly: no easy traceability, painful consistency across many files, and no simple way to slice by subgroup. See how to do thematic analysis in Excel for the method and its ceiling.
How do I keep coding consistent across a team?
Write clear code definitions, code a shared subset independently, compare agreement, and reconcile differences by tightening definitions rather than overruling each other. A shared, well-defined codebook is the single biggest driver of consistency.
Ready to code your transcripts faster without losing the audit trail? Try Skimle for free and let AI do the systematic first pass while you keep control of interpretation, with every code traceable to its source quote.
Want to learn more? If you are an academic researcher, see how Skimle fits your workflow, and read our guides on how to code qualitative data and thematic analysis methodology.
About the authors
Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile
Sources
- Using thematic analysis in psychology. Braun & Clarke (2006), Qualitative Research in Psychology
- How Many Interviews Are Enough? An Experiment with Data Saturation and Variability. Guest, Bunce & Johnson (2006), Field Methods
- Code Saturation versus Meaning Saturation: How Many Interviews Are Enough? Hennink, Kaiser & Marconi (2017), Qualitative Health Research



