Data saturation is the point at which collecting additional data no longer produces new themes, codes, or insights relevant to your research question. In practice, it means your dataset is sufficiently rich to answer what you set out to investigate. But saturation is not a universal rule, and it is not an alternative to thinking carefully about your sample. The right number of interviews depends on your methodology, your question, and what kind of saturation you are aiming for.
What is data saturation in qualitative research?
Data saturation occurs when additional data no longer generates new categories, themes, or meaningful variation relevant to your research questions. The concept originates in grounded theory, where it signals that your theoretical categories are fully developed and further sampling would be redundant. It has since been adopted (sometimes loosely) across qualitative traditions as a shorthand for "enough data."
The key word is "relevant." Saturation is not about running out of things participants could say. People can always tell you more. Saturation means you have reached a point where more data would not materially change your analysis or deepen your understanding of the phenomenon under study.
This matters because sample size in qualitative research is rarely determined statistically. Instead of calculating power and picking a number before you begin, qualitative researchers often treat sample size as provisional, starting with a plan and adjusting based on what the data actually shows.
How many interviews for saturation? Key studies at a glance
| Study | Context | Saturation point |
|---|---|---|
| Guest, Bunce & Johnson (2006) | 60 interviews, homogeneous sample | ~12 interviews (most codes by 6) |
| Hennink, Kaiser & Marconi (2017) | Code vs meaning saturation | Code ~9; meaning ~16–24 |
| Francis et al. (2010) | Theory-based interview studies | 10 + 3 confirmatory (stopping criterion) |
| Hennink & Kaiser (2022) | Systematic review of empirical tests | Commonly 9–17 interviews |
Where does the concept come from?
Data saturation was first formalised by Barney Glaser and Anselm Strauss in their 1967 work on grounded theory. In that tradition, theoretical saturation is a methodological requirement: you sample until your emerging theory is fully developed and no new categories are needed to explain the data. Saturation and sampling are intertwined; you follow the data to wherever the theory leads.
The concept was later borrowed by other qualitative traditions, often stripped of its grounded theory meaning and reapplied as a general criterion for stopping data collection. This borrowing has created confusion. What means something precise in grounded theory becomes vague when applied to thematic analysis or phenomenology, where sampling logic differs considerably.
What did Guest, Bunce and Johnson (2006) actually find?
The most-cited empirical study on saturation is Guest, Bunce and Johnson's 2006 paper "How many interviews are enough? An experiment with data saturation and variability," published in Field Methods. Their finding (that thematic saturation occurred within the first twelve interviews, with basic elements present as early as six) has become the go-to justification for small qualitative samples worldwide.
It is worth understanding what they actually did. They analysed 60 in-depth interviews conducted with women in two West African countries on a well-defined topic (sexual behaviour and risk). They then tracked when new thematic codes stopped appearing as they analysed each successive interview. Saturation, by their measure, arrived by interview 12.
That finding is meaningful in context. Their study used a relatively homogeneous sample on a bounded topic with semi-structured interviews following a consistent guide. The more homogeneous your participants and the more bounded your research question, the faster saturation tends to arrive. When researchers cite "12 interviews" as a universal threshold, they are generalising from a specific study to all qualitative research, a step the authors themselves would not support.
If you are working on a study with diverse participants, an open-ended research question, or a topic where experiences vary considerably by context, twelve interviews is unlikely to be enough. The Guest Bunce Johnson study gives you a floor for well-defined studies, not a ceiling for ambition.
What is the difference between code saturation and meaning saturation?
This distinction, introduced by Hennink, Kaiser and Marconi in a 2017 Qualitative Health Research study, is one of the most useful conceptual contributions to the saturation debate.
Analysing 25 in-depth interviews with HIV patients and healthcare users in Atlanta, Hennink and colleagues found that:
- Code saturation (when all the thematic categories have been identified and no new codes are appearing) arrived at around 9 interviews. At this point, 91% of codes had been identified and the codebook had stabilised.
- Meaning saturation (when each code is richly understood, with sufficient variation, nuance, and context) required 16 to 24 interviews, depending on the conceptual complexity of the code.
The distinction matters enormously in practice. If you stop collecting data when you stop seeing new codes, you may have heard everything, but you have not necessarily understood it. Meaning saturation requires enough interviews to explore how and why participants describe their experiences differently, not just what they report.
Concrete codes ("feel well," "time," "travel") tended to reach meaning saturation quickly, often by interview 9 or 10. Conceptual codes ("disclosure," "HIV stigma," "trust in the healthcare system") required substantially more data to understand fully, reaching meaning saturation only between interviews 16 and 24.
The implication for researchers is that the right saturation target depends on your codes. Studies examining concrete, behavioural phenomena can reach saturation with fewer interviews than studies exploring complex social processes, identity, or institutional experience.
For a practical guide to how many interviews qualitative research studies typically require, see our dedicated guide.
How does grounded theory treat saturation differently?
In grounded theory, theoretical saturation is not a sample size criterion. It is the end-state of an iterative process. You collect data, analyse it, identify emerging categories, then collect more data deliberately designed to develop or challenge those categories (a process called theoretical sampling). You stop when your theoretical categories are fully developed and no new properties or dimensions emerge.
This is meaningfully different from checking whether new codes appear in interview 13 versus interview 12. Grounded theory saturation is a claim about theory, not about coverage. It requires you to be able to say not just "I have not heard anything new" but "my theoretical framework is complete and would not be materially changed by further sampling."
Grounded theory methodology therefore demands ongoing analysis during data collection, not a retrospective audit. You cannot code 20 interviews in a batch and then claim theoretical saturation. The iterative loop between collection and analysis is the method.
Why do Braun and Clarke reject saturation for reflexive thematic analysis?
Virginia Braun and Victoria Clarke have argued, in their 2019 paper "Reflecting on reflexive thematic analysis" in Qualitative Research in Sport, Exercise and Health, that saturation is fundamentally incompatible with reflexive thematic analysis (RTA) as they conceptualise it.
Their argument runs as follows. In reflexive TA, themes are not discovered in the data. They are constructed by the researcher, who brings their own theoretical position, questions, and interpretive perspective to the material. Because themes are constructed rather than found, the logic of saturation ("no new themes are appearing") does not translate. A different researcher, with a different orientation, would construct different themes from the same dataset. There is no objective threshold to reach.
Using saturation as a stopping rule also encourages researchers to frame their analysis in positivist terms: as if themes were waiting in the data to be uncovered, and sufficient data collection would eventually uncover all of them. Braun and Clarke reject this framing entirely. For RTA, sample size decisions should be made in advance based on the research question, the depth of analysis intended, and practical constraints, not treated as a moving target to be determined by when themes stop appearing.
This creates a genuine tension in the literature. Many reviewers and ethics committees still expect saturation to be invoked and justified. Many journal editors treat it as a marker of rigour. PhD supervisors often require it. Braun and Clarke's position asks researchers to resist that expectation and instead articulate a different kind of justification for their sample, one rooted in depth and analytical purpose rather than coverage thresholds.
If you are using reflexive thematic analysis, our dedicated guide covers what rigour means within that framework and how to write up your sample size decisions in ways that reflect the methodology accurately.
What did Francis et al. (2010) add to the debate?
Francis et al. (2010), writing in Psychology and Health, took a more pragmatic approach aimed at researchers doing theory-based interview studies. Their contribution was a structured stopping criterion: specify in advance an initial analysis sample (for example, ten interviews), then monitor subsequent interviews in sets of three. If three consecutive interviews produce no new themes or information relevant to your theoretical framework, you have reached saturation.
This approach has the advantage of being transparent and auditable. Rather than retrospectively claiming saturation at the end of a project, researchers prospectively define what saturation means for their study and demonstrate they reached it. The three-interview stopping criterion has been widely adopted, particularly in applied health and psychology research.
The limitation is similar to the broader critique: this works best for studies with bounded research questions and theory-driven frameworks. For more open-ended, exploratory research, the stopping criterion may close down inquiry prematurely.
What type of saturation applies to your study?
The right approach to saturation depends on your methodology.
| Research approach | Saturation concept | What it means |
|---|---|---|
| Grounded theory | Theoretical saturation | Theoretical categories fully developed; new sampling would not extend the theory |
| Applied qualitative / health research | Operational saturation (Francis et al. approach) | Pre-specified criterion: N initial interviews, then stopping after 3 without new themes |
| Descriptive or thematic approaches | Code saturation (Hennink et al.) | No new codes appearing; basic coverage achieved |
| Rich, contextual thematic analysis | Meaning saturation (Hennink et al.) | Full nuance and variation within codes understood |
| Reflexive thematic analysis (Braun & Clarke) | Saturation not applicable | Sample size justified on depth and analytical purpose, not coverage thresholds |
| IPA | Not typically invoked | IPA uses small purposive samples; depth over coverage |
The practical implication: when your committee, your ethics board, or a journal reviewer asks you to justify saturation, knowing which concept you are invoking, and being clear about what it means for your methodology, is more defensible than citing "12 interviews per Guest et al."
How do you know when you have reached saturation in practice?
For researchers working in traditions where saturation is appropriate, here is what the process looks like in practice.
Analyse as you go. Saturation is a claim about your analysis, not your data collection. You cannot determine whether you have reached it by looking at interview count alone. You need to be actively coding and generating categories during fieldwork, not at the end.
Track new codes explicitly. Keep a running codebook and note which interview each code first appeared in. When the rate of new code introduction drops significantly and then stops entirely across several interviews, you have evidence of code saturation.
Revisit existing codes for depth. Code saturation is the floor, not the ceiling. Once no new codes are emerging, ask whether your existing codes are richly understood. Are there sub-types, contradictions, or contextual variations you have not yet captured? If so, you have not reached meaning saturation.
Consider participant variation. If your sample includes meaningfully different groups (different occupations, ages, backgrounds, levels of experience), saturation should hold across those groups, not just in aggregate. You may reach saturation within one group while another remains under-explored.
Document your decision. When you write up your methods, explain the saturation criterion you used, when you reached it, and what you did to verify it. A simple saturation table (showing which codes appeared in which interviews) is a practical and auditable way to demonstrate this.
Tools like Skimle can assist here. When you are coding qualitative data systematically, you can track the frequency and distribution of codes across interviews and identify when your codebook has stabilised. This makes the saturation decision transparent and reproducible rather than an intuitive judgement made after the fact. Academic researchers working on PhD studies can see how Skimle fits into a rigorous qualitative workflow.
How many interviews is actually enough?
There is no single answer, but some evidence-based guidance exists.
For well-defined, relatively homogeneous samples on bounded topics: code saturation may arrive by interview 9-12, consistent with both Guest et al. and Hennink et al. For richer contextual understanding, plan for 16-24.
For diverse samples or exploratory questions: budget for more. Studies exploring complex social processes often require 25-50 interviews, and some published studies go higher.
For grounded theory: sample size is not decided in advance. You collect until theoretical saturation, which depends entirely on how quickly your categories develop.
For reflexive TA, IPA, and other interpretive approaches: use the principle of "information power" (Malterud et al. 2016), a concept that accounts for sample specificity, the quality of participant-researcher interaction, the depth of analysis intended, and the theoretical framework. A small, rich, well-selected dataset is analytically superior to a large, diffuse one.
Our guide on qualitative research sample size covers these practical ranges in more detail across different methodologies.
Frequently asked questions
What is data saturation and why does it matter?
Data saturation is the point in qualitative data collection when additional interviews, observations, or documents no longer produce new codes, themes, or meaningful insights relevant to your research question. It matters because it provides a principled criterion for stopping data collection, one grounded in analytical sufficiency rather than an arbitrary number.
Is 12 interviews always enough for saturation?
Not in general. Guest, Bunce and Johnson (2006) found saturation occurring around 12 interviews in a specific, bounded study with a relatively homogeneous sample. Studies with diverse participants, complex research questions, or exploratory designs routinely require more. The number 12 is a minimum reference point for well-defined studies, not a universal rule.
What is the difference between code saturation and meaning saturation?
Code saturation means no new thematic categories are appearing in your data. Meaning saturation means each category is richly and fully understood, including its nuances, variations, and context. Hennink et al. (2017) found code saturation arrived at around 9 interviews, while meaning saturation required 16-24. Meaning saturation is the higher bar and the one most studies need to reach.
Should I use saturation as a criterion for reflexive thematic analysis?
Braun and Clarke argue no. In reflexive TA, themes are constructed by the researcher rather than discovered in the data, so the logic of saturation ("no new themes are appearing") does not apply. Instead, they recommend justifying your sample size on the basis of depth, analytical purpose, and the nature of your research question. However, many reviewers and ethics committees still expect saturation to be referenced. If you are in that position, it is worth explaining clearly which concept you are invoking and why it fits (or does not fit) your methodology.
How do I report saturation in my thesis or paper?
Report the saturation criterion you used, when you reached it, and how you verified it. For example: "Following Francis et al. (2010), we set an initial analysis sample of ten interviews and applied a stopping criterion of three consecutive interviews without new themes. Saturation was reached at interview 17, confirmed by two subsequent interviews that produced no additional codes." A supplementary table showing code distribution across interviews strengthens the claim considerably.
Try Skimle for your qualitative research
Analysing interviews systematically and tracking saturation doesn't have to mean weeks of manual coding.
Try Skimle for free and see how AI-assisted thematic analysis, with full traceability from every insight back to the source quote, can make your saturation decisions transparent, auditable, and faster to reach.
Want to go deeper on methodology? Read our guides on thematic analysis, how to code qualitative data, reflexive thematic analysis, and qualitative research methods.
About the authors
Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile
Sources
- How many interviews are enough? An experiment with data saturation and variability. Field Methods (2006)
- Code saturation versus meaning saturation: How many interviews are enough? Qualitative Health Research (2017)
- Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health (2019)
- To saturate or not to saturate? Questioning data saturation as a useful concept for thematic analysis. Qualitative Research in Sport, Exercise and Health (2021)
- What is adequate sample size? Operationalising data saturation for theory-based interview studies. Psychology and Health (2010)
- A simple method to assess and report thematic saturation in qualitative research. Guest, Namey & Chen (2020), PLOS ONE



