For most AI-moderated interview studies, 30 to 50 completed interviews per segment you want to compare is enough to find the main themes, and 100 to 200 in total is enough to say with some confidence how common each theme is. Beyond that, extra interviews mostly repeat what you already know, unless you are slicing the data into many groups or hunting for rare cases.
That answer surprises people from both directions. Traditional qualitative researchers are used to 12 to 20 in-depth interviews and suspect that 200 is a waste. Teams buying AI interviewing platforms are told that more is always better, because the marginal interview costs a few dollars. Both instincts are partly right, and the difference comes down to what question the study has to answer.
Why does sample size change when interviews are AI-moderated?
In a human-moderated study, sample size is set by cost. A moderator, a recruiter, incentives and transcription push the cost of one in-depth interview into the hundreds of dollars, so most commercial studies stop at 15 to 30 interviews and rely on saturation arguments to justify the number.
AI moderation removes most of that cost. With a tool like Skimle Ask, an interviewer can run hundreds of conversations in parallel, in any language, by text or voice, with follow-up questions on every answer. The constraint moves from fieldwork to two other places: recruiting enough of the right people, and analysing what comes back without drowning in it.
When interviews are cheap, the question is no longer "how few can we get away with?" but "how many do we need for the decision we are making?" That is a better question, and it has a different answer for different study types.
What does the saturation research say?
The best-known evidence comes from academic studies of saturation, the point where new interviews stop producing new themes.
- Guest, Bunce and Johnson (2006) coded 60 interviews and found that 92% of the codes they eventually identified had appeared within the first 12 interviews.
- Hennink, Kaiser and Marconi (2017) separated two kinds of saturation. Code saturation, where you have heard every theme at least once, came at about 9 interviews. Meaning saturation, where you understand each theme well enough to explain it, took 16 to 24.
- Hennink and Kaiser (2022) reviewed 23 empirical studies and concluded that 9 to 17 interviews, or 4 to 8 focus groups, usually reach saturation for a fairly homogeneous group and a narrow research question.
We cover these studies in depth in our guides to how many interviews are enough and data saturation. The important caveat for market research is in the small print: these numbers hold for one homogeneous group and one narrow question. Commercial studies rarely have either.
Why might you need 200 interviews instead of 20?
Saturation tells you when you have heard every theme. It does not tell you how common each theme is, whether it differs between segments, or whether a theme mentioned by 3 people out of 20 matters. Those are exactly the questions marketing and product teams ask.
The chart at the top of this post shows the typical shape. Most themes appear in the first batch of interviews. Later batches add fewer and fewer new themes, but they make the counts behind each theme far more reliable.
There are four situations where going well beyond saturation pays off:
- You need to compare segments. Saturation applies per group. If you want to compare new customers with churned customers across three markets, that is six groups, and each needs enough interviews to stand on its own. Thirty per cell gets you to 180 quickly.
- You need prevalence, not just presence. "Some customers mention delivery delays" is a qualitative finding. "Roughly a third of churned customers mention delivery delays, against a tenth of retained ones" is a finding a board acts on. Counting themes needs larger numbers, which is why quantifying qualitative data has become a standard step in commercial reporting.
- You are looking for rare but important cases. A safety issue, a fraud pattern or an unmet need held by 5% of users will not appear in 20 interviews. In 200 it appears about 10 times.
- The decision is expensive. If the study informs a launch, a price change or a repositioning, the cost of 150 extra AI interviews is small next to the cost of a wrong call.
If you run customer or consumer research, see how this fits the market research and customer insights workflow.
When are 200 interviews too many?
More interviews are not free just because fieldwork is cheap. Three costs remain.
- Recruitment and incentives. Panel costs scale with sample size, and hard-to-reach audiences (CFOs, rare-disease patients, niche B2B buyers) may simply not exist in those numbers. Pushing a panel past its real supply is also how fraudulent respondents get in, which we cover in our guide to spotting fake respondents in AI interviews.
- Respondent goodwill. Every interview asks for 15 to 30 minutes of someone's time. Running 400 when 120 would answer the question is poor practice, especially with customers.
- Analysis. Reading 200 transcripts by hand takes weeks. Without a tool that codes every transcript systematically, a large sample becomes a pile of text that gets skimmed, which defeats the purpose of collecting it.
If your question is exploratory ("what do people even think about this category?") and your audience is one fairly uniform group, 20 to 40 AI-moderated interviews are plenty. Spend the savings on a second wave once you know what to probe.
How do you size an AI-moderated interview study?
Here is the rule of thumb we use with clients. Treat the numbers as starting points, not targets.
| Study goal | Typical size | Why |
|---|---|---|
| Exploratory discovery in one group | 20 to 40 | Code saturation plus a margin for meaning |
| Theme prevalence in one group | 80 to 150 | Stable percentages for the top 10 themes |
| Comparing 2 to 4 segments | 30 to 50 per segment | Saturation per cell, counts comparable across cells |
| Comparing 5+ segments or markets | 25 to 40 per cell, 200+ total | Wide coverage, deeper probes on the cells that differ |
| Finding rare cases (5% incidence) | 150 to 300 | Enough to see the case 8 to 15 times |
Then apply four checks before you launch:
- Write down the comparison you will show. If the final slide compares three segments, size for three segments. If it doesn't, don't.
- Plan a checkpoint. Field the first 30 to 50 interviews, analyse them, and see whether new themes are still appearing. AI interviewing makes a two-stage design easy, and it often saves half the sample.
- Fix the minimum per cell. Never report a theme share for a group of fewer than about 20 respondents. Report it as a qualitative observation instead.
- Budget the analysis, not just the fieldwork. A study you cannot analyse properly should be smaller.
How do you know you have reached saturation in a large AI interview study?
In a 15-interview study, a researcher can feel saturation: by interview 12, nothing surprises them. In a 200-interview study nobody has read every transcript closely, so you need a visible measure instead.
The practical method is to analyse in batches. Code the first batch, then check how many new categories each following batch adds and whether the share of interviews in each category is still moving. When a batch of 25 adds no new categories and the top themes' shares move by less than a few percentage points, you are there.
This is where analysis tooling matters. Skimle codes every interview line by line into a category structure, keeps the verbatim quote behind each insight, and counts how many respondents sit in each category. Because the metadata from your screener (segment, market, customer tenure) is attached to each interview, it also ranks which group differences are largest, as in the screenshot below. That turns "we interviewed 200 people" into "here is what separates churned customers from loyal ones, with the quotes to show it".

What changes in reporting when n=200?
A large qualitative sample tempts teams to report it like a survey. Resist that, but don't swing to the other extreme either.
- Report shares for big themes, with the base. "34% of churned customers (n=62) raised onboarding" is fair. Avoid decimals and significance tests unless the design supports them.
- Keep the quotes. The point of qualitative interviews is the why. Every number on a slide should sit next to a verbatim quote that explains it.
- Say what was not found. With 200 interviews, the absence of a theme is itself evidence. "No respondent mentioned price as a reason to leave" is a strong finding.
- Separate saturation from prevalence. State whether each finding is "we heard this" or "this is how common it is". Mixing the two is the most common reporting error in qual at scale.
For more on putting large qualitative studies in front of decision-makers, see presenting qualitative research findings to executives.
Frequently asked questions
Is 200 AI interviews the same as a survey of 200 people?
No. A survey asks everyone the same closed questions and gives you precise percentages on a narrow set of variables. AI-moderated interviews ask open questions with follow-ups, so you learn reasons and language you didn't anticipate. You can count themes across 200 interviews, but the counts describe open answers and should be reported with that caveat.
How many AI interviews do I need per segment?
As a starting point, 30 to 50 completed interviews per segment you plan to compare. Fewer than about 20 per segment is too thin to report theme shares, though it still supports qualitative observations.
Do AI-moderated interviews reach saturation faster than human interviews?
Not necessarily. Saturation depends on the diversity of the group and the breadth of the question, not on who asks. AI interviewers do ask follow-ups more consistently, which can make each interview more complete, but they may probe less creatively than an experienced human moderator. Our comparison of AI and human interviewing covers the trade-offs.
How long does it take to analyse 200 AI interview transcripts?
By hand, several weeks. With a systematic AI analysis tool such as Skimle, the coding runs in under an hour, and the researcher's time goes into reviewing categories, checking quotes and writing the story, typically one to two days for a study of that size.
Planning a large qualitative study? Try Skimle for free to run AI-moderated interviews with Skimle Ask and analyse every transcript with traceable quotes. Plans start from around $23 (€20) per user per month.
Related reading:
- How many qualitative interviews are enough?
- Data saturation in qualitative research
- Skimle Ask: qual at scale
- Best AI interview tools in 2026
About the authors
Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile
Sources
- How many interviews are enough? An experiment with data saturation and variability - Guest, Bunce and Johnson (2006), Field Methods
- Code saturation versus meaning saturation: how many interviews are enough? - Hennink, Kaiser and Marconi (2017), Qualitative Health Research
- Sample sizes for saturation in qualitative research: a systematic review of empirical tests - Hennink and Kaiser (2022), Social Science & Medicine



