How to use AI responsibly in qualitative market research: 7 rules for insights teams

Responsible AI in qualitative market research: 7 rules covering where AI helps, why synthetic respondents fail, and what human-in-the-loop really requires.

Cover Image for How to use AI responsibly in qualitative market research: 7 rules for insights teams
Share this article:

Use AI for the work that scales badly: transcription, coding, theme detection and synthesis across a large corpus of real human data. Do not use it to invent respondents. Responsible practice needs three things in place: a named human who reviews every theme, traceability from each finding back to the source quote, and the ability to override the model. Tools like Skimle are built around exactly that loop.


The insights profession has settled into an odd position. AI is now near-universal in the research workflow and deeply distrusted at the same time, and both of those things are correct.

This guide is for market researchers and customer insights people inside companies who need a defensible position on AI, whether that is for a procurement review, a client conversation, or an internal policy nobody has written yet. We will cover where AI genuinely earns its place in qualitative work, what the evidence says about synthetic respondents, and the seven rules that make the difference between AI-assisted research and AI-flavoured guesswork.


Where does AI actually help in qualitative market research?

AI helps most where the work is mechanical, high-volume and currently done badly because nobody has the hours.

Qualitative analysis has always had a scale ceiling. A researcher can read 20 interviews carefully. At 60, memory and consistency start to fail. At 200, the analysis becomes a summary of whichever transcripts were read most recently. The ceiling is clerical rather than intellectual, and clerical ceilings are exactly what machines lift.

Here is a realistic division of labour:

TaskWhat AI does wellWhat the human must still doRisk if unsupervised
TranscriptionFast, multi-language, speaker-separated transcriptsSpot-check names, jargon, product termsSilent mistranscription of the terms that matter most
First-pass codingApply a category framework consistently across hundreds of documentsDefine the angle to take, judge edge cases, add categories the model missedCategories that fit the model's priors rather than the market
Theme detectionSurface patterns across a corpus too large to hold in memoryDecide which patterns are meaningful and which are artefactsPlausible themes with no real support in the data
Cross-segment comparisonSlice findings by metadata (segment, tenure, geography, wave)Choose the cuts that matter commerciallyOver-reading differences in small subgroups
Synthesis and draftingAssemble a first draft with quotes attachedOwn the interpretation and the recommendationA polished narrative nobody has verified
Generating respondentsNothing you should rely onRecruit real peopleFabricated evidence presented as market truth

The pattern is consistent. AI is good at applying a structure across more material than a person can hold in their head. It needs judgement and "research taste" for deciding what the structure should be, and it is worse than useless at manufacturing the material itself.

This is why we have argued before that manual interview coding is too slow to be the default, while also insisting that rigour means being willing to return a negative result. Both positions follow from the same view of what the technology is for.


Why should you not trust synthetic respondents?

Synthetic respondents are AI-generated personas that answer your survey or interview as if they were real people. The pitch is irresistible: 1,000 responses by lunchtime, no panel cost, no recruitment lag. We have written a full assessment in synthetic respondents in research: promise, pitfalls and when to use in 2026. The short version is that the published evidence has got steadily worse for them, not better.

They are wrong at the level that matters. Verasight's January 2026 report, Can Large Language Models Replicate Survey Data Across Topics?, put GPT-5.2 against 2,000 nationally representative US adults across 52 questions. The mean absolute error was 14.5 percentage points across 305 response options, rising to 23.4 points on healthcare questions. On multi-answer questions the model ignored 31% of the available options entirely, including one option that 40% of real humans selected and the model never picked once.

They flatten variance and exaggerate identity. A cross-domain benchmark published in July 2026, When Synthetic Users Fail by Chen, Zhu and Zheng, found that models treat demographics as vastly more predictive than reality. In one case political views explained roughly 1.5% of the real variation in a response but around 67% of the model's variation. Applied to segment targeting, the models inflated between-segment gaps by 2 to 4.7 times, and pointed at the wrong segment in 50% of cases on US social attitudes and 72% on cross-cultural values. Notably, the more capable model in the set stereotyped more, not less, so this does not get fixed by waiting for the next release.

The replication record is poor. Jim Lewis and Jeff Sauro's review of 12 peer-reviewed synthetic user experiments at MeasuringU, published April 2026, found that only 21% of attempted replications of classic social science studies succeeded. Survey studies matched high-level means but produced inaccurate subgroup means, artificially small standard deviations and unreliable regression coefficients.

Put those together and you get a specific failure mode, which is the dangerous one for commercial research. Synthetic data looks right in aggregate and is wrong everywhere you would actually make a decision: the subgroup, the segment, the minority view, the surprising finding. A synthetic study will reliably tell you what you already assumed about your market, dressed up with quotes.

The problem with synthetic respondents is not that they are obviously fake. It is that they are plausibly fake, confidently wrong at the subgroup level, and biased towards the consensus already present in their training data.

There is a narrow legitimate use: pressure-testing a discussion guide, rehearsing an interview, or sanity-checking whether a question is comprehensible before it goes to field. Treat that as a rubber duck, not a respondent. Nothing generated this way should enter a findings deck.

Another use case where synthetic respondents can help is when you are researching things already well known - for example how does this B2B SaaS landing page check the boxes that typical B2B SaaS landing pages do? But for this use case... why bother with simulating hundreds of respondents when basic AI tools already hold the summarised taste of those people and can give essentially the same answers on what to fix and how?


What does human-in-the-loop actually mean?

"Human in the loop" is the phrase everyone uses and few define. In qualitative analysis it has a precise meaning: at every point where the system makes an interpretive judgement, a researcher can see that judgement, check its basis, and change it.

Sounds great in theory, but the problem in practice is that without proper tooling the role of the human quickly becomes to blindly hit the "Enter" key to accept the suggestions. Or, that they start to judge just the output in terms of it making sense or not, instead of also understanding the underlying reasoning and process. In the latter case, they start to simply apply their pattern recognition and biases instead of truly be a part of the loop.

Real human-in-the-loop requires three concrete properties from your tooling.

Visibility. You can see every coded passage, not just the theme summary. If the tool tells you 34% of enterprise customers raised integration reliability, you can open that number and read all of the passages behind it. In Skimle, every insight links back to its source document, which is the whole point of the working with insights view.

User-friendly editability. You can rename a category, merge two that are duplicates, split one that has collapsed two distinct ideas, and remove a coded passage that does not belong. An analysis you cannot easily correct is a report, not an analysis. The manual editing needs to be equally accessible as the AI-one, and the AI-assistance should remain available even after the user has manually altered something. "My way or high-way" type AI easily leads to humans just accepting the AI answers to save them from extra work. Skimle's categories view exists so that the AI's first pass is a starting point rather than a verdict.

Directability. You can impose your own framework rather than accept the model's. Predefined category analysis lets you bring the coding frame you already use across studies, or start from Skimle's suggestions when you genuinely want the data to speak first. Choosing between those two postures is a methodological decision, and the tool should let you make it.

If a tool gives you a summary and no way to interrogate or change it, the human is not in the loop. The human is downstream of the loop, reading its output... or worst case just talking about loops in some conference while in reality absent from the actual customer understanding...


The 7 rules

1. Never generate the data you are trying to learn from

Use AI on material that real people produced: interviews, open text, support conversations, sales calls, reviews, community posts. The moment the model is producing the respondent rather than processing the respondent, you have left research and entered simulation. If you cannot name the humans behind a finding (in aggregate, with consent), the finding is not evidence.

Every theme in a deliverable should be traceable to specific passages in specific documents. This is the single most valuable discipline in AI-assisted work, because it converts an unverifiable assertion into a checkable one. When a stakeholder pushes back on a finding, being able to show the six quotes behind it in ten seconds ends the argument. We make the case at length in two-way transparency: creating confidence in AI.

3. Review the coding, not just the summary

The summary is the most persuasive and least reliable artefact an AI produces. Reviewing it tells you whether the writing is good. Reviewing the coded passages tells you whether the analysis is right. Budget time for the second. A practical minimum: read every passage behind your top three themes, and read a random sample of 20 passages from elsewhere in the corpus.

4. Assume the model has a house view and check against it

Language models carry the dispositions of their training data, which for most commercial models means Western, English-language, online-expressed opinion. That matters when you are researching a market in Poland, a customer segment of field engineers, or anyone whose way of talking is not well represented online. Read our guide to bias in AI-assisted qualitative analysis for the specific mechanisms and mitigations. The practical check: look hardest at the findings that flatter your hypothesis. In Skimle we do our best to hide possible bias markers from the LLM: for example when comparing answers from men and women, we simply label them "group A" and "group B" for the AI to ensure the differences are more from the data and less from it's inherent biases.

5. Watch for the failure modes specific to LLMs

Hallucinated quotes, context truncation on long documents, and black-box aggregation are all real and all detectable. Ask your vendor how each is handled, and test it: upload a long transcript, ask for a quote, and check the quote exists verbatim. We cover the mechanics in AI qualitative analysis: hallucinations, context and the black box.

6. Keep one person accountable for the interpretation

Someone signs the findings. Not the tool, not "the AI said". This sounds like a governance formality and is in fact the thing that keeps standards up, because an accountable analyst reads the material.

7. Be transparent about method with your audience

Say what was AI-assisted and what was not. A methods note of three sentences ("42 interviews, transcribed automatically and checked; first-pass coding by Skimle against a predefined framework; all themes reviewed and edited by the research lead; every finding traceable to source") buys more credibility than silence, and costs nothing. Clients and internal stakeholders are already assuming AI was involved. Being specific is the advantage.


How do you write an AI policy for an insights team?

Most teams do not need a twelve-page document. They need a one-page answer to four questions.

What data may be processed with AI, and where does it go? Name the tool, the hosting region and the retention policy. For European teams this is usually the first question procurement asks. Where personal data is involved, decide whether transcripts are pseudonymised before analysis, and see how to anonymise interview transcripts for the practical steps.

What is AI allowed to do, and what is reserved for humans? The table earlier in this post is a reasonable starting template. Copy it and edit the last two columns to reflect your standards.

What is banned outright? For most insights functions this list is short: no synthetic respondents in findings, no AI-generated quotes, no client deliverable whose themes have not been reviewed by a named researcher.

How is AI use disclosed? Internally, in the deck. Externally, in the methods section. Agencies working under ESOMAR guidance should align the wording with their existing code commitments.

That is a policy. It fits on a page and it will survive contact with a procurement questionnaire.


Does responsible AI slow you down?

Less than people expect, because the review work replaces reading work that was never actually happening at scale.

Consider a typical 60-interview consumer study. Manual transcription and coding runs to several weeks of analyst time, and in practice the analysis leans on the 15 or 20 interviews the lead read closely. The AI-assisted version front-loads structure: transcripts in hours, first-pass coding across all 60 documents, then a review cycle where the researcher checks the framework, corrects the coding, and reads the passages behind each theme. The review is real work. It is also work on the full corpus rather than a convenience sample of it.

The gain is not only speed. It is coverage. A finding that appears in 4 of 60 interviews is invisible to a researcher who read 18 of them, and visible to one whose analysis touched all 60. Weak signals in small subgroups are frequently the commercially interesting ones, which is the same reason synthetic data is dangerous: it is precisely the subgroup level where the models fail.

For the tooling landscape, see our complete comparison of qualitative data analysis tools and the more focused interview analysis software comparison for 2026. If you work in consumer insights or at an agency, the customer and market researchers use-case page covers how this fits a commercial research workflow end to end.


What about AI-moderated interviews?

Collecting data with AI sits in a different category from generating it. An AI interviewer talking to real people is still collecting real human responses, so the objections that sink synthetic respondents do not apply. The questions that do apply are about quality of moderation: does the interviewer probe well, does it follow interesting tangents, does it know when to stop.

Our comparison of AI interviewing versus human interviewing sets out where each performs better. In brief, AI moderation works well for structured-but-open topics at volume (concept reactions, post-purchase experience, employee sentiment) and less well for sensitive, exploratory or highly technical conversations where a skilled human reads the room. Skimle Ask is our implementation of this, and the introduction to Skimle Ask covers what it is and is not for.

The rule stays the same. Real respondents, always. AI can help you reach more of them and can help you make sense of what they said. It should never stand in for them.


Frequently asked questions

Is it acceptable to use AI for coding qualitative data in commercial research?

Yes, provided the coding is reviewable and reviewed. AI-assisted coding applies a framework more consistently across a large corpus than a team of analysts working separately, which is a rigour gain rather than a compromise. The conditions are that a researcher defines or approves the framework, every coded passage links to its source, and the analyst can correct errors. Coding you cannot inspect is the problem, not coding assisted by a machine.

Can synthetic respondents ever be used legitimately?

For rehearsal, yes. Testing whether a discussion guide flows, whether a question is ambiguous, or whether a survey is too long are all reasonable uses, because you are testing your instrument rather than measuring a market. What they cannot do is produce findings. The published evidence shows acceptable aggregate accuracy paired with severe subgroup error, inflated segment differences and suppressed variance, which makes them unfit for exactly the decisions research is commissioned to inform.

How do I explain our AI use to a client or procurement team?

Answer four things: which tool, where the data is hosted and for how long, what the AI does versus what a human does, and how findings are traced back to source. A three-sentence methods note covering these will satisfy most reviewers. Vagueness causes far more procurement friction than AI use itself.

Does AI-assisted analysis reduce the researcher's role?

It moves the effort. Less time goes to transcription, mechanical coding and searching for a half-remembered quote, and more goes to framework design, interpretation and challenging what the data appears to say. The judgement work grows as a share of the job. The clerical work shrinks. Most researchers we speak to consider that a good trade.

How much of an AI-generated analysis should I check?

At minimum, every passage behind your headline findings plus a random sample from the rest of the corpus. If the random sample throws up coding you disagree with more than occasionally, the framework is wrong and needs fixing before you go further. Our 20-question AI qualitative analysis checklist is a useful pre-publication review.


Ready to run AI-assisted analysis you can actually defend? Try Skimle for free. Upload your interviews or open-text feedback, run the analysis inductively or against your own framework, and trace every theme back to the exact passage that produced it.

Related reading: See our guides on synthetic respondents in research, bias in AI-assisted qualitative analysis, and how to analyse customer interviews at scale.


About the authors

Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile

Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile


Sources

Dig deeper to your data with Skimle

Skimle collects, analyses and categorises interviews, survey responses, reports and other qualitative data automatically. Our modern qualitative analysis software combines a rigorous and transparent workflow with the speed of AI.

Upload text or audio, remove sensitive data with Skimle Anonymise, automatically create categories and sub-categories, explore the data across documents and export the data to seamlessly fit your workflow. Built by professionals for professionals, with full privacy and GDPR compliance.

Free trial · No credit card required · Full plans from €20/month