Every vendor demo for an AI research tool includes the same reassurance, delivered in roughly the same tone: don't worry, a human stays in the loop. It's become such a standard line that it barely registers as a claim anymore, more of a compliance phrase than a design decision. Which is exactly the problem, because in a lot of tools, what "human in the loop" actually means in practice is an accept button at the end of a process you had no visibility into and no way to meaningfully change.
Raw data goes in. A black box does something. A themes summary comes out. The human's role is to click through it, and unless something looks obviously wrong, most people do exactly that.
Why does "human in the loop" so often mean so little?
The phrase describes a position in a workflow, not a capability. Technically, if a person clicks "accept" before a result gets used, a human was in the loop. That's a low bar, and most AI tools clear it without providing anything close to genuine oversight.
Three failure patterns show up repeatedly in tools that use the phrase but don't build for it.
You can't see inside the process. The tool takes documents in and produces themes out, with nothing in between to inspect. There's no way to check whether a category is built on three quotes or thirty, whether the whole corpus was considered or just a sample, or what got left out. Without that visibility, "review before accepting" is not a meaningful instruction, because there's nothing substantive to review.
Editing is technically possible and practically painful. Some tools do let you change a category name or reassign a quote, but the interface treats it as an edge case rather than a core workflow: three clicks become fifteen, changes don't propagate consistently, or the AI "helpfully" reverts your edit on the next pass. Manual control that costs ten minutes per correction gets used rarely, which functionally means it doesn't exist.
Experts default to accepting because checking is expensive. Give a domain expert a plausible-looking AI output and a black box they can't easily interrogate, and the common outcome is that they spot the occasional result that looks obviously wrong and accept everything else, not because they've verified it, but because verifying it properly would take longer than doing the analysis manually in the first place. This isn't a character flaw. It's a predictable response to a workflow that makes real scrutiny expensive.
What would real human-in-the-loop control actually require?
Three things, and all three need to be genuinely low-friction, not just technically available.
Full transparency across the analysis, not just the output
A themes summary with supporting quotes attached is the minimum bar, not the whole picture. Real transparency means being able to open any source document and see exactly what was coded and what wasn't, in both directions: from a finding back to its evidence, and from a document forward to what the analysis did and didn't pick up. We've written more on why this second direction is the one that actually catches missed patterns, rather than just confirming the ones the tool already found.
Manual editing that doesn't fight you
Categories should be mergeable, splittable, and renameable directly, with coding decisions individually reassignable, and none of it should require a support ticket or a multi-step workaround. Editing in the categories view is meant to work the way editing a spreadsheet cell works: click, change, done. AI should assist that process (suggesting a merge, flagging a possible duplicate) without ever making the manual version harder to reach than the automated one.
Workflows the expert actually guides
The strongest version of human-in-the-loop design doesn't position the human as a final checkpoint after AI has already decided everything. It positions the researcher as the one steering the process throughout: choosing whether to start from predefined categories or let them emerge from the data, deciding which passages warrant a closer manual read, directing an agentic analysis toward the specific research question that matters for this project rather than a generic theme extraction. AI does the mechanical first pass. The expert decides what the analysis is actually for.
A test for telling the real thing from the checkbox
Before trusting a vendor's "human in the loop" claim, or before trusting your own team's process, run through this:
- If I disagree with a category, can I fix it in under a minute, or does it require a workaround?
- Can I see what the analysis did not pick up, not just what it did?
- If I made an edit yesterday, is it still there today, or did the next AI pass quietly revert it?
- Does the tool make it easier to check the work than to skip checking it?
- Would an expert under deadline pressure actually use the review step, or would they rubber-stamp it because checking properly takes too long?
That last question is the decisive one. A control feature that's technically available but practically expensive gets used the way most people use terms and conditions: skipped, because the cost of engaging with it exceeds the perceived risk of not engaging with it. Real human-in-the-loop design has to account for how experts actually behave under time pressure, not how they behave in a vendor's demo environment where nobody's on deadline.
Where this fits in a broader evaluation
This is one piece of a larger buying decision, not the whole of it. If you're running a formal software evaluation, our RFP checklist for qualitative analysis software covers control and editability alongside six other domains (analytical capability, traceability, data governance, collaboration, export, and commercial model) with a scoring template you can use directly. Treat this piece as the diagnostic question to ask before you get that far: is "human in the loop" here a real design commitment, or a line in the pitch deck.
For teams evaluating specifically in a market research or customer insights context, see how Skimle fits that workflow, and for the design principles behind building genuine control into an AI research tool, our 9 design criteria for AI qualitative analysis tools goes deeper into what the academic critiques of AI-assisted qualitative research actually require.
Try Skimle with real editable control
If you want to test whether "human in the loop" holds up under the five-question check above, the fastest way is to try it on a real project rather than a demo.
Try Skimle for free and edit a category, then check whether the change actually stuck.
Related reading:
- How to evaluate qualitative analysis software: an RFP checklist
- Two-way transparency: creating confidence in AI to make it useful for real work
- 9 design criteria for AI qualitative analysis tools
About the authors
Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile



