To evaluate qualitative analysis software properly, score every vendor across 7 domains: analytical capability, traceability, control, data governance, collaboration, export and lock-in, and commercial model. Run a pilot on your own corpus with timed tasks rather than watching a demo. Weight traceability and control highest, because they determine whether the output survives a stakeholder challenge.
Many software evaluations in insights teams go the bad way: Three demos, a feature comparison spreadsheet, a price negotiation, and a decision driven by whichever vendor's interface looked cleanest in a 40-minute call. Chances are, eighteen months later the tool is used by two people for one thing while you are paying enterprise rates for the whole thing.
The failure is structural. A demo shows you a vendor's best data in a vendor's best workflow. It cannot tell you what happens when you load 300 messy transcripts in three languages, disagree with the coding, and need to explain a finding to a sceptical commercial director.
This checklist is the alternative. It is deliberately vendor-neutral, because its main use is to be forwarded inside your organisation to people who will assume any framework we wrote is rigged. Where our own product has a relevant answer, we say so and flag it as ours.
Why do software evaluations go wrong?
Three predictable mistakes.
Feature checklists reward breadth over depth. A grid of ticks favours the tool that does thirty things adequately over the one that does the six things you actually do extremely well. Insights work is narrow and deep: you code, you compare, you evidence, you report. A tool that does those four things properly beats one with a video highlight-reel editor you will never open.
Demos test the vendor, not the tool. The person driving has run that flow four hundred times on data chosen to flatter it. What you learn is that the vendor's employee is fast. The only meaningful test is you, on your data, doing your tasks, timed.
Procurement is treated as a final hurdle rather than a criterion. Teams pick a tool and then discover the DPA is unacceptable, the hosting is in the wrong jurisdiction, or the training-data policy will not clear legal. Bring the governance questions to the first conversation, because they eliminate vendors faster than anything else and there is no point evaluating features on a tool your organisation cannot use.
In many ways the market for AI assisted qualitative analysis resembles the Wild West. The category is new, growing and evolving. Many try to enter it with half-baked solutions that are thin wrappers on top of AI, or try to add a sprinkle of AI features to a legacy product just to tick the boxes.
As a buyer it makes sense to get smart and consider the options out there wisely. Here is what we would evaluate.
The 7 evaluation domains
| # | Domain | What it determines | Suggested weight |
|---|---|---|---|
| 1 | Analytical capability & rigour | Whether it does the analysis you actually do | 20% |
| 2 | Traceability and verification | Whether findings survive a challenge | 20% |
| 3 | Control and editability | Whether the analysis ends up yours | 15% |
| 4 | Data governance and security | Whether you are allowed to use it | 15% |
| 5 | Collaboration and workflow fit | Whether the team adopts it | 10% |
| 6 | Export and lock-in | What happens when you leave | 10% |
| 7 | Commercial model | Cost flexibility and TCO | 10% |
Weights are a starting point. Adjust them, but be wary of moving traceability or control below 15%, because those two determine whether the tool produces defensible research or plausible output. Everything else is convenience by comparison.
1. Analytical capability
The question is not "does it have AI" but "does it support the analysis my studies require". Most tools are optimised for one shape of work: video-first UX research, survey open text, academic manual coding, or large-corpus text analysis. Buying against the wrong shape is the most common expensive mistake.
Tools that start from a specific channel or use case are often optimised well for that end-to-end workflow which makes them compelling. But the actual analytical capability might just be an afterthought (we are aware of tools that literally feed the pages to a chatbot and ask it to "identify the themes"... something prone to bias and hallucinations and not justifying the price), and the tool is unusable for any instances where you need to deviate from the locked process (e.g, bring existing transcripts to your project instead of starting with recruiting respondents and setting up the interviews).
Ask:
- Can I apply my own existing coding framework, or only accept generated categories?
- Can I run inductive and deductive analysis on the same corpus?
- What is the practical limit on corpus size, and what happens at 500 documents?
- How does it handle mixed document types in one project: transcripts, survey verbatims, tickets, call recordings?
- What languages, and can it analyse a mixed-language corpus in one project?
- Can I restructure a coded corpus under a different framework without recoding?
- How does the coding actually work: does the tool build a tag structure and apply it to content, and are the outputs organised around those tags? Can I look at coded documents to know what was coded where?
- Which steps are AI-led and which are human-driven, and what checks sit between them? Can the human easily challenge and edit decisions without being thrown to full manual mode afterwards?
- Which models does it use, and is a model chosen per step or is everything run through one?
- Do my researchers need to be technical to get good results out of it?
That second-to-last question separates vendors who have built a workflow from vendors who have wrapped a chatbot. A considered answer describes a sequence of narrow steps with a verification stage between them, and names the methodological tradition it was modelled on. A vague answer about "advanced AI" is telling you there is no workflow to describe. The last question matters for adoption: if getting a decent result depends on prompt craft, the tool will be used well by one person and badly by everyone else.
Test in the pilot: load a corpus with at least three document types and run the same question two ways, once letting the tool generate the structure and once imposing yours.
2. Traceability and verification
This is where tools separate most sharply, and it is the domain a stakeholder challenge will find. When a commercial director says "who actually said that", you need an answer in seconds.
Ask:
- Show me a claim in a report, and take me to the exact passage in the source document.
- Are quotes verified against the source, or generated by the model? What stops a fabricated quote?
- Show me a document and tell me which passages were used and which were ignored.
- If I add twelve documents, does the existing analysis update?
- Does the tool ever sample my corpus rather than analysing all of it, and does it tell me when it does?
That last question matters more than it sounds. Many tools retrieve a subset of your documents through an embedding search and analyse that, without reporting which documents never entered the analysis. The output looks corpus-wide and is not. We set out the underlying argument in design criteria for AI-assisted qualitative analysis tools and the specific failure modes in hallucinations, context and the black box.
3. Control and editability
Cheap to test and highly predictive of whether the tool gets used. If revising the AI's output is slow, researchers stop revising it, and the analysis becomes whatever the model produced.
Ask:
- Rename a category, merge two, split one, and delete a category and assign twenty passages from it. Time each.
- Can I undo? What is not reversible?
- Can I code manually alongside the AI, or is manual coding a second-class path?
- Who has the final say on category structure, me or the system?
Test in the pilot: deliberately disagree with the tool. Restructure its output into the framework you would have built by hand and record how long it took. This single test predicts adoption better than any other.
4. Data governance and security
Bring this to the first call. It is the fastest way to shorten a vendor list, and the answers are either documented or they are not.
Ask:
- Where is data stored and processed? If the customer is in the EU, can you guarantee the data and processing will be within EU?
- Is customer data used to train models, ever, including in aggregate?
- What is the retention policy, and can I hard-delete a project?
- Which subprocessors touch the data, and is there a current list?
- When an AI feature calls a third-party model, whose data processing terms apply, and can that be restricted or switched off for a sensitive project?
- Will you sign our DPA, or do you have a GDPR Article 28-compliant one?
- Is single-tenant or private deployment available, and at what price?
- Do you support integrated anonymisation or pseudonymisation steps in the analysis workflow?
On that last point, stripping names is rarely sufficient. Participants remain identifiable through role, location, employer, relationships or a combination of demographics, which is why pseudonymisation needs to be consistent across a corpus rather than per document.
Where Skimle stands, since we are one of the vendors you might be scoring: all data is hosted and processed inside the EU, on AWS, in an environment configured by a certified third party using row-level security and two-factor authentication. Customer data is never used to train models. DPAs are available, and our privacy policy and terms of service are the documents to start from. Some public sector customers run Skimle in their own private cloud, and a few are exploring running it against locally hosted models. Pseudonymisation is built in.
5. Collaboration and workflow fit
The domain that decides whether the tool is used by a team or by one enthusiast.
Ask:
- How does a second researcher review or challenge someone else's coding?
- Can a stakeholder be given read access without a full seat?
- Is there an API and/or MCP interface for anything we need to automate?
- Is there a viewer or observer role, or does everyone who needs visibility need a full paid seat?
- What does onboarding look like for someone who uses the tool twice a year, as against someone in it weekly?
- What onboarding do you actually provide: live sessions, documentation, in-product guidance, or a link to a help centre?
The viewer question decides more of your bill than it looks. If findings are only useful once the product team and two executives can read them, a tool with no read-only role turns five researcher seats into eleven. Ask whether a read-only export or shareable report exists as an alternative, since that solves the same problem without the seats.
Be sceptical of repository features you will not maintain. A research repository that nobody curates becomes a folder with extra steps. Building a research repository that people actually use covers what makes the difference.
6. Export and lock-in
Ask before signing, not at renewal. The answer tells you how confident the vendor is that you will stay by choice.
Ask:
- Can I export coded data with codes intact, not just a PDF of a report? For example, do you support open source REFI-QDA format, so the material can move to e.g., NVivo, MAXQDA or ATLAS.ti?
- Do you support easily readable formats like Word documents with all text coded or Excel table with quotes per theme?
- If we cancel, what happens to the data and for how long can we retrieve it?
- Are exports available on all plans or gated to enterprise?
REFI-QDA support is a useful signal well beyond its practical use. A vendor that implements the interoperability standard is telling you it does not rely on trapping your data.
7. Commercial model
Three models dominate, and they scale very differently.
| Model | Scales with | Suits | Watch for |
|---|---|---|---|
| Per seat | Headcount | Stable teams, heavy collaboration | Stakeholders you want to give read access to |
| Credit or volume | Analytical throughput | Variable study volume, small teams | Cost spikes on a big study |
| Flat platform fee | Nothing, until tiers | Predictable budgeting | Paying for capacity you do not use |
Model the cost at three points: today, at double your current volume, and with five stakeholders added as viewers. Per-seat pricing that looks cheap for four researchers gets expensive the moment you want the product team to read findings.
Seat elasticity is the question project-based teams forget to ask. If your live analysis work swings between nobody and six-plus researchers depending on what is in the pipeline, annual per-seat pricing is a poor fit and you will spend months paying for seats nobody opens. Ask specifically:
- How often can seats be added and removed: any time, monthly, or only at renewal?
- What happens mid-cycle? Is a seat added in week two charged pro rata or as a full month?
- Is there a minimum commitment, or a floor on seat count?
- Can spend be attributed to a specific project rather than smeared across the year?
That last question matters most for agencies and consultancies who recharge to clients. A hybrid structure answers it better than any pure model: a light per-seat subscription so the whole team keeps access and a shared allowance, plus consumption packs you draw down during the months a large project is running. The subscription keeps the tool available; the packs attach cost to the project that caused it.
Note also that tooling should be a small share of total research spend. Our breakdown of what a qualitative study actually costs puts software at 5 to 15% of a study; if an evaluation is consuming attention out of proportion to that, the fieldwork budget is the thing that deserves the scrutiny. You want to make sure the tool you pick is helping your research dig deeper and faster to their data because the quality of outcomes and the Total Cost of Ownership to create them is what actually matters.
How to run a pilot that tells you something
A two-week pilot with a defined protocol is worth more than six demos. The protocol matters, because an unstructured trial becomes an unstructured play.
- Use a corpus you already analysed. You know what the answer should be, which is the only way to judge quality. Twenty to forty documents is enough.
- Write three tasks before you start. One coding task, one comparison task ("how do segment A and B differ on X"), one evidence task ("find the three strongest quotes for this claim").
- Time each task, including the correction work afterwards. Uncorrected output is not a deliverable.
- Deliberately disagree. Restructure the categories into your framework and record the friction.
- Break it. Load a badly formatted transcript, a document in another language, and a file twice the size of the rest. See what it reports rather than what it silently drops.
- Have a second person review the first person's output. If they cannot follow how a finding was reached, no stakeholder will.
- Score against the seven domains while the friction is fresh.
Give the same protocol to every vendor. A tool that wins on a comparable task list has actually won.
Two things to establish before you start. First, ask whether the vendor will support running one live project in parallel with your incumbent. A parallel run on real work is the strongest evidence available, because you get the same corpus analysed both ways and a direct comparison of the output rather than an impression of the interface. A vendor confident in the product will agree readily; reluctance is itself informative.
Second, check whether a self-serve trial exists before you book anything. Several tools in this category offer a free tier or trial with no card required, which lets one researcher form a view in an afternoon and means the formal pilot starts from an informed shortlist rather than a vendor's suggestion.
5 red flags
No answer on quote verification. If a vendor cannot explain what prevents a fabricated quote reaching your report, the answer is nothing.
Cannot show what was not coded. A tool that only shows you what it found cannot help you notice what it missed, and systematic omission is a harder failure to catch than a wrong code.
Export gated behind the top tier. Charging for the ability to leave is a statement about the product's confidence.
Vague training-data language. "We take privacy seriously" is not a policy. "Customer data is never used to train models, and here is the DPA clause that says so" is.
Results that change materially on identical inputs. Some variation is inherent in these systems. Wholesale reshuffling on a rerun means you cannot build a repeatable process on it, and a repeatable process is what a methods section describes.
A scoring template you can copy
Score each domain 1 to 5, multiply by the weight, and total. Force yourself to write one sentence of evidence per score, drawn from the pilot rather than from the demo.
| Domain | Weight | Vendor A | Vendor B | Vendor C |
|---|---|---|---|---|
| Analytical capability | 20% | |||
| Traceability and verification | 20% | |||
| Control and editability | 15% | |||
| Data governance and security | 15% | |||
| Collaboration and workflow fit | 10% | |||
| Export and lock-in | 10% | |||
| Commercial model | 10% | |||
| Weighted total | 100% |
A vendor scoring below 3 on governance is disqualified regardless of total, because that is a permission question rather than a preference. The same applies to traceability for any team whose work is challenged by stakeholders or reviewed externally.
For the current landscape, our complete comparison of qualitative data analysis tools and interview analysis software comparison for 2026 cover the main options, and qualitative research tools for market research agencies is the agency-specific version. If you work in consumer insights, the customer and market researchers use-case page covers the workflow this sits inside.
Frequently asked questions
Should we run one pilot or several in parallel?
Parallel, with the same corpus and the same three tasks. Sequential pilots are unreliable because you get better at the tasks as you go, which flatters whichever tool you tried last. Parallel pilots cost more calendar attention and produce a comparison you can actually defend.
What if procurement demands certifications no vendor in this category has?
Common, and worth surfacing early. The qualitative analysis category is smaller than general SaaS and certification coverage is patchier than procurement expects. The productive route is to establish what risk the certification is standing in for, usually data residency, access control and breach handling, and to satisfy those directly through the DPA, hosting commitments and a security questionnaire. Certification is evidence, not the requirement itself.
How do we document the decision for an audit trail?
Keep the scoring matrix, the pilot task definitions, and the evidence sentence behind each score. If your organisation reports on research methods externally, this also feeds the methods documentation: reporting standards work such as COREQ+LLM is moving towards requiring an account of which tool was used, what it did, and how outputs were verified. A completed scoring matrix answers most of that for free.
Is a general-purpose AI assistant a legitimate option to include?
Include it as a baseline, because someone will ask. It will score adequately on capability and poorly on traceability, verification, governance and auditability, which is exactly the point the matrix exists to make. Scoring it fairly alongside the specialists is more persuasive than excluding it.
Want to run this checklist against us? Try Skimle for free and use the pilot protocol above on your own corpus. Bring a study you have already analysed, run the three tasks, and time the corrections.
Related reading: See design criteria for AI-assisted qualitative analysis tools for the argument behind the traceability and control domains, what a qualitative study actually costs for where software sits in a research budget, and responsible AI in qualitative market research for the policy side.
About the author
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile
Disclosure: Skimle is one of the tools this checklist can be used to evaluate. The framework is written to be applied to us on the same terms as anyone else. Bring it on!
Sources
- ISO 20252 Standard (Market Research) - Insights Association
- Extension of COREQ to Large Language Models (COREQ+LLM): Protocol for a Multiphase Study, Fehring and colleagues, 2025, JMIR Research Protocols
- AI in Market Research: Five rules to live by, Andersson, August 2025, Research World
- User research budget planning: costs, allocation and ROI, March 2026 - CleverX



