Reflexivity on the role of AI in research: am I staying in form AND raising the bar?

Reflecting on AI is a good practice, whether you use it or not. I argue to consider both the risks of losing ownership, but equally the opportunities of improving quality.

Cover Image for Reflexivity on the role of AI in research: am I staying in form AND raising the bar?
Share this article:

I propose that reflexivity about AI in qualitative research needs two sides. Staying in form asks whether the researcher still does the thinking, stays close to context and participants, and challenges AI output. Raising the bar asks whether AI is used to find counter-evidence, spot patterns manual coding misses, test alternative framings and free time for participants. A reflexive journal should record both.

A high jumper who only watches their run-up never finds out how high they can go. One who only watches the bar eventually knocks it off. Most of the current guidance on AI and qualitative research covers the first worry, and for good reason. My argument in this post is that the second deserves equal space in your reflexive practice, and I set out the questions I use for both.


Why does reflexivity now need to cover AI?

Reflexivity in qualitative research has always meant examining how the researcher shapes the analysis: their background, assumptions and position relative to the people they study. When a generative AI tool codes transcripts, proposes themes or drafts a summary, it becomes part of that shaping. Kathryn Wheeler, comparing how CAQDAS packages MAXQDA and NVivo and ChatGPT handled more than 1,300 open-ended survey responses, describes reflexivity as now "a distributed practice, shared (often unequally) across researchers, tools, and infrastructures".

The tools are already in the room. According to Wiley's ExplanAItions survey of 2,430 researchers, 84% of researchers used AI tools in 2025, up from 57% a year earlier. In the same survey, concern about inaccuracies and hallucinations rose from 51% to 64%.

The field is also sharply divided on whether AI belongs in interpretive work at all. In a position statement in Qualitative Inquiry, 419 experienced qualitative researchers from 32 countries, led by Tanisha Jowsey with Virginia Braun and Victoria Clarke among the authors, rejected AI for reflexive qualitative approaches, arguing that methods such as reflexive thematic analysis require "a subjective, positioned, and reflexive researcher". Whatever your own position, the debate makes one thing clear: if you do use AI, you need to be able to explain how it affected your analysis, and reflexivity is the discipline for doing that.

Margaret Roller makes this point in a recent Research Design Review article on reflexivity and AI. Working from the Total Quality Framework she developed with Paul Lavrakas, she writes:

"It cannot be overstated that researchers' reflections on the Credibility, Analyzability, Transparency, and Usefulness of their GenAI-enabled research is a necessary ingredient to the integrity of their outcomes."

She recommends the reflexive journal as the place to do this work, because it lets the analyst see "how they are interfacing with GenAI and the impact this is having on the process and the outcomes."


What does "staying in form" and "raising the bar" mean?

Think of the two sides as two different failure modes. The first is the one most researchers worry about: handing over judgement to a tool and producing work that is fluent but shallow. The second gets less attention: using a powerful tool only to do the old study a bit faster, and leaving the extra depth on the table.

Staying in formRaising the bar
The riskAI takes over the thinking; context, nuance and participants' voices get flattenedAI is used only to speed up the same study, so findings are no better than before
The core questionAm I still the one doing the analysis?Is my analysis better than it would have been without AI?
What it protectsCredibility, integrity, the researcher's interpretive roleDepth, breadth, robustness and usefulness of the findings
Typical evidence in a journalWhere I checked, challenged or overruled the toolWhat I tested, explored or discovered that I would not otherwise have done
What failure looks likeThemes that read well but cannot be traced back to the dataA thesis or report that is indistinguishable from one done in 2019, only delivered sooner

The diagram below illustrates the type of questions to ask. Maybe print it next to your journal or bedstand :)

Reflexivity in AI-assisted qualitative research: staying in form checklist beside raising the bar checklist


Staying in form: which reflexive questions should you ask?

These questions draw on several sources: Margaret Roller's reflexivity questions for AI-assisted analysis and the Total Quality Framework from Applied Qualitative Research Design (Roller and Lavrakas, 2015), Lester and Paulus's four areas for reflection, Wheeler's work on technological reflexivity, and research on how AI affects critical thinking. I have rephrased them and grouped them into five themes.

1. Who is doing the thinking?

  • Whose judgement shaped this theme: mine, grounded in the data and the research question, or the tool's framing?
  • When my interpretation shifted, was it new evidence that moved me, or a persuasive AI summary?
  • Did the tool's defaults (its level of detail, its labels, the order it showed results in) steer my choices without my noticing?

This is the hardest question to answer truthfully, because the effect is gradual. A CHI 2025 survey of 319 knowledge workers by Hao-Ping Lee and colleagues at Microsoft Research and Carnegie Mellon found that higher confidence in AI was associated with less critical thinking, while higher confidence in one's own expertise was associated with more. A change of mind is not a problem in itself. Good analysis changes the analyst's mind all the time. The question is whether the data changed it, or the fluency of the tool's summary. Tool design matters too: Wheeler's comparison of three software packages shows that different architectures nudge researchers toward different analytic strategies, so note which defaults you accepted.

At minimun I would call for the tools to have full two-way transparency in the core: can you trace all summaries back to the original quotes; and can you see from source material what was coded where?

2. Did I verify and challenge the output?

  • Which AI claims did I check against the transcripts, and how?
  • Where did I disagree with the tool because it oversimplified, misread or overreached, and what did I do about it?
  • Are my prompts and settings recorded as methodological decisions, so I could walk a sceptical examiner through every step?

The instruction to "find the main themes" is a choice of granularity and orientation, in the same way a coding framework is, and belongs in the methods record. Verification is only possible if every claim can be traced back to the passages behind it. A theme with no visible evidence trail can be accepted or rejected, but not checked. The 20-question checklist for AI-assisted analysis covers what to verify before anything is published, and our post on bias in AI-assisted qualitative analysis describes the specific distortions to look for.

3. Am I still reading for context?

  • Am I reading whole transcripts, or only the excerpts the tool hands me?
  • What does a summary leave out: who said it, when, in what setting and in what tone?
  • Would participants recognise the circumstances they described in my findings?

Summaries strip context by design. A sentence such as "participants felt unsupported by management" may be accurate across 20 interviews and still miss that half of them meant their line manager and half meant the executive team. Reading full transcripts, not only coded excerpts, remains part of the job. Hamilton and colleagues, comparing human and ChatGPT analyses of interviews with guaranteed income recipients, found that each picked up themes the other missed, and that the AI's main weakness was limited contextual understanding.

4. Have I protected participants' voices?

  • Does each participant's account still read as theirs, or have individual stories been averaged into a composite?
  • What did anonymisation remove that mattered for interpretation?
  • Would taking my interpretations back to participants (member checking) strengthen or test the analysis?

Anonymising transcripts before they reach an AI tool is often an ethical requirement, but it is not neutral. Replacing a hospital name with "[organisation]" can erase the fact that two participants worked at the same ward. Member checking is one way to test whether an AI-assisted interpretation still resonates with the people it describes; Birt and colleagues give a useful critique of when it adds real credibility and when it is a token gesture.

5. Am I still analysing with other people?

  • When did I last talk the analysis through with colleagues, a supervisor or participants?
  • Has working with the tool replaced conversations that used to challenge my reading?
  • How has that changed the analysis and the conclusions?

Qualitative analysis has always been partly social: the debrief after an interview, the argument over a code, the supervisor who asks "where is that in the data?". When a tool can produce a full first pass overnight, those conversations can quietly disappear. Jessica Nina Lester and Trena Paulus, writing for the CAQDAS Networking Project, warn that skipping the struggle with the data can mean "skipping past the part where actual growth occurs". They propose reflecting across four areas (methods, technologies, humans and knowledge), which is a useful wider frame for this whole section.

If you work in academic research and want to see how a tool can support these checks rather than replace them, see how Skimle fits academic research workflows. Every category is traceable to the quotes behind it, and every AI decision can be opened, questioned and changed.


Raising the bar: is AI making your research better?

Staying in form is necessary, but it is a defensive standard. It asks whether AI has made your research worse. That is why I add a second set of questions, because any tool or method should ultimately be chosen by the researcher to produce better outcomes. A researcher who uses AI responsibly but only to save time has passed the first test and skipped the second.

These are the five questions I ask myself.

1. Have I used AI to challenge my thinking and find counter-evidence?

Confirmation bias is the oldest problem in qualitative analysis. Once you have a favourite theme, you start seeing it everywhere. AI is very good at the tedious part of disconfirmation: searching 40 transcripts for every passage that contradicts your emerging argument, or asking "what would someone who disagrees with this theme point to?".

In practice: after drafting each theme, ask the tool explicitly for the strongest counter-evidence and the participants who do not fit, then read those passages in full. In Skimle, agentic analysis produces counter-evidence alongside each analytical theme, so the disconfirming cases are on the page rather than left for you to remember to search for. Treat AI-surfaced counter-evidence like any other AI output: read the passage in context before you let it change an argument.

The framing depends on your paradigm. In a post-positivist design, "counter-evidence" tests whether a theme holds. In reflexive thematic analysis, the same search is better described as looking for divergent accounts and tensions that enrich the interpretation. Either way, the question is whether you actively looked for what does not fit.

2. Have I used AI to discover patterns I might have missed manually?

Differences between segments, shifts over time, a theme that only appears among one group of participants: these are the patterns that manual coding of 30 or more interviews tends to miss, because no one can hold the whole dataset in their head at once. Systematic comparison across participant metadata or over time is exactly the kind of task AI handles well and humans handle slowly.

In practice: once the coding is stable, compare every theme across your key attributes (role, site, cohort, wave) and ask which differences are large and which are noise. Then go back to the transcripts to understand why. Be careful not to turn this into fishing: with 30 interviews, a theme mentioned by five people in one group and two in another is a lead to explore, not a finding. Note in the journal which comparisons you planned and which you discovered along the way. The screenshot below shows a ranked view of how themes differ between groups in Skimle.

Skimle metadata comparison ranking qualitative research themes by how strongly they differ between participant groups

3. Have I used AI to spend more time with people and with the material?

The time saved on transcription and first-pass coding is only valuable if it goes somewhere. The best destination is usually upstream: more interviews, a second round with key participants, more careful data collection, or longer thinking about what the findings contribute to theory or practice.

In practice: write down at the start of the project where the saved hours will go, and check at the end whether they did. If you used to conduct 15 interviews because coding 25 was unmanageable, consider whether your sample size should now be set by saturation rather than by your coding capacity.

4. Have I explored alternative ways to structure the findings?

Initial coding schemes are sticky. By the time a manual codebook has been applied to 20 transcripts, restructuring it feels too expensive, so the first framing often becomes the final one. With AI, re-coding the same data under a different framework takes minutes rather than weeks.

In practice: before settling on a structure, try at least one alternative. Code the data inductively and then against a theoretical framework, or organise findings by process stage instead of by stakeholder. Compare which structure explains the data better. Even if you keep your first scheme, you can now defend it against a real alternative rather than a hypothetical one.

5. Have I used AI to elevate the role of my collaborators, especially junior ones?

In many research teams, junior members spend months on mechanical coding while the interpretive work stays with senior researchers. AI can shift that balance. When first-pass coding is fast, a research assistant or PhD student can spend their time on the interesting parts: questioning themes, writing analytic memos, testing alternative explanations.

In practice: give collaborators shared access to the analysis and a defined interpretive role, such as owning the counter-evidence review for one theme. This also answers the fifth staying-in-form question from the other side: AI can make analysis more collaborative, if the team decides to use it that way. One caution: junior researchers still need to learn to code by hand. Lester and Paulus's point about growth happening in the struggle with the data applies most to them, so have new team members code a few transcripts manually before they review AI-assisted coding.


What does a two-sided reflexive journal entry look like?

Here is a worked example. A PhD researcher is studying how ward nurses experience twelve-hour shift patterns, using 24 semi-structured interviews across three hospitals and reflexive thematic analysis supported by an AI tool. After the first full pass of coding, the journal entry reads:

Staying in form. The tool proposed "fatigue" as a single overarching theme. Reading six full transcripts, I disagreed: two distinct experiences are being merged, physical exhaustion at the end of a shift and the cumulative "never recovering" across a run of shifts. I split the theme. I also noticed that de-identification replaced ward names, which hid that four of the most negative accounts came from the same night-shift team. I restored a pseudonymised ward identifier. I have not discussed the coding with my supervisor for three weeks, which is too long; meeting booked.

Raising the bar. I asked for counter-evidence to my draft theme "shift length erodes care quality" and found five participants who argued the opposite: fewer handovers meant better continuity. This is now a tension in the analysis rather than a footnote. Comparing themes by hospital showed that "rostering fairness" appeared almost entirely at one site, which I would not have noticed coding transcript by transcript. I used the time saved on transcription to schedule follow-up conversations with four participants, two of them from the counter-evidence group.

Neither half would be enough alone. The first shows the researcher is in control of the analysis. The second shows the analysis is stronger than it would otherwise have been, and it gives the examiner something concrete to evaluate in the methods chapter.


How do you build the two-sided check into a research project?

Reflexivity is most useful when it is routine rather than retrospective. The table below suggests when to ask which questions.

StageStaying in formRaising the bar
Before analysisWhat will I read myself in full? What will de-identification remove?Where will the saved time go? Which alternative frameworks will I test?
After first-pass codingWhich AI-proposed codes did I accept, change or reject, and why?What counter-evidence exists for my emerging themes?
During theme developmentAm I still reading full transcripts, or only excerpts?Which segment or time differences have I checked? Have I tried a second structure?
Before writing upCan every claim be traced to quotes I have read? Have I discussed the analysis with others?Is this analysis deeper, broader or more robust than a manual one would have been?
In the methods sectionHow did I verify and challenge the tool?What did AI let me do that I could not have done otherwise?

The last row matters for publication. Most journals now ask for disclosure of AI use, and a disclosure that only lists what the tool did reads as defensive. One that also explains what the tool made possible (more interviews, systematic counter-evidence review, comparison across sites) gives reviewers a reason to see AI use as a strength. Our guide to writing up a thematic analysis covers how to report methods in detail.


What if you have decided not to use AI at all?

That is a legitimate position, and the 419 signatories of the Qualitative Inquiry statement make a serious case for it in reflexive approaches. Reflexivity about the decision still applies. The "raising the bar" questions are not really about AI: they ask whether you looked for counter-evidence, compared across groups, tested alternative structures and gave collaborators an interpretive role. Those are good questions for any analysis. AI simply makes some of them easier to answer, which is why they are worth asking every time a researcher decides how to work.

With AI tools now available and opening up the possibility to spend a larger share of time talking with people and considering the insights, doing transcription and anonymisation automatically, making it faster to explore alternative framings and increasing the throroughness of discovering patterns... deciding not to use AI is a choice you need to justify. I believe the quality bar is going to go up, and researchers need to be able to pass it.

Skimle has argued elsewhere that, used well, AI can strengthen rather than erode analytical thinking. The two-sided journal is how you find out whether that is true for your own project.


Frequently asked questions

What is reflexivity in AI-assisted qualitative research?

It is the practice of examining how an AI tool, alongside the researcher's own background and assumptions, shaped the data analysis and findings. It covers both risks (losing context, accepting output uncritically, drifting away from the research team) and opportunities (finding counter-evidence, discovering patterns, testing alternative structures) and documents both in a reflexive journal.

What questions should I ask in a reflexive journal when using AI?

Ask two sets. To check you are staying in form: am I still doing the thinking, did I verify and challenge the output, am I reading for context, have I protected participants' voices, and am I still analysing with other people? To check you are raising the bar: did I look for counter-evidence, find patterns across segments or time, spend more time with participants, try alternative structures and involve collaborators more?

Is using AI compatible with reflexive thematic analysis?

This is contested. Braun, Clarke and over 400 co-signatories argue it is not, because reflexive thematic analysis depends on a positioned human interpreter. Others argue AI can support the process if the researcher stays responsible for interpretation and documents how the tool was used. If you use AI, a detailed reflexive account is what makes the choice defensible.

How do I report AI use in my methods section?

State which tool you used and for which steps, how you verified and challenged its output, and what it enabled that you could not otherwise have done. Specific examples (a theme you split, counter-evidence you found) are far more convincing to reviewers than general statements that "all AI output was checked".

How often should I write reflexive journal entries about AI?

At minimum after each major analytical step: before analysis, after first-pass coding, during theme development and before writing up. Short entries at each stage are more useful than one long reflection at the end, because they capture decisions while you still remember why you made them.


Ready to make AI raise the bar on your next study? Try Skimle for free and analyse your interviews with every theme traceable to its source quotes, counter-evidence on the page and metadata comparisons built in.

Want to go deeper on reflexivity and AI? Read our guides on reflexivity in qualitative research, the AI qualitative data analysis checklist and how to do thematic analysis with AI.


About the author

Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile


Sources

Dig deeper to your data with Skimle

Skimle collects, analyses and categorises interviews, survey responses, reports and other qualitative data automatically. Our modern qualitative analysis software combines a rigorous and transparent workflow with the speed of AI.

Upload text or audio, remove sensitive data with Skimle Anonymise, automatically create categories and sub-categories, explore the data across documents and export the data to seamlessly fit your workflow. Built by professionals for professionals, with full privacy and GDPR compliance.

Free trial · No credit card required · Full plans from €20/month