Data collection methods in qualitative research: 8 approaches compared

Qualitative data collection includes interviews, focus groups, observation, and more. This guide compares 8 methods on depth, feasibility, and analysis needs.

Cover Image for Data collection methods in qualitative research: 8 approaches compared
Share this article:

The main qualitative data collection methods are semi-structured interviews, focus groups, in-depth interviews, ethnographic observation, document analysis, online and AI-conducted interviews, diary studies, and case studies. Each method captures different types of evidence, suits different research questions, and creates different demands on analysis time and skill. Choosing the right method early shapes everything that follows.

Across academic fields, the prevalence of qualitative methods has at least doubled since 1996, with interviews now recognised as by far the most commonly used approach. But the rise in popularity has also brought more choice. Researchers designing studies today have access to methods that barely existed a decade ago, including AI-conducted interviews that can gather responses from hundreds of participants simultaneously.

This guide covers all eight primary data collection methods, when each is most appropriate, what sample sizes to plan for, and how each method shapes the analysis work that follows. There is also a decision table at the end to help you choose.

What is qualitative data collection?

Before comparing methods, it is worth being precise about what qualitative data collection means. Qualitative research gathers non-numerical data: transcripts, observation notes, documents, diary entries, images, and other material that captures meaning, experience, and context. The goal is to understand how and why, not just how many.

Unlike quantitative data collection, where the instrument (a survey, a sensor, a measurement tool) largely determines quality, qualitative data collection is shaped heavily by the skill of the researcher and the appropriateness of the method to the question being studied. A poorly designed interview guide will produce surface-level responses no matter how sophisticated the analysis software. A method mismatched to the research question will produce data that cannot answer it.

The eight methods below represent the primary ways researchers generate qualitative data. Many studies combine two or more, particularly case studies, which are explored in the final section.


1. Semi-structured interviews

Semi-structured interviews are the workhorse of qualitative research. The interviewer prepares a topic guide with core questions and prompts, but the conversation is allowed to follow threads that emerge naturally. This balance between structure (ensuring comparability across participants) and flexibility (allowing depth and unexpected findings) makes semi-structured interviews suitable for a very wide range of research questions.

When to use them: When you need to understand individual perspectives, experiences, or reasoning in depth; when your research question is exploratory; when you are working across a moderately heterogeneous population where follow-up questions will differ between participants.

Sample size: Most semi-structured interview studies reach thematic saturation with somewhere between 12 and 30 interviews, though this depends on population homogeneity and the scope of the research questions. A useful finding from recent saturation research is that theme saturation is typically reached around 9 interviews, while meaning saturation, where no new interpretive dimensions are emerging, typically requires closer to 24. See our guide on qualitative research sample size for more detail.

Analysis complexity: Moderate to high. Good semi-structured interview data requires careful thematic analysis, which typically means reading transcripts multiple times, coding by hand or with software, and iterating through themes. See how to analyse interview transcripts for a step-by-step process.

Practical considerations: Interviews typically run 45-90 minutes. Transcription adds significant time; a one-hour interview produces roughly 8,000-10,000 words of transcript. AI transcription tools have reduced this burden considerably. See best AI transcription tools for research in 2026 for a current comparison. For guidance on building the conversation guide, see how to write a perfect interview guide.


2. Focus groups

Focus groups bring together 6-8 participants to discuss a topic together, usually for 60-120 minutes. The defining feature is group dynamics: participants respond to each other, build on each other's ideas, and sometimes challenge each other's assumptions. This makes focus groups particularly powerful for studying social norms, shared meanings, and how people articulate positions in a group context.

When to use them: When you want to understand how people think and talk about a topic collectively; when social consensus or group negotiation of meaning is relevant to your research question; when you need to generate ideas or test concepts with a target group.

Sample size: Most focus group studies run 3-6 groups, with each group comprising 6-8 participants. Saturation research suggests that 4-8 focus group discussions reach saturation for relatively homogeneous populations. Total participant numbers therefore typically range from 18 to 48.

Analysis complexity: Moderate, with some specific challenges. Focus group transcripts are harder to code than individual interview transcripts because group dynamics, turn-taking, and overlapping speech create analytical layers that one-to-one interviews do not. See how to analyse focus group transcripts for method-specific guidance.

When not to use them: Focus groups are not well suited to sensitive or stigmatised topics where participants will not speak freely in front of others, to heterogeneous groups where status differences may inhibit junior or marginalised voices, or when you need detailed individual experience rather than shared meaning. See focus groups vs individual interviews for a fuller comparison.


3. Unstructured (in-depth) interviews

Unstructured or in-depth interviews give the participant maximum freedom to direct the conversation. The researcher may begin with a single opening question ("Can you tell me about your experience of...") and then follow wherever the participant leads, using active listening and minimal prompts. This approach is associated with phenomenological and narrative traditions in qualitative research.

When to use them: When capturing the participant's own frame of reference is essential; when you want to understand lived experience as the participant constructs it; when imposing a researcher-determined structure would distort the data. IPA (interpretive phenomenological analysis), narrative inquiry, and oral history research typically use this approach.

Sample size: Smaller than semi-structured interviews. IPA studies, for example, typically use 4-10 participants. The depth of a single in-depth interview compensates for smaller sample size. See our guide on interpretive phenomenological analysis for specific sample size guidance.

Analysis complexity: High. The richness of in-depth interview data is also its analytical challenge. Each transcript may contain multiple layers of meaning, narrative structure, and reflexive positioning. Analysis is typically slower and more interpretive than semi-structured interview analysis.

Practical considerations: In-depth interviews often run 90 minutes to 2 hours or longer. They require skilled interviewers who are comfortable with long silences, can resist filling conversational space, and can track complex narratives in real time. See informal interview questions and unstructured interview design for question examples.


4. Ethnographic observation

Ethnographic observation involves the researcher entering a setting and systematically observing behaviour, interactions, and context over time. There are two broad forms: participant observation, where the researcher takes an active role in the setting (working alongside participants, for example), and non-participant observation, where the researcher observes without joining in.

Ethnography remains the gold standard for understanding what people actually do rather than what they say they do, a gap that is often significant. Interview data captures reported experience; observation data captures enacted behaviour.

When to use them: When the research question concerns practice, routine, or context that participants may take for granted or struggle to articulate in an interview; when the setting itself is part of what you are studying; when you are interested in tacit knowledge or informal processes.

Sample size: Ethnography is typically measured in time rather than participants. Field studies might run from a few weeks to several years, depending on the phenomenon and the tradition. Focused ethnographies in applied settings (healthcare, workplace studies) might run 3-6 months.

Analysis complexity: High. Ethnographic data includes field notes, photographs, informal conversations, documents collected in situ, and sometimes interview transcripts. Integrating these different data types into a coherent analysis requires sustained interpretation. According to a systematic analysis of qualitative research across academic disciplines, ethnography is the least common of the four main qualitative methods and is concentrated primarily in arts, humanities, and social sciences.

Practical considerations: Participant observation raises distinct ethical issues around informed consent, role clarity, and the researcher's influence on the setting. Non-participant observation is less disruptive but may miss tacit or behind-the-scenes behaviour.


5. Document analysis

Document analysis uses existing textual material, rather than data generated specifically for the study, as its primary evidence base. The documents might include policy papers, meeting minutes, organisational reports, published media, historical records, interview transcripts from previous studies, or any other written or recorded material relevant to the research question.

When to use them: When the research question concerns institutional processes, historical change, official discourse, or policy; when access to participants is limited or impossible; when you want to triangulate findings from other methods with documented evidence; when studying what organisations produce and communicate officially.

Sample size: Measured by document count rather than participants. A systematic document analysis might code 50 policy documents, 200 news articles, or a corpus of annual reports spanning 20 years. The size depends on the scope of the question and the richness of individual documents.

Analysis complexity: Moderate to high, depending on the approach. Content analysis of documents can be structured relatively systematically. Critical discourse analysis, which examines power relations and ideological assumptions embedded in texts, is significantly more interpretive. Tools like Skimle can accelerate document analysis considerably, particularly for large corpora where manual reading would take weeks. See our AI document analysis guide for practical approaches.

Key advantage: Non-reactive. Unlike interview data, documents are not produced in response to the researcher's presence, which removes some types of social desirability bias. They also represent what was actually communicated or decided, not retrospective reconstruction.


6. Online and AI-conducted interviews

Online qualitative interviews, conducted via video conferencing or text-based platforms, have become standard since 2020. Research comparing in-person and online qualitative interviews finds that online interviews are broadly equivalent in data quality for most research purposes, with advantages in geographic reach, scheduling flexibility, and lower cost.

AI-conducted interviews represent a more significant methodological development. Tools like Skimle Ask can deliver asynchronous text-based or conversational interviews to large numbers of participants simultaneously, with AI probing for elaboration and following up on responses dynamically. This shifts the feasibility frontier: a study that would require 300 interviewer-hours to conduct manually might be completed in days.

When to use them: Online video interviews are suitable for most research that would previously use in-person interviews. The exception is research where physical presence, embodied interaction, or observation of non-verbal behaviour is analytically important.

AI-conducted interviews are particularly suited to: large-sample qualitative studies where breadth matters; research in which anonymity encourages more candid disclosure; customer and employee research where scale and speed are priorities; and educational or commercial contexts where access to skilled interviewers is limited.

Sample size: Online interviews follow the same saturation logic as in-person semi-structured interviews (typically 12-30 for themed analysis). AI-conducted interviews can scale to hundreds or thousands of participants, enabling qualitative depth at a scale previously only possible with surveys. See the Skimle Ask documentation for how AI interview projects are structured.

Analysis complexity: Moderate. AI interview data tends to be more structured than in-person interview transcripts because the conversational flow is guided. This can make analysis more tractable. Large-scale AI interview datasets are well suited to AI-assisted thematic analysis.

Considerations and limitations: AI-conducted interviews may produce shallower responses than skilled human interviewers for complex or emotionally sensitive topics. The technology still performs best when questions are clear and the topic does not require extended emotional rapport. Researchers should pilot AI interview instruments with a small sample before full deployment.


7. Diary studies and experience sampling

Diary studies ask participants to record their experiences, thoughts, or behaviours at regular intervals over a defined period, capturing phenomena as they unfold rather than through retrospective recall. Experience sampling method (ESM) is a variant that prompts participants at random or semi-random moments during the day, typically via smartphone, to report on what they are doing, thinking, or feeling in that moment.

The defining methodological advantage is reduced recall bias: participants record close to the event, before memory reconstructs and edits the experience. Research on ESM confirms that this temporal proximity produces data that diverges meaningfully from what participants would report in a retrospective interview.

When to use them: When temporal sequence or change over time matters; when phenomena are distributed across many small episodes rather than concentrated in notable events; when studying everyday experience, routine behaviour, emotion, or wellbeing; when recall bias in retrospective interviews would compromise data quality.

Sample size: Diary studies are typically smaller in participant number but produce large amounts of data per person. A study running for 3 months with 20 participants who each submit weekly entries produces 240 data points. Attrition is a real concern: researchers should build extra recruitment into the sample plan, as completion rates in published diary studies typically run from 60% to 85%.

Analysis complexity: Moderate to high. The longitudinal structure of diary data enables temporal analysis that other methods cannot: tracking change, identifying turning points, and comparing within-person patterns over time. But this power comes at analytical cost. Integrating entries across participants and time points requires careful data management.

Practical considerations: Diary study data collection is demanding on participants. Sustained engagement over weeks or months requires good design (brief, clear entry formats), regular researcher contact to maintain motivation, and authentic acknowledgement of participant contributions. Over-demanding diary protocols produce attrition and incomplete data.


8. Case studies

Case study research investigates a phenomenon within its real-world context, typically using multiple sources of evidence in combination. A case might be an organisation, a project, a programme, a community, a policy, or an individual. Robert Yin's foundational work defines the case study as appropriate when the "how" and "why" of a contemporary phenomenon is the focus and when the researcher has little control over events.

The crucial point about case studies is that they are not a single method but a research design that combines multiple methods. A typical case study might include semi-structured interviews with key informants, document analysis of organisational records, and observation of meetings. This triangulation across data sources is what gives case studies their distinctive analytic power.

When to use them: When you need to understand a complex phenomenon in depth within its context; when the boundaries between the phenomenon and its context are not clearly separable; when multiple perspectives on the same events are analytically important; when you want to develop or test theory against a real-world instance.

Sample size: Measured by number of cases. Single-case studies justify depth; multiple-case studies (typically 2-6) enable replication logic and comparison. Within each case, the number of interviews, documents, and observations depends on what access the case permits and what questions require.

Analysis complexity: High. Integrating multiple data types, multiple informants with potentially conflicting accounts, and contextual documentation into a coherent analysis is demanding. The risk of selective use of evidence is real and must be managed through systematic coding and transparent reporting.

Practical considerations: Case study access depends on gatekeepers. Negotiating entry, managing ongoing relationships with participants, and navigating organisational sensitivities are as important as the research skills. Allow more time than you think for access negotiation, particularly in institutional settings.


How do the 8 methods compare?

The table below summarises the key dimensions for choosing between methods.

MethodBest forTypical sample sizeAnalysis complexityAI-assisted?
Semi-structured interviewsIndividual experience, perspectives, reasoning12-30 participantsModerateYes (transcription, coding, theming)
Focus groupsShared meaning, group norms, idea generation3-6 groups (18-48 total)ModerateYes (transcription, group-level coding)
Unstructured/in-depth interviewsLived experience, narrative, phenomenology4-10 participantsHighPartial: AI handles transcription; deep analysis is researcher-led
Ethnographic observationBehaviour, tacit knowledge, contextWeeks to years of fieldworkHighLimited: field notes can be coded in Skimle; observation itself is manual
Document analysisInstitutional discourse, historical change, policy20-500+ documentsModerateYes (strong fit for AI-assisted corpus analysis)
Online/AI interviewsLarge-scale qualitative data, speed, anonymity20-500+ participantsModerateYes (purpose-built for AI-assisted collection and analysis)
Diary studies / ESMTemporal change, everyday experience, recall-sensitive topics10-30 participantsHighPartial: AI can code entries; study management is manual
Case studiesComplex phenomena in context, theory building1-6 cases (multiple data sources per case)HighYes (AI assists with document and transcript coding within each case)

Which method should you choose?

The single most important question is: what kind of evidence does your research question require?

If the question is "how do employees experience organisational change?", semi-structured or in-depth interviews are natural choices. If it is "how do teams actually work together during change?", observation would produce more valid data than asking people to describe what they do. If it is "what have organisations officially communicated about their change processes?", document analysis is the right tool.

Secondary considerations include:

Access. Can you get into the settings or reach the participants the research requires? Document analysis and AI-conducted interviews have lower access barriers than ethnography or in-person focus groups.

Time and resources. Ethnography and in-depth case studies are resource-intensive. Diary studies require participant commitment over time. Online and AI-conducted interviews can compress the data collection phase dramatically.

Analysis capacity. Be honest about what you can analyse. A corpus of 300 AI interview transcripts is more manageable than 300 one-to-one interview recordings, but it is still a large analysis task. Plan your analysis approach before you collect data.

The nature of the phenomenon. Distributed, everyday experience is poorly captured by retrospective interviews; diary methods or observation suit it better. Contested or politically sensitive topics may be better explored in one-to-one interviews than focus groups. Institutional or policy phenomena invite document analysis.

For guidance specific to academic contexts, see how Skimle supports academic researchers. For customer research applications, the process differs in important ways from academic research; see our customer discovery interviews guide for a practical workflow.


How does method choice affect analysis?

This is an underappreciated dimension of research design. The method you use to collect data shapes what is possible in analysis, and researchers sometimes find themselves with data that cannot answer their question because the collection method constrained what was said or observed.

Interview data (all variants) is suited to thematic analysis, grounded theory, IPA, narrative analysis, and discourse analysis. The transcript is the primary analytical unit. AI tools handle transcription efficiently; thematic coding is where AI assistance adds most value.

Focus group data requires attention to group dynamics as well as content. Who said what, in response to whom, and how the group reached (or failed to reach) consensus are all analytically relevant. Standard thematic analysis needs adapting for this.

Observation and ethnographic data are typically analysed through ethnographic writing, grounded theory, or interpretive frameworks. Field notes are not transcripts; they are already partially interpreted records that carry the researcher's framing.

Document data suits content analysis, critical discourse analysis, and narrative analysis. The analytical unit might be a sentence, a paragraph, a document, or a corpus, depending on the question.

Diary and ESM data enable temporal and within-person analysis that other methods cannot. Mixed-methods approaches that combine quantitative ESM data (frequency, timing, intensity ratings) with qualitative diary entries are increasingly common.

Multi-method case study data requires a synthesis approach. Researchers typically analyse each data source separately first, then integrate findings across sources, looking for convergence and productive divergence.

For all transcript-based methods, AI-assisted analysis has reduced the time burden substantially. Skimle's approach, for example, maintains a complete audit trail from theme back to source excerpt, which is important for methodological transparency and researcher confidence. The key discipline is ensuring that the speed of AI-assisted analysis does not outrun the depth of researcher interpretation.


Frequently asked questions

What is the difference between primary and secondary qualitative data collection?

Primary qualitative data collection means generating new data specifically for your research: conducting interviews, running focus groups, doing observation. Secondary data collection uses existing data gathered for another purpose: archived interviews, published reports, government records, social media posts. Document analysis is the primary secondary data collection method in qualitative research.

Can you use multiple data collection methods in the same qualitative study?

Yes, and this is common. Triangulation across methods, collecting both interview data and documentary evidence, for example, strengthens the validity of findings by showing that different methods converge on the same conclusions. Case studies are specifically designed as multi-method approaches. The caution is that each additional method multiplies analysis work; be realistic about what you can manage within the resources available.

How does the choice of data collection method affect ethical requirements?

All qualitative data collection requires informed consent, but what that means varies by method. Observation in public settings may not require consent from everyone observed, while participant observation in a workplace almost certainly does. Document analysis of publicly available materials typically requires no individual consent. AI-conducted interviews require the same ethical standards as human-conducted ones: clear consent, data protection, the right to withdraw, and honest representation of the research purpose.

How long does qualitative data collection take?

It varies enormously. Online AI interviews can collect responses from hundreds of participants in days. Ethnographic fieldwork runs for months or years. A typical semi-structured interview study collecting 20 interviews, transcribing, and reviewing them might take 8-12 weeks from first recruitment contact to final transcript. Build more time than you expect into research timelines, particularly for recruitment, which almost always takes longer than planned.

Which qualitative data collection method produces the most credible evidence?

No single method is universally superior. Credibility depends on fit between method and question. Observation produces more credible evidence about what people do than retrospective interviews do. In-depth interviews produce more credible evidence about individual lived experience than a survey open-text response. Case studies produce credible evidence about complex phenomena in context. The appropriate standard is methodological coherence, whether the data collection method can actually generate the evidence needed to answer the research question.


Ready to collect and analyse your qualitative data?

Try Skimle for free and work with your interview transcripts, focus group recordings, documents, and diary entries in a single platform. Skimle handles transcription, AI-assisted thematic coding, and synthesis, with a full audit trail from every finding back to the source excerpt.

Related reading:


About the authors

Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile

Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile


Sources

Dig deeper to your data with Skimle

Skimle collects, analyses and categorises interviews, survey responses, reports and other qualitative data automatically. Our modern qualitative analysis software combines a rigorous and transparent workflow with the speed of AI.

Upload text or audio, remove sensitive data with Skimle Anonymise, automatically create categories and sub-categories, explore the data across documents and export the data to seamlessly fit your workflow. Built by professionals for professionals, with full privacy and GDPR compliance.

Free trial · No credit card required · Full plans from €20/month