Discovering themes in the data using metadata variables - advanced analysis with Skimle

How to use Skimle's metadata fields and Visualisations to find where segment, time period or any other variable explains real differences in what people say.

Cover Image for Discovering themes in the data using metadata variables - advanced analysis with Skimle
分享这篇文章:

You have 300 survey responses, 60 interview transcripts, or a year's worth of customer feedback. You've run your thematic analysis and you can see the main themes clearly. But now the real questions start arriving: are younger customers saying something different from older ones? Did sentiment shift after the product update in November? Do public sector respondents have different priorities from private sector ones?

This is where many qualitative tools leave you stuck. While with some tools you can sort data by metadata variables (for example manually in Excel, or with ATLAS.ti's tools), you still need to read through everything again and try to spot the differences.

Skimle's metadata features solve this. They let you attach structured variables to your documents and then automatically surface where those variables explain meaningful differences in what people are saying. Think of it as pivot tables for qualitative data: the same analytical power, but applied to themes and language rather than numbers.

This article explains how the feature works, when to use it, and what it looks like in practice.


What metadata is in Skimle

In Skimle, metadata refers to structured descriptive information attached to each document in your project. A document might be an interview transcript, a survey response, a customer support ticket, or any other text. Metadata is everything you know about that document beyond its content: who produced it, when, in what context, from which segment.

Some examples:

  • For customer research: response date, product used, subscription tier, country, age group
  • For employee interviews: department, tenure, role level, office location
  • For academic fieldwork: interview date, participant gender, organisation type, industry sector
  • For document analysis: publication date, author organisation, document type, region

How do you get metadata values into your data?

There are four ways, and most projects end up using more than one.

Import data as .csv with metadata fields

The easiest way is to bring in data as .csv tables (these can be exported easily from Excel). When loading the data into Skimle, you define which columns to treat as metadata, which are content, and which to ignore.

For a customer feedback dataset, you could set the respondentID column as the title, NPS scores, products used and time periods as metadata, and the written feedback as content. This way Skimle runs the thematic analysis on the open text answers only, and the other variables become available to explain differences in the verbatims. See supported formats for the details of how CSV import works.

Automatically recognised from imported documents

If you are importing interview transcripts, reports or other documents, Skimle automatically identifies system metadata including creation date, authors and organisation, alongside the original filename and an editable document name. These fields cannot generally be deleted, but most can be edited.

Create custom metadata with AI

In the Metadata section, click Add new metadata fields. For each field you give a name (for example "Gender", "Age group", "Region"), a description explaining what the field represents, and a value type (text, number, date or URL). The description is what the AI uses to decide what to extract, so be specific: it is the difference between a field that fills reliably and one that returns noise.

Enable AI generation and Skimle searches each document for the relevant information and extracts a value. You can add several fields at once before saving. One thing that catches people out: to start generation you need to scroll to the bottom of the dialogue and click Generate metadata values. For larger datasets this takes a few minutes as Skimle works through each document.

Every generated value is editable afterwards, which matters. AI-inferred metadata should be reviewed like any other AI output, especially for fields where a wrong value would quietly distort a comparison.

Edit metadata directly in the metadata editor

The metadata editor is a spreadsheet-like table with documents as rows and metadata fields as columns, and it is the fastest route for anything the other three methods get wrong. Click any cell to type or change a value; changes save automatically. Click a field name in the header row to change its value type, which is worth doing for anything you want to treat as a number or a date, since those types unlock the quantitative and timeline analyses further down this article.

You can also import metadata from a spreadsheet at any point after upload. Skimle matches rows to documents by reference number or document name, previews the new fields, updated cells and deletions, and applies them on confirmation. Columns that do not yet exist in the project are created automatically.

The trick worth knowing: export your metadata to Excel, edit the values there (adding new columns if you need them), and import the file back. For bulk corrections across a large corpus this is far quicker than editing cell by cell. Full detail is in the adding metadata documentation.


Discovering patterns with metadata

Once your metadata is populated, Skimle gives you two distinct routes to what it reveals: automatic pattern detection inside your category summaries in the categories view, and the Visualisations section for interactive exploration.

Automatic pattern detection in the categories view

Run Analyse metadata from the categories view and Skimle examines every metadata field to determine which ones produce meaningful differences across your insight categories.

It works through two complementary methods. The first uses the semantic embeddings of your insights to check whether the content expressed by different metadata groups is genuinely different, rather than merely labelled differently. The second uses a statistical test to check whether insights from certain metadata groups cluster disproportionately in certain subcategories. A field is treated as significant only when both signals agree, which is what keeps spurious findings out.

Where a field clears the threshold, Skimle surfaces it in the category summary, ranked by significance, with a plain-language description of what characterises each group. Under an "Age group" field you might see that younger participants focused on usability while older participants prioritised reliability. If your documents carry dates, Skimle also compares recent against earlier responses automatically and describes any temporal shift it finds.

This appears directly in the category summary, alongside the main themes, so you see the "what" and the "who" in one place. Fields that show no meaningful difference are not shown at all.


Exploring metadata in Visualisations

Visualisations is where you interrogate the patterns yourself. It is organised into five lenses listed down the left-hand navigation, each holding several blocks.

LensWhat it answers
OverviewWhat stands out across the whole project
CategoriesHow the codebook behaves: sizes, splits, co-occurrence
MetadataHow categories and insights split across each field's values
DocumentsPer-document coverage and where insights fall
TimelineHow the material moves over time

The lens you are on is held in the page address, so you can bookmark or share a link straight to a particular view.

The Metadata lens

This is the one to reach for first when you have a segmentation question.

Metadata distribution lists your fields with their number of values and coverage. Pick a field to see which values are most common and how the documents spread across them, then break the field down by category to see the composition of each group.

Categories by metadata is the block that replaces what used to be a simple heatmap. It is a full-width grid of categories against the values of one metadata field, coloured by how frequent each pairing is. Scan it and uneven distribution jumps out: a cell that stands out tells you a particular group is over-represented in that category. Click any cell to open the insights behind it, which is the part that matters, because a pattern you cannot read the verbatims for is a number rather than a finding.

Comparison contrasts a group you pick against the rest of the corpus and ranks the categories where your selection stands out most. You choose each side by documents or by metadata values, and if you leave Group B unset it defaults to everything else. Three metrics are available: percentage-point difference (the gap in share between the groups), times as common (how many times more common a category is in one group), and odds ratio. Clicking a metadata value elsewhere in Visualisations seeds the comparison, so moving from "that looks odd" to a quantified contrast is one click.

The Timeline lens

Available when your documents carry a date field. A Dates from picker chooses which field places documents on the time axis, once for the whole lens.

Timeline trend shows how categories, subcategories or metadata values move over time, in five chart forms: lines, split, bars, matrix, and area. A sudden spike in a particular value, or a gradual shift in the balance between groups, usually corresponds to something real: a product change, a policy announcement, a seasonal shift in who is engaging with your service. Grouping all observations together hides exactly this.

Compare periods is the more decisive tool for a before-and-after question. Mark an earlier and a later range on a calendar strip and Skimle shows which categories grew and which shrank between them. Use compare periods when you are contrasting two stretches of time and comparison when you are contrasting two groups.

What the other lenses add

Overview opens with key statistics for the project (corpus, structure, metadata, density), insights by category, a word cloud you can switch between the words your documents use and the entities they name, a metadata distribution block, and a live comparison. It also folds away a Terminology reference explaining every control, which is worth reading once.

Categories holds category distribution, a co-occurrence heatmap showing which categories turn up in the same documents, and a scatter that plots every document or category as a dot so outliers are easy to spot.

Documents gives you a per-document ranking by how much each is coded, a document codeline showing where each category's quotes fall from the start of a document to the end, and a full grid of documents against categories.

Controls worth knowing

Most blocks share the same controls: Level (roots, categories or subcategories), Group by, View (what each bar represents), Measure (count by insights or by documents), Metric (counts or share of total), and Show, Sort and Colour. Categories keep the same colour everywhere and their subcategories get distinguishable shades of it, so related subcategories read as a family across every chart.

Three features apply across all five lenses. Filter narrows the whole of Visualisations to a subset of documents, categories or metadata values, and every block updates to match. Export copies any chart as an image, downloads it as a PNG, or downloads the underlying CSV, so figures go straight into a report. And selections carry between lenses, so clicking through from a document or category keeps your place. There is a per-block reference of which dimensions each one supports in the Visualisations settings documentation.


A business example: analysing customer feedback over time

Suppose you run a media streaming service offering three products (Music, Magazine, Video) and you collect ongoing subscriber feedback. You have 300 responses gathered across December 2025, January 2026 and February 2026.

After running your analysis on open text responses, you have categories like "App performance", "Content quality", "Pricing" and "Customer support". You have set up three metadata fields: product, response date, and NPS score.

What the automatic analysis surfaces:

In "App performance", Skimle flags the Product field as highly significant. Music and Video subscribers dominate the category and their feedback is semantically distinct from Magazine subscribers'. The temporal comparison detects that performance-related feedback from Music and Video subscribers peaks sharply in December 2025, then drops in January and February. You know why: there was a service outage over the holiday period.

In "Content quality", the temporal analysis surfaces something different. Magazine subscriber feedback has shifted in tone across the three months. Recent responses are still largely positive, but a small group of February respondents use language around editorial bias and changing editorial direction, distinct from the earlier, more uniformly positive feedback. Skimle flags this as a borderline-significant temporal shift and describes it in plain language.

What Visualisations adds:

In timeline trend, the December spike in "App performance" from Music and Video subscribers is unmistakable: nearly 60% of the category's insight volume in December against under 30% in the surrounding months. Export the chart as a PNG and it goes to the product team as evidence of the outage's qualitative impact.

In categories by metadata, cross-referencing NPS bracket against categories shows detractors concentrated in "App performance" and "Pricing" while promoters cluster in "Content quality" and "Ease of use". Click the detractor and "Pricing" cell and you are reading the verbatims behind it.

Then use comparison to quantify it: set Group A to detractors, leave Group B as the rest of the corpus, and read the result as times as common. "Pricing" being three times as common among detractors is a sharper line for a leadership deck than a heatmap cell.

The whole analysis takes less than an hour rather than days. And every insight in every chart traces back to the original text, so you can read the verbatim feedback behind any pattern. This is what rigorous AI-assisted analysis looks like in practice: a structured system you can interrogate at every level rather than a black box.


An academic example: interview data with demographic variables

Metadata analysis is equally useful in research settings. Suppose you have conducted 45 semi-structured interviews with managers from large and small organisations, across public and private sectors, about their experience of digital transformation.

You set up metadata fields for organisation size (large/medium/small), sector (public/private/NGO), and interview date, to track whether narratives shifted during your data collection period. After importing your audio transcripts and turning them into text with Skimle and running the analysis, your categories include "Leadership and vision", "Resource constraints", "Employee resistance" and "Technology choices".

The metadata analysis might reveal:

  • Organisation size is highly significant for "Resource constraints": small organisations' responses are semantically distinct from large organisations', with small organisations discussing procurement and budget cycles while large organisations discuss internal governance and approvals.
  • Sector is significant for "Technology choices": public sector respondents cluster around compliance and interoperability, private sector around speed to market and vendor relationships.
  • The temporal field shows that interviews conducted before a major regulatory announcement describe "Leadership and vision" differently from later ones, suggesting the announcement reshaped how managers frame transformation priorities.

Each is a substantive finding that would have taken days of manual comparative coding to identify. And you might have missed the patterns you were not expecting, since every comparison you run by hand costs effort and you only run the ones you thought of. Here they surface as part of the normal workflow, and once the AI has spotted a pattern you can dig into the data and collect more evidence bearing on it the very next day you speak to an informant. That is theoretical sampling with a shorter feedback loop, which is why this pairs well with grounded theory work.


When metadata analysis is most useful

Metadata analysis earns its place when:

  • You have a large enough dataset for group comparisons to be meaningful (Skimle needs at least a few insights per group to run the analysis)
  • Your research question involves comparing across segments: demographic groups, customer tiers, regions, time periods
  • You are running longitudinal or repeated studies and want to track change rather than describe the current state
  • You are doing due diligence or competitive research where documents come from different sources and you want to know whether source explains variation in what is being said

It is less useful for small, homogeneous datasets where the variables do not explain much variation. In those cases Skimle tells you that no significant fields were found rather than manufacturing a result, which is the behaviour you want.

One caution worth stating. A metadata field that reaches significance tells you a difference exists; it does not tell you why. Small groups are the usual trap: a difference across eight documents in one segment is a lead to investigate, not a finding to report. Click through to the insights and read them before you write anything down.


How it relates to the broader Skimle workflow

Metadata analysis sits downstream of your core thematic analysis. You do not need it to get value from Skimle, since the categories, summaries and document-level insights are useful on their own. But it adds a layer of analytical depth that is genuinely difficult to achieve at scale with qualitative data.

If you work across multiple languages, Skimle's multilingual analysis works alongside metadata: you can hold documents in several languages in one project and still run comparisons across them. For teams collecting data before analysis, Skimle Ask generates transcripts with participant metadata already attached, ready for this kind of structured analysis. If you are combining several feedback sources in one project, combining customer insights across feedback channels covers why a channel field is the most valuable metadata you can add. And for workflows where metadata-structured insights feed AI agent pipelines, see our guide on agentic chat and MCP.


Ready to discover patterns in your qualitative data? Try Skimle for free and see how metadata analysis gives you the power of pivot tables in Excel, applied to themes, language and meaning. Whether you are working with customer feedback, interview transcripts or document archives, you can stop re-reading your data and start understanding it.

Want to learn more about Skimle's analysis features? Read our guide on how Skimle's end-to-end workflow handles qualitative data and how to set up and export your projects.


About the authors

Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile

Olli Salo is a co-founder at Skimle and former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile

用 Skimle 深入挖掘您的数据

Skimle 会自动收集、分析并归类访谈、问卷回答、报告及其他定性数据。我们这款现代定性分析软件,把严谨透明的流程与 AI 的速度结合在一起。

上传文本或音频,用 Skimle Anonymise 去除敏感数据,自动生成类别与子类别,跨文档探索数据,并把数据导出为契合您工作流程的格式。由专业人士为专业人士打造,充分保护隐私并符合 GDPR。

免费试用 · 无需信用卡 · 完整方案每月 20 欧元起