Code and affinity-map research evidence
Use AI for a traceable first pass across an approved qualitative corpus, then have researchers verify, reorganise, and name the themes.
Why
The workflow problem
Large qualitative corpora make first-pass coding slow and inconsistent, while an untraceable AI summary can hide contradictory or minority evidence.
AI can retrieve repeated language and propose provisional clusters across a closed corpus, reducing clerical load. It cannot determine what participants meant or what matters; the trade-off is speed against the risk of plausible but selective themes.
Concrete output
What you will produce
An evidence-linked affinity map with a codebook, theme definitions, outliers, source references, and a record of unresolved interpretations.
How
Run the workflow
Prepare an approved, traceable corpus
The researcher removes or protects personal data, assigns stable source IDs to transcripts and notes, and writes a short analytic question plus an initial human code seed. Check that every item is in scope and can be reopened by the review team.
Generate provisional codes and clusters
Give the approved corpus, analytic question, and source-ID convention to the AI. Ask for candidate codes and clusters only when each claim links to exact excerpts; treat the result as a draft working surface.
Propose a cited first-pass codebook
Analyse only the supplied research corpus. Analytic question: [insert] Source-ID convention: [insert] Initial human code seed: [insert, if any] Return a table: candidate code | definition | supporting source IDs and short verbatim excerpts | contradictory or outlier excerpts | uncertainty. Then propose rough clusters. Do not infer facts beyond the corpus, invent quotes, rank themes, or replace source IDs with participant labels.Rework the affinity map against raw evidence
Researchers and relevant observers inspect each proposed cluster against its source excerpts, move or split notes, preserve disconfirming material, and rename themes in their own words. Check that the map contains observations before interpretations.
Record findings and next questions
The research lead writes concise findings, links each to the verified map, and records gaps, disagreements, and follow-up questions. Approve only findings that a reviewer can trace back to the source corpus.
Human–AI partnership
Who contributes what
AI contribution
Provisional text coding, clustering, and retrieval across the supplied corpus. Its outputs are hypotheses and pointers, not findings.
Human responsibility
Set the analytic question and data boundary; protect participant data; retain familiarity with the corpus; interpret meaning; resolve disagreement; preserve exceptions; and approve findings.
Stop rule
Stop AI use and return to the corpus when a code, quote, or cluster lacks a verifiable source link; when an Other cluster absorbs material; or when lexical similarity is being treated as conceptual analysis.
Check before use
Watch out
A fluent cluster label can overstate consensus, and de-identification may not make sensitive research safe for every tool. Sample clusters across participants and contexts, inspect negative cases, and confirm that access, retention, and consent allow the chosen system.
Evidence base
What supports this workflow
GOV.UK describes affinity sorting as a collaborative process grounded in observed behaviour and verbatim quotes. NN/g says AI can suggest codes for text data but may miss, misinterpret, or manufacture insights. Harvested practice accounts support a traceable first pass, not autonomous analysis; claimed time savings are not independently generalisable.
- GOV.UK · Analyse a research session
Defines evidence-grounded affinity sorting, including verbatim observations, collaborative grouping, and human interpretation into findings.
- Maddie Brown and Kate Moran · Accelerating Research with AI
Supports AI as a starting point for text coding while explicitly warning that researchers must review for missed, misinterpreted, or manufactured insights.
- Rodion Sorokin · Teaming Up with AI
Harvested named-client practice account that documents AI-supported transcript synthesis and the need to check every cited quote against the source.