Generative AI has moved from a curiosity to a genuine part of the qualitative analysis workflow in a few years, and the methodological literature evaluating it is starting to catch up. This guide covers what current studies actually show about AI-assisted coding — not the marketing claims, the measured performance — and why the emerging consensus lands on human-in-the-loop workflows rather than full automation.
Where AI performs well
The clearest finding across recent studies is that large language models do well on deductive coding — applying a predefined, well-specified set of categories to text — often reaching reliability close to human coders on concrete categories (Xu et al., 2025; Kang et al., 2025; Prescott et al., 2023). This makes sense: deductive coding is closer to classification than interpretation, and classification against a clear definition is exactly the kind of task LLMs tend to handle well.
Where it doesn’t
Inductive coding — where the categories themselves need to emerge from the data — is a different story. One direct comparison found ChatGPT and Bard reached moderate consistency with human-identified themes on inductive coding, up to 71%, but dropped substantially on more granular, fine-detail coding, down to 36–47% agreement with human analysts (Prescott et al., 2023). Left to prompt a model without constraints, researchers also report a real risk of hallucinated or decontextualized themes — patterns the model reports that don’t actually trace back cleanly to the source text (Kang et al., 2025; Xu et al., 2025).
That gap is well studied enough now that there’s a name for what’s missing: current systems can expedite early-stage analysis but lack the interpretive depth human coders bring, and show measurably lower reliability when working autonomously (Prescott et al., 2023; Ayik et al., 2026). Notably, a 2026 comparative study evaluating ChatGPT, QInsights, ATLAS.ti AI, and MAXQDA AI Assist against human thematic analysis found this pattern held across multiple commercial AI tools, not just general-purpose chatbots (Ayik et al., 2026) — the gap isn’t a quirk of one product, it’s a property of where the technology currently is.
What actually mitigates the risk
The techniques that measurably reduce hallucination and improve grounding share a common idea: keep the model anchored to the actual source text rather than reasoning freely. Multi-agent architectures — where one agent proposes themes and another checks them back against the raw transcript — and interfaces that force every AI-suggested theme to stay visibly linked to its source passage both improve reliability over unconstrained prompting (Kang et al., 2025; Xu et al., 2025). The TAMA framework is one documented example: a human-in-the-loop, multi-agent system that generates, evaluates, and refines candidate themes against source transcripts before a human reviews them (Xu et al., 2025).
The human-in-the-loop consensus
Across every study reviewed here, the recommendation converges on the same shape: AI as a drafting and pattern-surfacing assistant, with a human researcher reviewing, correcting, and making the final interpretive call — not AI coding autonomously and a human rubber-stamping the output (Yan et al., 2023; Xu et al., 2025; Prescott et al., 2023). The consistent finding across this literature is that automation can’t substitute for a researcher’s active, transparent analytical role — it can only make that role faster to execute.
A practical starting point
Based on where the evidence currently is, a reasonable working approach is: let AI assist with deductive coding against a codebook you’ve already validated, use it to surface candidate themes in inductive work but treat every suggestion as a hypothesis to check against the transcript yourself, and never adopt an AI-suggested theme or code without being able to point to the specific passage that justifies it.
AI assistance in Byleron QDA: AI-suggested codes and themes stay linked to the exact passage they came from, so reviewing a suggestion means checking it against the real quote right there — never accepting a pattern you can’t trace back to the data yourself.
References
- Ayik, B., Gu, D., Zan, Y.-X., Kim, S., & Kim, W. L. (2026). Human vs. AI: Evaluating Thematic Analysis With ChatGPT, QInsights, ATLAS.ti AI, and MAXQDA AI Assist. Qualitative Inquiry. https://doi.org/10.1177/10778004251412874
- Kang, D., Han, Z., Tian, J., Zhang, M., & Rzeszotarski, J. M. (2025). ThemeViz: Understanding the Effect of Human-AI Collaboration in Theme Development with an LLM-enhanced Interactive Visual System. Proceedings of the ACM on Human-Computer Interaction, 9, 1–29. https://doi.org/10.1145/3757675
- Prescott, M. R., Yeager, S., Ham, L., Rivera Saldana, C. D., Serrano, V., Narez, J., Paltin, D., Delgado, J., Moore, D. J., & Montoya, J. (2023). Comparing the Efficacy and Efficiency of Human and Generative AI: Qualitative Thematic Analyses. JMIR AI, 3. https://doi.org/10.2196/54482
- Xu, H.-R., Yi, S., Lim, T., Xu, J., Well, A., Mery, C., Zhang, A., Zhang, Y.-J., Ji, H., Pingali, K., Leng, Y., & Ding, Y. (2025). TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews. ACM Transactions on Computing for Healthcare. https://doi.org/10.1145/3828752
- Yan, L., Echeverría, V., Fernández-Nieto, G., Jin, Y.-Q., Swiecki, Z., Zhao, L., Gašević, D., & Martínez-Maldonado, R. (2023). Human-AI Collaboration in Thematic Analysis using ChatGPT: A User Study and Design Recommendations. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3613905.3650732

