Human vs. machine: Testing AI’s ability to synthesize and analyze research
Summary
Nielsen Norman Group (NNG) has conducted and continues to conduct extensive research testing various large language model (LLM) tools designed for research synthesis and analysis. Our goal was to determine whether these AI-powered tools could meaningfully accelerate the work of experienced UX researchers. Through rigorous testing across multiple models and specialized research tools, we’ve found that while a few tools provide modest speed improvements for experienced researchers, none come close to replacing human expertise in research synthesis and analysis. The core problem is that these tools consistently exhibit critical flaws: they hallucinate findings, fail to identify meaningful patterns in qualitative data, cannot adequately consider nuanced research questions, and produce only superficial, high-level summaries of participant behavior. What makes this particularly dangerous is that these AI-generated outputs often have the veneer of legitimate research results—they look professional and sound plausible. However, closer inspection reveals significant gaps, inaccuracies, and missed insights that would mislead stakeholders and result in poor design decisions. The appearance of competence masks fundamental limitations that make these tools unreliable for serious research work. While we’ve found several places in the research process that can benefit from LLM usage, analysis and synthesis consistently falls short. In this talk, I can share the specific research we’re doing and explain what actually works.
Key Insights
-
•
AI tools frequently produce insight-shaped outputs but often lack the rigor and accuracy of trained human researchers.
-
•
AI moderators cannot currently assess user behavior beyond spoken words, missing key usability observations like failed or inefficient tasks.
-
•
Contextual elements such as environmental interruptions are critical in research but are invisible to AI tools.
-
•
Synthetic users generated by AI tend to produce overly positive, unrealistic feedback that can mislead product teams.
-
•
AI excels at finding semantic connections and grouping codes in large, already coded qualitative datasets quickly.
-
•
Meta-analysis of large repositories using AI can uncover recurring user themes, like change aversion, much faster than manual methods.
-
•
Integrating AI with organizational systems to pull in diverse data sources improves context but requires expert setup and is not yet simple.
-
•
AI’s context window limitations cause it to forget earlier input, affecting the accuracy of multi-turn interactions.
-
•
Even trained researchers must use AI outputs cautiously, vetting insights to maintain research quality.
-
•
Effective user research depends on human synthesis, collaboration, and contextual understanding, areas where AI currently fails.
Notable Quotes
"AI can generate insights, but it does not do them as well as a moderately trained human researcher."
"There is a world of difference between what a participant says and what they actually do, and AI misses that completely."
"AI tells you what you want to hear, which is dangerous if you’re making product decisions based on synthetic feedback."
"Our job as researchers is not making reports or interviewing users; it’s providing actionable, correct insights."
"AI tools are incentivized to produce final deliverables, but that’s an output, not the essence of research."
"AI is pretty good at finding semantic patterns among codes after human researchers have done the initial coding."
"Nobody is going to be satisfied by insight-shaped answers or high-level summaries masquerading as breakthroughs."
"AI cannot notice body language, tone, or environmental context during a research session."
"Using AI to scan large archives of research is a game changer for meta-analyses, even if it’s imperfect."
"Well-set-up AI systems pulling data from multiple company sources will have more context, but it’s still limited compared to human understanding."
Or choose a question:
More Videos
"We need to figure out how to assess design projects and designers better because employers need to make choices, but design resists simple scoring."
Nathan Shedroff Hugh Dubberly Thomas J. McLeishHow Will Design be Taught When the Schools Shut Down?
May 8, 2026
"We replaced the exhausting time estimation game with data-driven models that give us more reliable results."
John Paul de Guzman10k Screens Later: How We Became a Data-Driven Design Organization
September 24, 2024
"It's almost impossible to do a survey without using some kind of tool nowadays."
Aleksandra Korczynska Caroline Jarrett Justyna ParmeeSurvey Tools
March 12, 2026
"Motivation and focus are especially challenging when working from afar."
Kristin SkinnerTheme 2: Introduction and Provocation
January 8, 2024
"Attendee cohorts meet daily in small groups with facilitators to share learning goals and deepen engagement."
Rachael Dietkus, LCSW Victor Udoewa Jennifer StricklandEverything You Need to Know about the Civic Design 2022 Call for Presentations
May 17, 2022
"Enterprises are places where one hand doesn’t know what the other hand is doing and sometimes despises it."
Uday GajendarThe Rise of Meta-Design: A Starter Playbook
May 19, 2022
"We approached feedback loops like a system equal parts structure and skill. They have to reinforce one another."
Vanessa VarinFeedback: The Other F-Word
September 10, 2025
"You don’t think of raising a child or teaching a student as solving a problem, it’s an ongoing process—and so is design today."
Hugh DubberlyProblems with Problems: Reconsidering the Frame of Designing as Problem-Solving
June 19, 2019
"Research is evolving. It’s not about ownership, it’s about impact."
Sean McKayCoexisting with non-researchers: Practical strategies for a democratized research future
March 11, 2025
Latest Books All books
Dig deeper with the Rosenbot
What strategies can researchers use to become catalysts for change using play?
What strategies help speed up the procurement and legal review of UX research platforms in organizations like LinkedIn?
Why is human storytelling stewardship still essential despite advances in AI and research democratization?