AI symptom checker matches or outperforms clinicians in national-scale diagnostic study

Google's SymptomAI correctly identified diagnoses as often as board-certified doctors — and its results correlated with real physiological data from wearables

Clinicians have long known that a good history takes you most of the way to a diagnosis. Now a large-scale study suggests an AI can conduct that history almost as well as they can. Google Research’s SymptomAI, tested across nearly 14,000 participants, produced differential diagnoses that board-certified clinicians preferred over their own peers’ assessments in more than half of cases. That’s not a marginal result. It’s a finding that warrants serious attention from anyone working at the intersection of AI and clinical care.

What the study actually tested

The study, published in a research paper titled “SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment,” enrolled 13,917 consenting participants across the United States. Each participant described their symptoms to one of five experimental AI agents, all built on Google’s Gemini Flash 2.0. The agents conducted end-to-end symptom interviews and generated a differential diagnosis, a ranked list of plausible conditions. Two weeks later, participants were asked to report any diagnosis they received from an actual healthcare provider.

That follow-up step is what makes this study methodologically credible. Rather than relying on synthetic patient vignettes or curated case studies, which is how most AI diagnostic tools have been evaluated, the researchers measured performance against real-world clinical outcomes. A panel of three board-certified clinicians then reviewed conversation transcripts, generated their own differential diagnoses, and ranked both the AI’s output and each other’s in a blinded comparison.

How the AI agents differed from each other

Participants were randomly assigned to one of five study arms, each reflecting a different interviewing strategy:

  • Dynamic Live and Dynamic Final: agents with full flexibility to ask any follow-up question they judged relevant
  • Fixed Canonical and Flexible Canonical: agents that drew from standard history-taking question sets taught in medical school
  • Base condition: a fully user-driven interaction with no AI-initiated follow-up, reflecting how most people currently use AI chatbots for health queries

Every agent-driven approach significantly outperformed the base condition. The message is clear: when the AI actively asks questions rather than waiting for the user to volunteer information, diagnostic accuracy improves substantially. That’s consistent with what clinicians already know. A passive listener rarely gets the full picture.

Where the AI performed best

The study measured accuracy using a top-5 metric, whether the participant’s actual diagnosis appeared somewhere in the AI’s list of five candidate conditions. On this measure, SymptomAI was rated more accurate than the clinician-generated differentials by the clinical raters. And the advantage was largest in cases where the clinicians themselves felt least confident. That’s a meaningful pattern. It suggests the AI may be particularly useful in ambiguous presentations, exactly the cases where a second opinion matters most.

Wearable data added a layer of validation

The researchers went further by cross-referencing SymptomAI’s diagnostic output with biometric data from participants’ Fitbit devices in the 30 days before their symptom report. For participants whose conversations led to an infectious disease diagnosis, wearable biosignals showed physiological changes consistent with immune activation in the days approaching symptom onset. The AI didn’t have access to this data during the conversation. So the correlation acts as an independent check on diagnostic plausibility, and it held up.

What this means for clinical practice

To be clear, SymptomAI is a research prototype. The diagnoses it generated were for research purposes only and do not replace clinical assessment. But the implications are hard to ignore. Millions of people face real barriers to accessing a clinician, whether geographic, financial, or systemic. A conversational AI that can conduct a structured history and generate a clinically credible differential could meaningfully extend the reach of diagnostic support, particularly in underserved settings.

There’s also a research infrastructure angle. Clinical-quality diagnostic labels are expensive to produce at scale. If AI symptom checkers can reliably generate them, that opens the door to population-level analyses of physiological data that are currently not feasible. The biosignal correlation work here is an early proof of concept for that possibility.

Still, questions remain about safety, equity, and how different populations with varying health literacy interact with these systems. The next step isn’t deployment. It’s replication, stress-testing, and honest evaluation of failure modes. But as a demonstration of what conversational AI can do in diagnostic medicine, this study sets a high bar.