Do symptom checkers actually work?

They are the most-used health software in the world, and they have been measured properly, more than once, over years. The results are worth knowing.

Symptom checkers are the most-used form of health software in the world. Someone types in what is wrong and gets back a list of possible causes and, usually, advice on how urgently to seek care.

They have also been studied properly, more than once, over years. The results are worth knowing before you rely on one at eleven at night.

The five year study

In 2022 a team published a follow-up in the Journal of Medical Internet Research, re-testing symptom checker apps against the same standardised patient cases they had used five years earlier. It is an unusually good design, because it measures whether the field is improving rather than just taking a snapshot.

The headline: "The median triage accuracy in 2020 (55.8%, IQR 15.1%) was close to that in 2015 (59.1%, IQR 15.5%)."

Triage accuracy here means getting the urgency right: emergency, see someone soon, or manage at home. Just over half, across the field.

Worse, it had not moved. "Triage performance of symptom checkers has, on average, not improved over the course of 5 years." And in two specific areas it went backwards: "It decreased in 2 use cases (advice on when emergency care is required and when no health care is needed for the moment)."

Those are the two questions people actually open these apps to answer.

The comparison that lands hardest

The authors also compared the apps against ordinary people with no medical training, making the same judgements on the same cases.

"Few apps outperformed laypersons in either deciding whether emergency care was required or whether self-care was sufficient." And then: "No apps outperformed the laypersons on both decisions."

Not one. For the central task these tools exist to perform, the software was not reliably better than a person guessing sensibly.

Diagnosis is a separate, weaker skill

A systematic review the same year separated the two things these tools do. Diagnostic accuracy, naming the condition, was low, with primary diagnosis correct in a range of roughly 19 to 37.9 percent. Triage accuracy was generally higher, ranging from 48.8 to 90.1 percent depending on the tool.

The review also makes a point that is easy to miss and important: correct diagnosis does not reliably translate into correct triage. An app can name the condition and still give you the wrong advice about what to do about it, and the reverse happens too.

Note also the width of those ranges. "Symptom checkers" is not one thing. The difference between the best and worst is larger than the difference between the average one and no tool at all.

How to use one sensibly

This does not make them worthless. It makes them a particular kind of tool with a known error profile.

  • Use it to generate questions, not conclusions. "Should I mention the calf pain?" is a good use. "It says it is probably nothing" is not.
  • Treat an urgent result as worth acting on, and a reassuring result as worth ignoring if you are still worried. The errors that matter are the falsely reassuring ones.
  • Never use one to decide against calling 911. The evidence on that specific judgement is the weakest in the whole literature.
  • Remember it does not know the person. It has no chart, no medication list, no history. It is reasoning about an average human being.

Why we take this seriously

We build a product where people type symptoms into a box, so this research is about our category, not somebody else's.

Two things in it shape how Dr.life is built. The first is that the tools studied were working without access to the person's record, which is a fundamentally harder problem than the one an integrated system faces. The second, and more important, is that a licensed clinician can enter the conversation. Software that must decide alone is being asked to do the thing this literature says it does worst.

We would rather point at these numbers than pretend a chat window has solved something that five years of measurement says is not solved.

Sources

  1. Triage Accuracy of Symptom Checker Apps: 5-Year Follow-up Evaluation. Schmieding ML, Kopka M, Schmidt K, Schulz-Niethammer S, Balzer F, Feufel MA. Journal of Medical Internet Research, 2022. Used for the median triage accuracy figures, the lack of improvement over five years, and the comparison with laypersons, all quoted directly.
  2. The diagnostic and triage accuracy of digital and online symptom checker tools: a systematic review. Used for the separate diagnostic and triage accuracy ranges and for the finding that correct diagnosis does not reliably produce correct triage.

Every source above was read before it was cited. Where the evidence is uncertain, this article says so rather than rounding it into advice.


More for families

Why seeing the same doctor matters

It feels like a preference, the way liking a particular barber is a preference. The research suggests otherwise, and the effect is larger than most people would guess.

When a parent refuses help

It looks like stubbornness. Understanding what is actually happening matters, because the usual family response makes it worse.

All articles


Dr.life is the app behind this

A healthcare AI with real doctors ready to join the conversation when needed, from Life Medical. Not live yet.