When healthcare algorithms get it wrong

The most important study about AI in medicine is not about a chatbot. It shows how this kind of failure really happens: quietly, at scale, with nobody intending it.

Software has been making decisions about American healthcare for far longer than the current conversation about AI suggests. Algorithms decide who gets flagged for extra care, who is prioritised, and which patients a health system looks at more closely.

The most important study in this field is not about a chatbot. It is about one of those older, duller systems, and it is worth knowing because it shows how this kind of failure actually happens: quietly, at scale, with nobody intending it.

The study

In 2019, Obermeyer and colleagues published an analysis in Science of a commercial algorithm used across US health systems to identify patients with complex needs who should receive additional support.

They found that Black patients assigned the same risk score as White patients were, in fact, considerably sicker. The consequence was not abstract: the researchers estimated the bias reduced the number of Black patients identified for extra care by more than half.

What makes the finding instructive is the cause. The algorithm did not use race. It predicted health care costs, and used predicted cost as a stand-in for health need. That sounds reasonable, and it is the sort of substitution a sensible engineer makes without a second thought.

But less money is spent on Black patients at the same level of need, for reasons rooted in access and history rather than health. So an algorithm trained to predict spending learned to predict who the system had historically spent money on, and concluded that patients it had underserved were healthier.

Reformulating the algorithm so that it no longer used cost as a proxy for need removed the bias.

Why this is the story worth understanding

Nobody wrote a racist rule. There was no malice and, as far as anyone can tell, no negligence in the ordinary sense. The system did exactly what it was built to do. The error was in the choice of what to predict, and it was invisible until somebody went looking.

That is the general shape of algorithmic harm in medicine. It rarely looks like a dramatic wrong answer. It looks like a defensible design decision that quietly encodes an existing inequity and then applies it consistently to millions of people, which is worse than a human doing it occasionally.

It is not one algorithm

An AHRQ evidence review on healthcare algorithms and racial and ethnic disparities documents the same pattern elsewhere. Race-based coefficients in the equations used to estimate kidney function were found to systematically overestimate kidney function in Black patients, with consequences for transplant eligibility and treatment decisions. Certain lung cancer screening criteria disproportionately excluded minority patients from preventive care. Risk prediction models across conditions from cardiovascular disease to intensive care showed differing accuracy between racial groups.

The review's recommendations are worth knowing because they describe what responsible practice looks like: developing race-free equations where race was being used as a crude proxy, auditing algorithms systematically for bias across demographic groups, involving patients and affected communities in how these tools are built and evaluated, and putting real governance and accountability around deployment.

What this means for AI in healthcare now

Modern language models are a different technology, but they inherit the same structural vulnerability, and arguably a worse version of it.

A risk score is at least auditable. You can inspect what it predicted, compare it against outcomes by group, and discover, as these researchers did, that it is wrong in a patterned way. A model that produces fluent prose is harder to audit, because there is no single number to compare, and its errors arrive wrapped in confident, reasonable-sounding language.

Three questions follow, and they are worth asking of any health system, insurer or product that uses this technology on you.

  • What is it actually predicting, and is that the same thing as the outcome anybody cares about? Cost is not need. Engagement is not health.
  • Has anyone checked whether it performs differently for different groups of patients? If nobody has looked, nobody knows.
  • Who is accountable when it is wrong, and is there a person in the loop who can override it?

Where we stand

This is one of the reasons Dr.life is built with a licensed clinician able to step into the conversation rather than as software that answers and closes the case. A clinician who knows the patient is the check on the machine, and the place accountability actually lands.

It is also why we would rather publish this study than not. Anyone selling healthcare AI should be able to tell you how their system fails, not only how it performs.

Sources

  1. Dissecting racial bias in an algorithm used to manage the health of populations. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Science, 2019, volume 366, pages 447 to 453. Used for the finding that Black patients at the same risk score were sicker, the estimate that this halved the number identified for extra care, the cost-as-proxy mechanism, and the effect of reformulating the algorithm.
  2. Impact of Healthcare Algorithms on Racial and Ethnic Disparities in Health and Healthcare. Agency for Healthcare Research and Quality evidence review. Used for the further examples including kidney function estimation and lung cancer screening, and for the recommended mitigations.

Every source above was read before it was cited. Where the evidence is uncertain, this article says so rather than rounding it into advice.


More for families

Why seeing the same doctor matters

It feels like a preference, the way liking a particular barber is a preference. The research suggests otherwise, and the effect is larger than most people would guess.

When a parent refuses help

It looks like stubbornness. Understanding what is actually happening matters, because the usual family response makes it worse.

All articles


Dr.life is the app behind this

A healthcare AI with real doctors ready to join the conversation when needed, from Life Medical. Not live yet.