The Watson Lesson

For an Executive Reader

What a $1 billion AI failure teaches us about the limits of pattern matching, and why causal structure cannot be replaced by scale.

In 2011, IBM's Watson defeated the two greatest Jeopardy! champions in history. The demo was real and remarkable: a system that could parse natural language questions, search a vast knowledge base, and surface confident answers in milliseconds. IBM called it cognitive computing. The press called it the dawn of thinking machines.

IBM then made a $1 billion bet that Watson could transform healthcare. It would read every oncology paper ever published, absorb clinical guidelines, and recommend cancer treatments better than any human physician.

Watson was not magic. It was a sophisticated pipeline: natural language processing to parse the question, information retrieval to find candidate answers, and a probabilistic scoring layer to rank them. Impressive engineering. But at its core, it was a pattern-matching system: finding associations between words, documents, and outcomes.

In Pearl's framework, Watson lived entirely on Rung 1. It could answer what tends to co-occur. It could not answer what would happen if we intervened. It had no model of disease, physiology, or causation. It had correlations at scale.

Rung 1: Association Watson's ceiling. Given this patient's profile, what do the records say tends to happen? A retrieval answer, not a causal one.
Rung 2: Intervention The question Watson could not answer. If we administer this treatment, what will happen to this patient? Requires a model of mechanism, not a retrieval engine over text.
Rung 3: Counterfactual Further still. Given that this patient received treatment A and deteriorated, would treatment B have changed the outcome? Requires a structural causal model. Watson had none.

MD Anderson Cancer Center ended their Watson partnership after spending approximately $62 million. Memorial Sloan Kettering followed. Physicians reviewing Watson's treatment recommendations found a significant fraction they rated as unsafe or incorrect.

The diagnosis was structural, not cosmetic. Watson had been trained on synthetic cases constructed by a single institution, not on real patient outcomes across diverse populations. More fundamentally, it had no causal model of how treatments affect disease progression. It was matching clinical language to guideline text. When the language was ambiguous or the case was atypical, the pattern matcher had nothing to fall back on.

2011

Watson defeats Jennings and Rutter on Jeopardy! IBM announces the cognitive computing era.

2013

MD Anderson and Memorial Sloan Kettering sign high-profile oncology partnerships.

2017

MD Anderson ends partnership after ~$62M spent. Physicians rate a significant share of recommendations unsafe or incorrect.

2018

Memorial Sloan Kettering ends partnership. Internal documents surface showing Watson was trained on synthetic, not real, patient cases.

2022

IBM sells Watson Health to Francisco Partners. The cognitive computing vision is abandoned.

By 2021, IBM was quietly unwinding Watson Health. In 2022 the division was sold to Francisco Partners. The Watson brand persists on a handful of enterprise tools, a chatbot builder, a search product, but the vision of a reasoning machine had been abandoned.

What remained was the infrastructure, stripped of the claim. Watson Assistant is a workflow chatbot. Watson Discovery is a document search layer. Neither is marketed as reasoning.

Watson was marketed in Rung 3 language: reasoning, understanding, insight, while running Rung 1 machinery. The gap between the claim and the architecture is where the damage happened: in patient safety, in wasted institutional investment, and in a decade of misplaced confidence that scale and NLP were sufficient substitutes for causal structure.

The question Watson could not answer was simple: if we give this patient this treatment, what will happen? That is an interventional question: Rung 2 at minimum. Answering it requires a model of the mechanism, not a retrieval engine over medical text.

This is not a critique of IBM's engineers. It is a critique of deploying associational systems in domains where decisions have causal consequences: and where the cost of confusing correlation with causation is measured in lives.

The knowledge Watson was missing wasn't in the data. It was in the causal structure of disease.