Why Not Just Use a Bigger LLM

Executive Technical

Pearl proves a language model cannot compute a causal effect. Companies keep building with language models anyway. Both things are true, and the second one has real reasons behind it.

Judea Pearl's ladder has three rungs. Rung 1 is association: P(Y|X). This is correlation, and it is where LLMs and standard machine learning live. Rung 2 is intervention: P(Y|do(X)). This measures the effect of an action. Rung 3 is counterfactual: P(Yx|x′, y′). This asks what would have happened under different conditions.

Pearl proves that Rung 1 data cannot answer a Rung 2 or Rung 3 question. You need an explicit structural causal model: a graph, stated by someone who knows the domain, encoding which variable causes which. No amount of observational data substitutes for that graph. This is not a matter of opinion. It is a theorem.

Executives hear the theorem and push back anyway. Not because the math is wrong. Because building the graph looks slow and expensive next to an API call, and the objections that follow are practical ones, not mathematical ones. They deserve real answers.

The Domain Knowledge Bottleneck

A causal model needs a human expert to draw the graph: pricing, marketing spend, macro conditions, competitor behavior, all mapped by hand. Executives see this as slow and expensive. Feeding data into a deep learning model looks cheaper by comparison, because no one has to sit in a room and state what causes what.

The bottleneck is real. It is also smaller than it looks. Your experts already carry this graph in their heads: it is how they explain last quarter's numbers to the board. The work is elicitation, not invention, and it costs one structured conversation, not a research program.

Causal Discovery Doesn't Scale Cleanly

Some ask if an algorithm can learn the graph on its own. Automated discovery methods exist, the PC algorithm and Fast Causal Inference among them. They assume no hidden confounders. Real enterprise data is noisy and sparse, and that assumption breaks fast. Data science teams try these methods, then quietly drop them.

This is a reason to elicit the graph from an expert, not a reason to skip having one. The failure of automated discovery is an argument for the domain model, not against it.

Good-Enough Heuristics Win Over Proofs

Teams run an A/B test instead of building a formal model. A randomized trial gives a real Rung 2 answer, without do-calculus, and that is a legitimate substitute where you can run one. Most enterprise decisions do not offer that luxury. You cannot randomize last year's pricing decision after the fact, and you cannot A/B test a merger.

Where a trial is possible, run the trial. A causal model earns its cost on the decisions a trial cannot reach.

Liability and Executive Ownership

A causal model forces someone to write down an assumption: price increases do not affect quality ratings, for example. If that assumption turns out wrong, the model is evidence of what was believed and when. Black-box machine learning spreads the blame instead. Nobody signed anything, so nobody owns the error.

That is not a defect of causal modeling. It is the point. A model that names its assumptions can be checked, corrected, and improved. A black box that hides them cannot, and the absence of a paper trail is not the same thing as the absence of a mistake.

Set the objections aside. Even where none of them apply, companies still reach for an LLM first. Three reasons hold up.

Most enterprise data is unstructured: email, transcripts, PDFs, support tickets. A Bayesian network wants clean, structured input and an explicit graph. An LLM eats the mess with no schema required.

The Causal Cheat Code Pearl noticed this himself after GPT-4 came out. Human text already contains causal explanations. An LLM does not calculate P(Y|do(X)). It retrieves human reasoning about interventions from what it read. For most decision support, that recitation is good enough, and it is why an LLM sounds causally fluent without doing any causal computation at all.

A Bayesian network built for telecom churn cannot answer a supply chain question. An LLM handles sentiment, summarizing, code, and basic logic without being rebuilt. One is a specialist. The other is a generalist, and companies buy generalists first.

Bayesian Networks vs. LLMs where each one sits on the ladder
FeatureBayesian Network (causal)LLM
MechanismDirected graph, conditional independenceStatistical pattern matching over text
Rung reachedRung 2–3, with an explicit graph and do-calculusRung 1, though it recites Rung 2–3 language fluently
Data requiredStructured, with the graph designed by handUnstructured text and code at web scale
ExplainabilityEvery node and table can be inspectedOpaque, billions of parameters
Effort to stand upHigh: expert elicitation, upfrontLow: an API call
Read this table both ways. The BN column is not free. The LLM column does not reach Rung 2 or 3, no matter how it reads.

The LLM is the ear and the mouth. The model is the brains. A language model can speak fluently about any domain. It cannot know one. The .bayes file is the knowledge the LLM is missing: a causal map of the domain, auditable, versioned, and wrong in specific correctable ways.

Companies agree Pearl is right. An LLM cannot compute a true counterfactual for a scenario it has not seen, not without a structural causal model behind it. They choose speed anyway, and the choice does not have to be all or nothing.

The Working Pattern LLMs for broad, unstructured tasks and as the interface layer. Causal models reserved for narrow, high-stakes decisions: dynamic pricing, clinical trials, credit underwriting. Places where mistaking correlation for causation costs real money. This is the same split named on Grammar, Mood, and the Shape of Cognition: cognition = language_interface ∘ causal_model, not language times causal reasoning. The LLM does not get smarter about your business by getting bigger. It gets more fluent. Only the domain model gets smarter, because only the domain model is built from what your experts know.

This is also the answer to the liability objection from §02. The split makes ownership legible: the LLM's job is translation, and it can be wrong about phrasing. The causal model's job is the actual claim about the business, and someone signed it. You always know which layer to blame.

The objections above do not stay in a strategy memo. They show up mid-meeting, from a sharp stakeholder testing whether this is worth the budget.

Can't the algorithm just learn the causal graph automatically from our data?
AnyChatRung 1 · observed limitation
Automated discovery algorithms exist, but they assume no hidden confounders, an assumption your data almost certainly breaks. Run one on noisy enterprise tables and it will hand back a graph that looks precise and is not reliable. The faster, cheaper path is asking the people who already run this business what causes what. That conversation takes a day. The algorithm's false confidence takes longer to catch.
We already run A/B tests. Isn't that enough?
AnyChatRung 2 · do(X)
For the decisions you can test, yes, an A/B test is a real Rung 2 answer and you don't need a causal model to get it. Keep running it. A causal model earns its cost on the decisions you cannot randomize: last year's pricing call, a merger, a regulatory change. There, a trial is not on the table, and a graph is the only path to an intervention estimate.
If the LLM is only Rung 1, why does it sound like it understands cause and effect?
AnyChatProvenance
It read human writing that already explains cause and effect, and it is repeating that explanation back. It is not running do-calculus against a graph of your business. Ask it for the source and it cannot point to one, because there is no model underneath the answer, only text that resembled an answer during training.

The audit trail is the .bayes file, not the transcript of a fluent-sounding answer.

This page is the objection-handling companion to The Watson Lesson, which shows what happens when Rung 1 pattern matching is deployed as if it were causal reasoning. This page explains why companies do that anyway, and where the line needs to hold. It also names the same composition Pearl's split forces: cognition = language_interface ∘ causal_model, argued in full on Grammar, Mood, and the Shape of Cognition.

Upstream: nothing, this is the case made before a domain model gets built at all, the argument for why the elicitation conversation in the engagement structure is worth having. Downstream: once the objections are answered, the work becomes the domain model itself, elicited from your experts and encoded as an auditable .bayes file, not discovered by an algorithm and not left to an LLM's recitation of what causality usually sounds like.