Skip to main content

Essay

Same reasoning, different decisions

Sam Carter
Sam Carter12 Jun 2026 · 4 min read

Same records.

Two AIs.

Same reading. Different calls.


What we did

The previous experiment left us with three questions we could not settle in the same run. Was it the past decisions doing the work, or was it everything stacked on top of everything else? Was the cautionary instruction inside the policy carrying the change, or was it the policy text itself? And was the AI's failure to notice an inverted rule a quirk of one model, or something that holds across different AIs?

We ran a separate, controlled experiment for each question. We isolated past decisions — gave the AI nothing but the precedents. We isolated the cautionary instruction — gave the AI the policy with the instruction removed. And we ran the inversion test again, this time at a hundred records instead of fourteen, and on a second AI from a different company alongside the original.

Every part of the experiment was written down before any of it ran — what we expected to see, what would count as a finding either way, and what we would call a result we could not interpret. Every decision the AIs made was bound into a signed record at the moment of decision.


What we saw

Past decisions on their own barely moved the AI. When we showed it the precedents and nothing else, it committed to a verdict on three and a half percent of records — far below what we had set as our line for "this lever is the one doing the work". The pattern in the previous experiment had not come from the precedents alone. It had come from everything stacked together: the framing, the rules, and the precedents on top, working as one.

The cautionary instruction was nearly irrelevant. When we removed it from the policy text, the AI's behaviour barely changed — sixty-one percent of the previous denials stayed as denials, compared with fifty-seven percent in the original. The policy text itself was doing almost all the work. The cautionary instruction had been a sentence we had attributed weight to that the data did not.

Then the cross-model run. We ran the same hundred records through a second AI from a different company. Both AIs were given the same inverted policy. Both AIs read the policy the same way — reasoning against what the rule was supposed to say rather than what the page now said, in 93% of records for one and 100% of records for the other. The reasoning surface was nearly identical between the two models.

What was not identical was their decisions. One AI committed to a verdict on twenty-three percent of records. The other committed on eighty percent. On the same records, with the same prompt, the same policy. The reasoning travelled between the two models. The decisions did not.


Why it matters

For organisations evaluating AI for regulated work, this changes how the evaluation has to be designed.

Reviewing AI outputs traditionally means looking at the verdicts. Did the AI approve the records that should have been approved? Did it deny the records that should have been denied? Treat the verdict surface as the thing being measured.

The cross-model finding says that approach is incomplete. Two AIs can reason about a policy in almost exactly the same way — including making the same mistake when the policy has been altered — and still commit to verdicts on radically different fractions of records. The verdict surface and the reasoning surface are not the same property. Agreement on one does not imply agreement on the other.

In practical terms: if an organisation deploys AI for decision support and decides later to switch to a different AI — for cost, latency, or vendor reasons — the verdict surface may shift even when the reasoning surface looks unchanged. The methodology has to test both, separately, and not assume one implies the other. The infrastructure for this kind of evaluation already exists: every decision binds the prompt, the policy, the reasoning, and the verdict into one signed record. The discipline is to use it on both axes.


What's next

This was the third experiment in a series, and the disambiguation work that the series was set up to do is now done. The next piece of writing is a methods note that captures the methodology substrate across the three experiments — the pre-registration discipline, the signed receipts, the inversion check, the way both verdict and reasoning surfaces were measured.

The experiment after that is operational. We are setting up to do receipt-anchored evaluation in a live workflow, with a partner whose decisions and policies are real. The substrate findings from this series are about a specific policy and a specific corpus; the operational experiment is about whether the methodology survives contact with a live decision environment, on different records and a different policy.

The full paper from this experiment is below — the three controlled splits, the cross-model run, every signed record, and the verification flow.


Decision Assurance

Two AIs, the same records, every decision signed and verifiable from outside the firm that ran them.

Inspect a receipt

More from the Journal