If two AI models read a policy the same way, will they decide the same way?
We broke the previous experiment apart to find which piece was doing the work, then ran the same records through a second AI from a different company. The two models read the policy the same way and committed to verdicts on very different fractions of records.
Sam Carter12 Jun 20264 min readUpdated 10 Sept 2026
From the piece
“The same records, two AI models from different companies, the same reading of the policy — and very different willingness to reach a verdict.”
An organisation puts a set of records in front of an AI model for review and uses what it returns as decision support. Some time later the model is swapped for another one, for cost, latency or vendor reasons.
The reasonable-sounding check is whether the new model reads the policy the same way as the old one.
We ran that cross-model AI evaluation on public UK procurement records, and the check turned out to be incomplete. Two models read the policy almost identically and still committed to verdicts on very different fractions of the same records.
Three questions the previous experiment could not settle
The previous experiment left three questions we could not answer in the same run. Was it the past decisions doing the work, or everything stacked on top of everything else? Was the cautionary instruction inside the policy carrying the change, or the policy text itself? And was the AI’s failure to notice an inverted rule a quirk of one model, or something that holds across different AIs?
We ran a separate, controlled experiment for each. We isolated past decisions, giving the AI nothing but the precedents. We isolated the cautionary instruction, giving the AI the policy with the instruction removed. And we ran the inversion test again, this time at a hundred records instead of fourteen, and on a second AI from a different company alongside the original.
Every part of the experiment was written down before any of it ran: what we expected to see, what would count as a finding either way, and what we would call a result we could not interpret. Every decision the AIs made was bound into a signed record at the moment of decision.
Past decisions on their own barely moved the AI
When we showed the AI the precedents and nothing else, it committed to a verdict on three and a half percent of records — far below the line we had set in advance for “this lever is the one doing the work”.
So the pattern in the previous experiment had not come from the precedents alone. It came from everything stacked together: the framing, the rules, and the precedents on top, working as one.
The cautionary instruction was doing almost no work
When we removed the cautionary instruction from the policy text, the AI’s behaviour barely changed. Sixty-one percent of the previous denials stayed as denials, compared with fifty-seven percent in the original.
The policy text itself was doing almost all the work. The cautionary instruction had been a sentence we attributed weight to that the data did not.
The same reading of the policy, very different verdicts
Then the cross-model run. We put the same hundred records through a second AI from a different company, and gave both AIs the same inverted policy.
Both read the policy the same way, reasoning against what the rule was supposed to say rather than what the page now said — in 93% of records for one and 100% of records for the other. The reasoning surface was nearly identical between the two models.
Their decisions were not. One AI committed to a verdict on twenty-three percent of records; the other committed on eighty percent. Same records, same prompt, same policy. The reasoning travelled between the two models. The decisions did not.
What this does not establish
This is one corpus and one policy: public UK procurement records, read under a single policy. The paper states its own limitations — among them that the cross-model result is not a verdict on either model, and that the substrate findings belong to this corpus rather than to AI review in general.
It is a research preview and it has not been peer-reviewed. It does not establish how every model will behave in every workflow.
Both surfaces have to be measured, separately
For an organisation evaluating AI for regulated work, this changes how the evaluation is designed.
Reviewing AI outputs traditionally means looking at the verdicts: did the AI approve what should have been approved, and deny what should have been denied? The verdict surface is treated as the thing being measured.
The cross-model finding says that approach is incomplete. Two AIs can reason about a policy in almost exactly the same way — including making the same mistake when the policy has been altered — and still commit to verdicts on radically different fractions of records. Agreement on the reasoning surface does not imply agreement on the verdict surface.
In practical terms: if a model is swapped for another one, the verdict surface may shift even where the reasoning surface looks unchanged. The methodology has to test both, separately, and not assume one implies the other. The infrastructure for that already exists in the experiment itself — every decision binds the prompt, the policy, the reasoning and the verdict into one signed record. The discipline is to use it on both axes.
What comes next
This was the third experiment in a series, and the disambiguation work the series was set up to do is now done. The next piece of writing is a methods note capturing the methodology substrate across the three experiments: the pre-registration discipline, the decision receipts, the inversion check, and the way both verdict and reasoning surfaces were measured.
The experiment after that is operational. We are setting up receipt-anchored evaluation in a live workflow, with a partner whose decisions and policies are real. The substrate findings from this series are about a specific policy and a specific corpus; the operational experiment asks whether the methodology survives contact with a live decision environment, on different records and a different policy.
The full paper is below — the three controlled splits, the cross-model run, every signed record, and the verification flow. A worked example of one of those records is on the receipts page.
Read next
The nearest pieces to this one: same kind first, then closest in time.
For search engines and answer enginesWhat this page carries
Structured data
BlogPosting + BreadcrumbList, with an Organization publisher and a Person author. No FAQPage: this piece has no FAQs, and none are padded in to reach a count.