Essay
What finally changed the AI's mind

Three rounds of added rules: nothing.
Past decisions on similar records: the agent committed at once.
Full policy text on top: it pulled half of those back.
What we did
We took the same 283 UK government purchase records from the first experiment and put the same AI in front of them — five times.
Each round, we added one more layer of the governance context an organisation might give an AI agent. Round one was the bare records, with no rules at all. Round two added a one-paragraph description of how the buyer is supposed to operate. Round three named the actual rules — publication windows, sign-off limits, the usual checks. Round four added something different: past decisions on similar records, with the verdict each one received. Round five added the full policy text on top of everything else.
We wrote down what we expected to see before we ran any of it: as the AI got more context, it should get more decisive — fewer "needs review" verdicts, more straight answers. Every decision the AI made was bound into a signed record at the moment it made it, so anyone could replay the entire experiment later.
What we saw
For the first three rounds, almost nothing changed. The AI hedged on 97-100% of records — "needs review" on nearly everything, whatever rules we showed it. More context, in the form we expected to matter, did almost nothing.
Then round four — past decisions on similar records — and the AI committed in a single jump. It denied 107 records at once, more than a third of the corpus. Not gradually, not partially. One round of past decisions and the hedging broke.
In round five, with the full policy text added on top, the AI eased about half of those denials back to "needs review". The shape is not a steady climb from cautious to confident. It is a single sharp jump in one specific round, partially reversed in the next.
We then ran a sanity check. We inverted one of the rules in the policy text — changed "must publish within thirty days" to its mirror image — and gave that inverted version to the AI. We wanted to see whether the AI would notice the swap. It did not. It kept reasoning as if the original rule were still there, citing what the rule was supposed to mean rather than what the page now actually said. Its confidence did not budge.
Why it matters
For organisations building AI-assisted review into real workflows, this is two different findings landing at once.
The first is practical. If an AI agent is hedging too much, more rules will not fix it. What moves the AI is being shown real past decisions on similar cases — that is the lever that broke its hedging. Anyone designing an AI review workflow needs to be deliberate about which examples the AI sees, because those examples are doing more of the work than the rules themselves.
The second finding is harder. The inversion test showed that the AI's reasoning was anchored to what it thought the rule meant, not to what the page actually said. That is a failure mode you cannot see in normal operation, because the rule and the AI's understanding usually match. It only shows up when you put the AI in front of an inverted version of the rule and watch what happens.
A regulator, a board, or an auditor reviewing AI-augmented decisions needs both: the past examples the AI was shown at the moment of decision, and a way to verify the AI was reading what it was actually given — not what it expected to see.
What's next
This experiment opened three new questions that we could not settle in the same run. Was it the past decisions doing the work, or was it the combined weight of all the layers stacked together? Was the policy text's pullback driven by the rules themselves, or by the cautionary instruction we had included with them? And was the AI's failure to notice the inverted rule a quirk of one model, or something that holds across different AIs?
The third experiment in the series was designed to disambiguate each of those, with another model run alongside the original to test how much of this was the AI we chose. That paper is now live, and the next post covers what it found.
The full paper for this experiment is below — the five rounds, the inversion check, every signed record, and the verification flow.
More from the Journal
- Three regulators, three vocabularies, one unanswered questionAn EU policy process, UK public opinion, and a Singapore legal working group have arrived at the same requirement in…
- Same reasoning, different decisionsTwo AI models read the same UK procurement records and reasoned about them the same way. One committed to a verdict o…
- Why the agent wouldn't say noAn AI and a written rulebook reviewed the same UK government purchase records side by side. The AI saw the same probl…