Skip to main content

Article · AI agent hedging

Why wouldn't the agent say no?

An AI and a written rulebook reviewed the same UK government purchase records at the same moment. The AI named the same problems the rulebook did and almost never put a verdict on them. Changing the wording of one rule changed how often they agreed.

Sam Carter21 May 20263 min readUpdated 10 Sept 2026
A mechanical arm holding a red bar across a junction in a bright hall

From the piece

“The experiment was meant to see where the AI got things wrong. The more interesting answer was where it refused to make a call.”

Reads in full at ‘How we ran the comparison’

A public buyer publishes a purchase record — the kind every public buyer now has to publish under the new UK procurement law.

Two reviewers read the same record at the same moment. One is an AI asked to approve it, flag it for a human, or deny it. The other is MeshQu, checking that record against a written rulebook.

The question the experiment set out to answer was where the AI would get things wrong. The answer it produced was about something else: AI agent hedging, where the AI would not commit at all.

How we ran the comparison

The AI was asked to read each record and pick one of three outcomes: approve, flag for a human, or deny.

MeshQu checked the same records against the rulebook — when each had to be published, who was allowed to sign it off, and what evidence had to be left behind.

Both reviews were bound together into a single signed record for each decision, so a regulator, a board or an auditor could replay the pair later. One decision, two reviewers, one verifiable file.

This is a research preview. It is a good-faith application of a public regulatory framework to public data, published with its records and its limitations, and it has not been peer-reviewed. It does not establish how every model behaves in every workflow.

The agent almost never said no

MeshQu approved about half the records and denied about half. The AI approved a handful, flagged almost everything else for a human to look at, and issued no denial at all across the whole set.

Read side by side, the pattern is clearer than the headline gap suggests. The AI was not missing the issues; it named the same ones the rulebook did — published too late, signed off above the buyer’s limit, evidence missing. It simply would not commit. Where the rulebook reached a verdict, the AI reached for ”more context needed” every time.

So the headline disagreement was large, while the real disagreement — what each reviewer actually believed about the record in front of it — was small.

Changing one rule’s wording changed the agreement

A rulebook that treats missing evidence as a confirmed problem, and a cautious AI that defaults to ”flag for a human”, will look like they disagree on every borderline record even when they are seeing the same thing.

So we changed one rule. Instead of ”this absence is a violation”, it said ”this absence needs a human to read”. Nothing about the AI changed. The two reviewers then agreed on about ten times as many decisions as before.

The rulebook had started describing the situation more accurately, and the apparent disagreement collapsed. That is the finding the headline number hid: the harder problem is not whether the AI is capable, but how the rulebook is worded.

Demoting PROC-005-OPEN-TENDER from a critical-by-default DENY to a needs-more-context REVIEW band raises agreement from 7 of 283 (2.5%) to 82 of 283 (29%).

≈11× under counterfactual
Not trivial overlap. The agent's reasoning on those records names the same evidence gaps the rule fires on — open-procedure flag absent, direct-award justification not present — so the demoted band reflects what both systems had already observed. The paper reads it as a finding about policy authoring, not agent capability: the agent's REVIEW class encodes information a binary policy projects away.

Source: MRP-2026-02, §5.2, Table 2 — counterfactual verdict distributions

What the next experiment measures

This was the first experiment. The next one gives the AI the rulebook and the precedent it is supposed to be reading against, and measures how its caution moves once it is not reasoning from training data alone. That experiment is already in build; the records will publish alongside the paper, and each one can be verified independently.

The practical standard behind both is the same. When a regulator, a board or a customer asks how one AI-augmented decision was made, the answer cannot be a screenshot, a vendor brochure or a re-run six months later under a different model. It has to be a record made at the moment of the decision, signed in place, and checkable from outside the firm that made it.

If a governed agent sits in one of your workflows, start by writing down which rule it is being asked to apply and what a review verdict from it should mean — then keep the decision receipts for each decision the pair produces.

The full paper is below: method, results, the one-rule change, the verification flow, and every record we ran.

Read next

The nearest pieces to this one: same kind first, then closest in time.