Skip to main content

Research Preview · MRP-2026-03 · v1.1 · August 2026

When precedents commit AI and policy pulls it back

We showed one AI agent the same 283 UK procurement decisions five times, adding more of our governance rules each round — from nothing, up to the full policy.

Why giving an AI agent more governance context doesn’t make it steadily more decisive

Sam Carter27 May 202638 pagesPublic DistributionRead the PDF
Stepped tide-line panels behind a held glass vessel, the levels marked in gradient

The agent first commits to verdicts at scale at L3, not at L4.

107 DENYs at L3

L3 is the first rung carrying precedent receipts, and it moves 37.8% of records off the REVIEW spine in a single step.

Source: §3 The L3 break, and §2 Table 1 (P1 outcome row)

Adding the full policy text at L4 pulled commitment back rather than extending it: 46 of L3's 107 fresh DENYs reverted to REVIEW, concentrated on one ambiguous rule class where commitment fell from 29 of 40 records to 1 of 40.

46 of 107 reverted
Source: §4 The L3→L4 backoff, and §9 Synthesis
Reference
MRP-2026-03 · v1.1 · August 2026
Extent
38 pages · 1.3 MB
Published
27 May 2026
Classification
Public Distribution

Cite as

Carter, S. (2026). When precedents commit AI and policy pulls it back: a five-rung governance-context ladder on 283 procurement decisions, signed and verifiable. MeshQu Research Preview MRP-2026-03. DOI — not yet assigned.

Version of record

The paper is the version of record. This page frames it; it does not replace it.

Read the PDF · 38 pp · 1.3 MB
SHA-256 of the served file
441655517b3285384d4bc7c6fbd1ed5db57caadc222d26bcd0953ce7394f8e85
Master
Cleared master, byte for byte identical to the published file. Never a re-export.
Published at
/research/when-precedents-commit-ai-and-policy-pulls-it-back.pdf

Abstract

The paper's own abstract, verbatim — not a rewrite.

Regulated teams deploying AI agents cannot say what makes one commit to a decision rather than defer it, or whether more governance context reliably helps. We reused the frozen 283-record UK procurement corpus from MRP-2026-02 and ran the same agent over the same records five times, adding context one rung at a time: baseline, prose, named rules, precedent receipts, then full policy. Predictions and ladder content were locked before any evaluation call; every verdict was bound into an Ed25519-signed receipt anchored to Sigstore Rekor. The agent held at 97.5–100% REVIEW for three rungs, then committed 107 DENYs (37.8%) when precedent receipts appeared, and backed 46 off again under full policy. Commitment is non-monotonic — the precedent rung, not the policy rung, is where it breaks.

Key figures

Every figure names the section it comes from, and every one is re-derivable from the published corpus.

107 DENYs at L3

The agent first commits to verdicts at scale at L3, not at L4.

REVIEW runs 97.5% at L0, 100.0% at L1, 100.0% at L2, then 61.1% at L3 and back up to 74.2% at L4. L3 is the first rung carrying precedent receipts, and it moves 37.8% of records off the REVIEW spine in a single step.

§3 The L3 break, and §2 Table 1 (P1 outcome row)

46 of 107 reverted

Adding the full policy text at L4 pulled commitment back rather than extending it: 46 of L3's 107 fresh DENYs reverted to REVIEW, concentrated on one ambiguous rule class where commitment fell from 29 of 40 records to 1 of 40.

§4 The L3→L4 backoff, and §9 Synthesis

Two predictions falsified

The two pre-registered claims describing the ladder's expected shape — REVIEW falling monotonically and agreement rising monotonically — both broke, in the inverted direction.

Agreement ran 2.5% at L0, 0.0% at L1, 0.0% at L2, 38.5% at L3, then dropped 12.7 points to 25.8% at L4. The paper reports the break as positive evidence that the method worked: a direction was specified with a falsification band before the data existed, the data broke it, and the writeup names the direction it broke in.

§2 Table 1 (P1, P2) and §1.3

13 of 14

On the permuted-policy diagnostic the agent returned the same verdict as it had under the un-inverted policy for 13 of 14 records, reasoning against what the rule means rather than what it now literally says.

The paper is explicit that this is not the obvious sycophancy reading. The agent is not agreeing with the inverted policy; it is ignoring it.

§5 Inversion-blindness: the Permuted-Policy diagnostic, and Abstract

53.3% (73/137)

Three predictions held.

DENY commitment on MeshQu's DENY records reached 53.3% at L4 against a pre-registered floor of 30%; L4 agreement stayed at 25.8%, below the counterfactual ceiling carried over from E1; and 94.5% of the records that shifted from REVIEW at L0 to DENY at L4 involved the same ambiguous rule, against a 60% floor.

§2 Table 1 (P3, P4, P6)

11.3% vs ≥50%

The agent cited explicit rule codes on 11.3% of records at L4 against a predicted 50%, falsifying that prediction.

Partly a measurement-floor question: the taxonomy's lexicon requires explicit rule-code strings and the agent paraphrases policy provisions more than it cites them. The paper's own reading is that the gap is too large to be explained by lexicon conservatism alone.

§2 Table 1 (P5), with the reading in §6

7 ALLOWs withdrawn

Adding governance prose made the agent less committal, not more.

All seven L0 ALLOW records withdrew to REVIEW once the L1 prose framed the substrate as procurement governance.

§2, F009 discussion

Cited in this piece

The work this piece rests on — the author’s own list first, then everything it links out to.

  1. 01
  2. 02
    When AI hedges and policy commits

    We ran 283 real UK procurement decisions through both an AI agent and MeshQu’s policy engine at the same moment, binding every verdict to a signed receipt.

  3. 03
  4. 04
  5. 05

How to cite

APA from the paper's own recommended citation; BibTeX derived from the same fields, not authored twice.

APA · from the paper’s own recommended citation
Carter, S. (2026). When precedents commit AI and policy pulls it back: a five-rung governance-context ladder on 283 procurement decisions, signed and verifiable. MeshQu Research Preview MRP-2026-03.

DOI — not yet assigned.

BibTeX · derived from the same fields
@techreport{carter2026precedents,
  author      = {Carter, Sam},
  title       = {When precedents commit AI and policy pulls it back},
  subtitle    = {Why giving an AI agent more governance context doesn’t make it steadily more decisive},
  institution = {MeshQu},
  type        = {Research Preview},
  number      = {MRP-2026-03},
  year        = {2026},
  month       = {5},
  url         = {https://www.meshqu.com/research/when-precedents-commit-ai-and-policy-pulls-it-back}
}

The programme in sequence

Each experiment answers the question the one before it left open, with the white paper as the argument they test.

  1. 00

    The Decision Proof Gap

    AI governance frameworks describe how decisions should be made — but the moment a decision is actually executed, the evidence to defend it almost never exists in a form anyone can verify.

    White paper · WP-PROOF-01 · 35 pp.
  2. 02

    When AI hedges and policy commits

    We ran 283 real UK procurement decisions through both an AI agent and MeshQu’s policy engine at the same moment, binding every verdict to a signed receipt.

    Research paper · MRP-2026-02 · 26 pp.
  3. 03

    When precedents commit AI and policy pulls it back You are here

    We showed one AI agent the same 283 UK procurement decisions five times, adding more of our governance rules each round — from nothing, up to the full policy.

    Research paper · MRP-2026-03 · 38 pp.
  4. 04

    Precedents, policy, and commitment

    We re-ran the same procurement decisions through an AI agent with the governance context broken apart piece by piece — to find out which piece was doing the work.

    Research paper · MRP-2026-04 · 46 pp.