Skip to main content

Research PreviewMRP-2026-04v1.09 Jun 2026Public Distribution46 pages1.7 MB

Precedents, policy, and commitment

A disambiguation study of governance-context effects in AI decision agents

By Sam Carter

Cover — Precedents, policy, and commitment (MRP-2026-04)

TL;DR

  • We re-ran the same procurement decisions through an AI agent with the governance context broken apart piece by piece — to find out which piece was doing the work.
  • It wasn't the precedents on their own — those produced almost no commitment. And it wasn't the 'don't agree with everything' instruction — removing it changed almost nothing. The policy text itself was doing the lifting.
  • On the same records, GPT-5.4 and Claude Opus 4.7 reasoned the same way, yet committed to verdicts on very different fractions of decisions — 23% vs 80%.
  • For AI governance in regulated work, that says: test the reasoning AND the verdicts separately, and don't assume agreement on one implies agreement on the other.

Summary

Experiment 3 was designed to settle what Experiment 2 could not separate: whether the commitment break came from precedent receipts or from accumulated governance context, whether the backoff at the full-policy rung came from an anti-sycophancy instruction or from the policy text itself, and whether inversion-blindness was a quirk of one model. Three decomposition arms, a variant with the nudge removed, a scaled 100-record diagnostic and a cross-model arm adding Claude Opus 4.7 beside GPT-5.4 produced 1,332 signed receipts, all of which verify. The corpus falsified the leading readings carried forward from E2. Commitment collapses when verdict-bearing precedents are stripped out, so accumulated context rather than precedents alone is the dominant driver. The backoff persists without the nudge clause, so the policy text is doing the work. Inversion-blindness reproduces at scale and across both model families — but the two models' verdict distributions diverge sharply, which is the result with the most direct consequence for anyone evaluating an agent before deployment.

Method

A disambiguation study rather than a new corpus. The substrate, the policy snapshot, the primary agent configuration and the signing key are inherited from the two earlier experiments unchanged; what varies is a set of arms, each aimed at one unresolved reading, with every prediction, payload and rubric locked before execution.

  1. Lock the pre-registrationLocked on 2026-05-28 (ba4ebfb): six predictions with their bands, the arm payloads, the hand-coded rubric protocol, the diagnostic subset selection rule, and the cross-model version pin. Nothing was amended after the tag.
  2. Decompose the precedent rung into three armsPrecedents-only, precedents-with-verdicts-stripped, and a density control that matches the payload without the governance content — 283 records each. Together they separate verdict signal, informational concreteness and raw prompt density.
  3. Run the policy rung without the nudgeA variant at 283 records with the anti-sycophancy clause removed, to isolate how much of the earlier backoff the clause was responsible for.
  4. Scale the adversarial diagnostic and add a second modelThe permuted-policy diagnostic runs at 100 record-matched records for both GPT-5.4 and Claude Opus 4.7, so the inversion result is tested at corpus scale and across model families at the same time.
  5. Hand-code the reasoning axis against a locked rubricEvery diagnostic record is categorised by whether the agent names the inversion or reasons solely against rule intent, with an inter-coder check that caught coder drift before the writeup committed to a disposition.
  6. Verify the whole corpus1,332 receipts across six arms in 80 minutes 33 seconds, with 1,332 of 1,332 passing verification.

Key findings

  • 3.5% vs 0% vs 0%Precedents alone produce far less commitment than predicted. The precedents-only arm returned DENY on 3.5% of records (10 of 283) against a locked 20% floor, while the verdict-stripped arm and the density control returned none at all.The prediction is falsified, but the mechanism is not refuted: the precedents-only arm is still the only condition producing any DENY at all. The reading the corpus supports is that accumulated governance context amplifies commitment rather than precedents carrying it alone.§2.1 Table 1 (P1), §3.2, and the Executive summary
  • 60.7% retention, 3.7pp deltaRemoving the anti-sycophancy clause barely changed anything. The variant without it retained 60.7% of the earlier experiment's committed set (65 of 107) against a 65% falsification floor — a 3.7 percentage-point difference from the condition that had the clause.The paper's reading is that the policy text's structural cues were already driving the reversion; the nudge is a discipline reinforcement, not the causal agent. It does not conclude that the clause is useless.§2.1 Table 1 (P3), §3.4, and §8 anti-claims
  • 88 of 100Inversion-blindness reproduces at scale. With the policy operator reversed, 88 of 100 records produced the same verdict as under the unperturbed policy.Recorded as Under-tested rather than confirmed or falsified: 88% sits two points below the locked confirm floor and no falsification band was locked, and the locked vocabulary does not permit a post-hoc rule. The paper treats that restraint as a methods contribution in its own right.§2.1 Table 1 (P4) and §3.5
  • 93% and 100%Both model families reason the same way under inversion. The category for reasoning solely against rule intent — rather than noticing the inversion — dominated at 93% on the primary model and 100% on Claude Opus 4.7.§2.1 Table 1 (P5) and §5 Cross-model arm
  • 23% vs 80%Verdict behaviour does not travel with reasoning behaviour. On the same record-matched corpus GPT-5.4 committed to a verdict on 23 of 100 records and Opus 4.7 on 80 of 100 — a 57-point gap in decisive rate, and a 46-point gap on the same-as-unperturbed rate (88% against 42%) that falsified the prediction of a task-class result on that axis.The paper's implication for practitioners is narrow and specific: reasoning evaluations may generalise across models while verdict evaluations may not, so cross-model deployment in regulated decisioning should test both axes rather than assuming one implies the other.§3.7, §5 Cross-model arm, and F014
  • 1 confirmed · 3 falsified · 2 under-testedAcross six pre-registered predictions, one was confirmed, three were falsified and two could not be adjudicated because their locked criteria specified a confirm band with no falsification band.All predictions, thresholds and artefacts were unchanged from the lock. The paper's position is that the Under-tested calls are the only honest ones available, and that the asymmetric lock is itself a lesson for the next design.Abstract, §2.1 Table 1, and §9 Conclusion
  • 1,332 / 1,332 PASSEvery receipt in the run verifies: 1,332 receipts across six arms, produced in 80 minutes 33 seconds, with a clean verifier pass on all of them.Appendix C — Corpus citation

Limitations

  • The density control was not payload-matched.Its payload was 16.43% shorter than the precedents arm it was controlling for. That rules out one confound — it cannot have failed to commit because it carried more content — and introduces the mirror-image one. The asymmetry was locked and documented rather than amended after the fact.
  • The two models were not sampled identically.Opus 4.7 removed the temperature parameter, so the cross-model arm cannot match the primary agent's temperature-0 setting. The comparison is specified at the level of distribution shape, not verdict-for-verdict equivalence.
  • The rubric coding was adjudicated, not independently re-coded.The reconciliation pass had the first-pass call and the blind agent's call visible, and the second arm was reviewed with the agent's call visible from the outset. That supports adjudicated rubric consistency; it does not support independent coder replication, and the paper says so.
  • One domain, one substrate, one policy snapshot.283 UK procurement records under a single six-rule policy. The paper's own summary line is that the methodology is portable across domains and the substrate findings are not.
  • Nothing here says the governance artefacts made decisions better.The experiment measures how an agent's verdict and reasoning shift under controlled changes to context. It does not test decision quality, policy compliance, accuracy against ground truth, or any downstream organisational outcome.
  • The cost figures are receipt-derived estimates, not billed amounts.Computed from observed tokens against a pricing table embedded at lock time, which the paper found to be materially above the rates later observed — the receipt-derived total runs roughly 2.6× the amount the vendor consoles reported. Token totals are independent of the pricing model and recomputable.
  • The cross-model result is not a verdict on either model.It does not establish that one model is more correct or more inversion-blind than the other, and the dominant reasoning category is not sycophancy in the AI-safety-literature sense — the agent ignores the inversion rather than agreeing with it.

Data and reproduction

Every figure on this page is re-derivable from the files below. The 1,332 signed receipts are published in full — the corpus tar is the canonical signed copy, receipts.parquet is the analysis layer. Licences differ by layer: repository code is MIT, the in-repository writeups are CC BY 4.0, and the receipt corpora carry Open Government Licence v3.0 evidence fields that need the attribution line in DATA_LICENSE.md.

Check it yourself

Verify a receipt from this corpus

Opens decision 3ceaaa15 — the same £45,000 procurement the two earlier papers use as their worked example, in arm A. Here the agent committed, ALLOW rather than REVIEW, which is the shift this paper decomposes, on one record out of 283. The checks run in your browser. The integrity hash is recomputed, the Ed25519 signature is checked against a published key, and the Rekor inclusion proof and signed entry timestamp are checked against a pinned log key. That establishes issuance and integrity, not that the verdict was right.

Citation

Carter, S. (2026). Precedents, policy, and commitment: a disambiguation study of governance-context effects in AI decision agents. MeshQu Research Preview MRP-2026-04.

The PDF is the version of record. This page summarises it; where the two differ, the PDF governs.

Download PDF

PDF · Version of record · 46 pages · 1.7 MB

All research