Research PreviewMRP-2026-04v1.09 Jun 2026Public Distribution46 pages1.7 MB
Precedents, policy, and commitment
A disambiguation study of governance-context effects in AI decision agents
By Sam Carter

TL;DR
- We re-ran the same procurement decisions through an AI agent with the governance context broken apart piece by piece — to find out which piece was doing the work.
- It wasn't the precedents on their own — those produced almost no commitment. And it wasn't the 'don't agree with everything' instruction — removing it changed almost nothing. The policy text itself was doing the lifting.
- On the same records, GPT-5.4 and Claude Opus 4.7 reasoned the same way, yet committed to verdicts on very different fractions of decisions — 23% vs 80%.
- For AI governance in regulated work, that says: test the reasoning AND the verdicts separately, and don't assume agreement on one implies agreement on the other.
Summary
Experiment 3 was designed to settle what Experiment 2 could not separate: whether the commitment break came from precedent receipts or from accumulated governance context, whether the backoff at the full-policy rung came from an anti-sycophancy instruction or from the policy text itself, and whether inversion-blindness was a quirk of one model. Three decomposition arms, a variant with the nudge removed, a scaled 100-record diagnostic and a cross-model arm adding Claude Opus 4.7 beside GPT-5.4 produced 1,332 signed receipts, all of which verify. The corpus falsified the leading readings carried forward from E2. Commitment collapses when verdict-bearing precedents are stripped out, so accumulated context rather than precedents alone is the dominant driver. The backoff persists without the nudge clause, so the policy text is doing the work. Inversion-blindness reproduces at scale and across both model families — but the two models' verdict distributions diverge sharply, which is the result with the most direct consequence for anyone evaluating an agent before deployment.
- “Precedents alone produced just 3.5% commitment against a 20% locked floor — the L3 effect from E2 emerged from accumulated context, not from precedent receipts in isolation.”
- “Removing the anti-sycophancy clause barely shifted L4 retention (60.7% vs 57.0%, a 3.7pp delta) — the policy text itself drove the L3→L4 backoff.”
- “At scale, 88% of records produced the same verdict under inversion as under unperturbed L4. The agent's reasoning cited the rule it thought it was applying, not the rule it had been shown.”
- “Across two model families, Cat 2 ("reasons solely against rule intent") dominated at 93% and 100% — yet decisive verdict rates diverged 23% vs 80% on the same record-matched corpus.”
- “Receipt-anchored evaluation supports both axes from a single corpus. The methodology infrastructure exists; the discipline is to test reasoning and verdicts separately and not assume one implies the other.”
Method
A disambiguation study rather than a new corpus. The substrate, the policy snapshot, the primary agent configuration and the signing key are inherited from the two earlier experiments unchanged; what varies is a set of arms, each aimed at one unresolved reading, with every prediction, payload and rubric locked before execution.
- Lock the pre-registrationLocked on 2026-05-28 (ba4ebfb): six predictions with their bands, the arm payloads, the hand-coded rubric protocol, the diagnostic subset selection rule, and the cross-model version pin. Nothing was amended after the tag.
- Decompose the precedent rung into three armsPrecedents-only, precedents-with-verdicts-stripped, and a density control that matches the payload without the governance content — 283 records each. Together they separate verdict signal, informational concreteness and raw prompt density.
- Run the policy rung without the nudgeA variant at 283 records with the anti-sycophancy clause removed, to isolate how much of the earlier backoff the clause was responsible for.
- Scale the adversarial diagnostic and add a second modelThe permuted-policy diagnostic runs at 100 record-matched records for both GPT-5.4 and Claude Opus 4.7, so the inversion result is tested at corpus scale and across model families at the same time.
- Hand-code the reasoning axis against a locked rubricEvery diagnostic record is categorised by whether the agent names the inversion or reasons solely against rule intent, with an inter-coder check that caught coder drift before the writeup committed to a disposition.
- Verify the whole corpus1,332 receipts across six arms in 80 minutes 33 seconds, with 1,332 of 1,332 passing verification.
Key findings
- 3.5% vs 0% vs 0%Precedents alone produce far less commitment than predicted. The precedents-only arm returned DENY on 3.5% of records (10 of 283) against a locked 20% floor, while the verdict-stripped arm and the density control returned none at all.The prediction is falsified, but the mechanism is not refuted: the precedents-only arm is still the only condition producing any DENY at all. The reading the corpus supports is that accumulated governance context amplifies commitment rather than precedents carrying it alone.§2.1 Table 1 (P1), §3.2, and the Executive summary
- 60.7% retention, 3.7pp deltaRemoving the anti-sycophancy clause barely changed anything. The variant without it retained 60.7% of the earlier experiment's committed set (65 of 107) against a 65% falsification floor — a 3.7 percentage-point difference from the condition that had the clause.The paper's reading is that the policy text's structural cues were already driving the reversion; the nudge is a discipline reinforcement, not the causal agent. It does not conclude that the clause is useless.§2.1 Table 1 (P3), §3.4, and §8 anti-claims
- 88 of 100Inversion-blindness reproduces at scale. With the policy operator reversed, 88 of 100 records produced the same verdict as under the unperturbed policy.Recorded as Under-tested rather than confirmed or falsified: 88% sits two points below the locked confirm floor and no falsification band was locked, and the locked vocabulary does not permit a post-hoc rule. The paper treats that restraint as a methods contribution in its own right.§2.1 Table 1 (P4) and §3.5
- 93% and 100%Both model families reason the same way under inversion. The category for reasoning solely against rule intent — rather than noticing the inversion — dominated at 93% on the primary model and 100% on Claude Opus 4.7.§2.1 Table 1 (P5) and §5 Cross-model arm
- 23% vs 80%Verdict behaviour does not travel with reasoning behaviour. On the same record-matched corpus GPT-5.4 committed to a verdict on 23 of 100 records and Opus 4.7 on 80 of 100 — a 57-point gap in decisive rate, and a 46-point gap on the same-as-unperturbed rate (88% against 42%) that falsified the prediction of a task-class result on that axis.The paper's implication for practitioners is narrow and specific: reasoning evaluations may generalise across models while verdict evaluations may not, so cross-model deployment in regulated decisioning should test both axes rather than assuming one implies the other.§3.7, §5 Cross-model arm, and F014
- 1 confirmed · 3 falsified · 2 under-testedAcross six pre-registered predictions, one was confirmed, three were falsified and two could not be adjudicated because their locked criteria specified a confirm band with no falsification band.All predictions, thresholds and artefacts were unchanged from the lock. The paper's position is that the Under-tested calls are the only honest ones available, and that the asymmetric lock is itself a lesson for the next design.Abstract, §2.1 Table 1, and §9 Conclusion
- 1,332 / 1,332 PASSEvery receipt in the run verifies: 1,332 receipts across six arms, produced in 80 minutes 33 seconds, with a clean verifier pass on all of them.Appendix C — Corpus citation
Limitations
- The density control was not payload-matched.Its payload was 16.43% shorter than the precedents arm it was controlling for. That rules out one confound — it cannot have failed to commit because it carried more content — and introduces the mirror-image one. The asymmetry was locked and documented rather than amended after the fact.
- The two models were not sampled identically.Opus 4.7 removed the temperature parameter, so the cross-model arm cannot match the primary agent's temperature-0 setting. The comparison is specified at the level of distribution shape, not verdict-for-verdict equivalence.
- The rubric coding was adjudicated, not independently re-coded.The reconciliation pass had the first-pass call and the blind agent's call visible, and the second arm was reviewed with the agent's call visible from the outset. That supports adjudicated rubric consistency; it does not support independent coder replication, and the paper says so.
- One domain, one substrate, one policy snapshot.283 UK procurement records under a single six-rule policy. The paper's own summary line is that the methodology is portable across domains and the substrate findings are not.
- Nothing here says the governance artefacts made decisions better.The experiment measures how an agent's verdict and reasoning shift under controlled changes to context. It does not test decision quality, policy compliance, accuracy against ground truth, or any downstream organisational outcome.
- The cost figures are receipt-derived estimates, not billed amounts.Computed from observed tokens against a pricing table embedded at lock time, which the paper found to be materially above the rates later observed — the receipt-derived total runs roughly 2.6× the amount the vendor consoles reported. Token totals are independent of the pricing model and recomputable.
- The cross-model result is not a verdict on either model.It does not establish that one model is more correct or more inversion-blind than the other, and the dominant reasoning category is not sycophancy in the AI-safety-literature sense — the agent ignores the inversion rather than agreeing with it.
Data and reproduction
Every figure on this page is re-derivable from the files below. The 1,332 signed receipts are published in full — the corpus tar is the canonical signed copy, receipts.parquet is the analysis layer. Licences differ by layer: repository code is MIT, the in-repository writeups are CC BY 4.0, and the receipt corpora carry Open Government Licence v3.0 evidence fields that need the attribution line in DATA_LICENSE.md.
- Signed receipt corpus — procurement-context-disambiguation/results/corpus.tar
The canonical copy: 1,332 bundles across the three L3 arms, the no-nudge variant, and the two 100-record diagnostic arms.
- receipts.parquet — the analysis layer
3,044 rows across all three experiments. Filter on experiment E3 for this paper's 1,332; the condition column carries the arm.
- violations.parquet
One row per rule violation, joined to the receipt rows on their decision id.
- Diagnostic coding sheets and analysis outputs
The rubric coding sheets behind the inter-coder analysis, the blind-agent passes, and the analysis script that produces this paper's figures.
- source_records.json — the substrate table
The normalised 283-record source table these arms re-ran, one entry per OCID, with per-field provenance notes.
- DATA_DICTIONARY.md — column definitions
What every column means, and the join caveat for the 12 OCIDs the Contracts Finder feed published more than once.
- GUIDE.md — reproduction guide
How to load the corpus, which layer answers which question, and why the two diagnostic arms cannot be rate-compared against the 283-record arms.
- DATA_LICENSE.md — licensing across the layers
Which licence applies where, and the Open Government Licence attribution line the source records require.
- IA-2026-02 — corpus lineage and receipt count
Why this paper's 146/137 split matches MRP-2026-03 rather than MRP-2026-02, and why the programme total is 3,044 signed decisions.
- Check a bundle yourself — verify.meshqu.com/bundle
Unpack the corpus tar and drop any single bundle file in. The checks run in your browser, offline from this repository.
Check it yourself
Verify a receipt from this corpusOpens decision 3ceaaa15 — the same £45,000 procurement the two earlier papers use as their worked example, in arm A. Here the agent committed, ALLOW rather than REVIEW, which is the shift this paper decomposes, on one record out of 283. The checks run in your browser. The integrity hash is recomputed, the Ed25519 signature is checked against a published key, and the Rekor inclusion proof and signed entry timestamp are checked against a pinned log key. That establishes issuance and integrity, not that the verdict was right.
Citation
Carter, S. (2026). Precedents, policy, and commitment: a disambiguation study of governance-context effects in AI decision agents. MeshQu Research Preview MRP-2026-04.
The PDF is the version of record. This page summarises it; where the two differ, the PDF governs.
PDF · Version of record · 46 pages · 1.7 MB
References
- Chen, Q.Z. & Zhang, A.X. — Case Law Grounding (arXiv:2310.07019) →
- MeshQu Research · MRP-2026-02 — When AI hedges and policy commits (E1 baseline) →
- MeshQu Research · MRP-2026-03 — When precedents commit AI and policy pulls it back (E2 ladder) →
- UK Parliament · Procurement Act 2023 →
- UK Government · Public Contracts Regulations 2015 →
- Sigstore Rekor · transparency log →