Research PreviewMRP-2026-04v1.09 Jun 2026Public Distribution46 pages1.7 MB
Precedents, policy, and commitment
A disambiguation study of governance-context effects in AI decision agents
By Sam Carter

TL;DR
- We re-ran the same procurement decisions through an AI agent with the governance context broken apart piece by piece — to find out which piece was doing the work.
- It wasn't the precedents on their own — those produced almost no commitment. And it wasn't the 'don't agree with everything' instruction — removing it changed almost nothing. The policy text itself was doing the lifting.
- On the same records, GPT-5.4 and Claude Opus 4.7 reasoned the same way, yet committed to verdicts on very different fractions of decisions — 23% vs 80%.
- For AI governance in regulated work, that says: test the reasoning AND the verdicts separately, and don't assume agreement on one implies agreement on the other.
Abstract
E2 left three open questions about why an AI agent's verdicts shift as governance context is added. E3 disambiguated them. Three L3 arms separated precedents from raw context density; an L4-without-nudge variant isolated the policy text from the anti-sycophancy clause; a scaled n=100 diagnostic re-tested inversion-blindness; a cross-model arm ran the same records on Claude Opus 4.7 alongside GPT-5.4. 1,332 signed Decision Receipts anchored to Sigstore Rekor at the v0.3 pre-registration lock. The headline finding sits cross-axis: both models exhibit the same inversion-blind reasoning pattern (Cat 2 at 93% and 100%) yet commit to verdicts on materially different fractions of records (23% vs 80%). Reasoning portability and verdict portability appear to be distinct properties.
- “Precedents alone produced just 3.5% commitment against a 20% locked floor — the L3 effect from E2 emerged from accumulated context, not from precedent receipts in isolation.”
- “Removing the anti-sycophancy clause barely shifted L4 retention (60.7% vs 57.0%, a 3.7pp delta) — the policy text itself drove the L3→L4 backoff.”
- “At scale, 88% of records produced the same verdict under inversion as under unperturbed L4. The agent's reasoning cited the rule it thought it was applying, not the rule it had been shown.”
- “Across two model families, Cat 2 ("reasons solely against rule intent") dominated at 93% and 100% — yet decisive verdict rates diverged 23% vs 80% on the same record-matched corpus.”
- “Receipt-anchored evaluation supports both axes from a single corpus. The methodology infrastructure exists; the discipline is to test reasoning and verdicts separately and not assume one implies the other.”
PDF · 46 pages · 1.7 MB
References
- Chen, Q.Z. & Zhang, A.X. — Case Law Grounding (arXiv:2310.07019) →
- MeshQu Research · MRP-2026-02 — When AI hedges and policy commits (E1 baseline) →
- MeshQu Research · MRP-2026-03 — When precedents commit AI and policy pulls it back (E2 ladder) →
- UK Parliament · Procurement Act 2023 →
- UK Government · Public Contracts Regulations 2015 →
- Sigstore Rekor · transparency log →