On a Wednesday in November, an AML monitoring model at a UK retail bank flags a £42,000 wire transfer to a small business in Cyprus. A Tier-2 alert is raised. An analyst clears it within forty minutes after a phone call to the relationship manager.
Eight months later, a regulator picks that payment out of a sample and asks why it was cleared.
At that point the bank’s AI policy is not what is under examination. One decision is: which policy applied that Wednesday, what the analyst considered, what the model returned, and whether the clearance met the requirements in force at the time.
The bank does not lack an AI policy. At board level the CRO can point to frameworks, principles and control structures; documentation is in place, risk maps are signed off, and controls are reviewed quarterly. Governance at that level can be described. Governance at a single decision has to be demonstrated, and those are different tasks.
The system moves on, and the decision disappears
Follow the alert forward.
AML alert · Tier-2
£42,000 wire to Cyprus SME
- Flagged at
- Wed, Nov · 09:14
- Cleared at
- 09:54 (40 min later)
- Analyst
- J. Okafor
- Resolution
- Cleared after RM call
- Audit trail
- Alert ID · timestamp · free-text comment
Eight months later the FCA’s S166 review picks the same payment from a sample of 200 and asks why this transfer was cleared.
The first-line risk team finds the alert ID, the timestamp, the analyst’s name and a free-text comment, alongside logs of the model run, events from the case management system and traces from the audit log. Those are fragments of activity, not decision proof. The input can be found and the outcome can be found; the reasoning is inferred.
Meanwhile the surrounding system has changed. The policy still exists in the policy repo, but the version that applied that Wednesday in November has been superseded twice. The model still exists, but it was retrained in February against a different feature set. Thresholds have shifted and context has changed.
So the original decision cannot be reconstructed. The team cannot verify it against policy as it stood at the time of the alert, and cannot determine whether the criteria the analyst applied were the criteria the policy required. The answer that goes back to the FCA is produced anyway: plausible, defensible, and partly invented.
Governance is tested one decision at a time
Governance is not static. Policies evolve, models are retrained, thresholds are adjusted, ownership shifts. The question that actually gets asked is not what is the policy now? It is what was the policy at the moment this decision was made? — and whether that can be demonstrated.
Even where evidence exists, it is rarely independent. The same case management system that recorded the alert produces the explanation, and the same model registry that ran the inference reconstructs the reasoning. That is a system describing its own behaviour: procedural rather than forensic.
Day to day this goes unnoticed. Systems operate, decisions are made, outcomes are delivered, and everything appears governed — until someone asks why. Why this customer was declined. Why this transfer was cleared. Why this loan was approved. Not in aggregate, not in a dashboard, not in a board pack, but in one specific case on one specific Wednesday in November.
There is nothing to rebuild
AI governance programmes in tier-one banks concentrate on better policies, more documentation and stronger framework alignment. The failure this describes is not in the policy. It is in the decision.
A governed system does not reconstruct a decision later. It captures the decision as it is made — input, policy, context and outcome bound together as a single signed artefact. When the FCA letter arrives, there is nothing to stitch together from logs, and nothing to approximate. The decision can be read.
Asked
Why was this £42,000 wire cleared on Nov 13?
→RCP-AML-90217✓ VerifiedResolved in 1.8 seconds
Be precise about what such a record settles. It shows what was assessed, under which policy version, with what outcome, and it can be checked without re-running the system that produced it. It does not establish that the inputs were accurate, that the policy was well designed, or that clearing the payment was the right call. A human reviewer’s judgement still has to be recorded, and still has to be questioned.
Make the eight-month-old question answerable
Policies create intent and frameworks create structure, but the decision is where governance is tested. A practical standard: any decision that could be selected from a sample eight months from now should leave a record another person can read without rebuilding it.
Start with one such decision — a Tier-2 alert, a cleared payment, an override. Name the policy version it was assessed against, capture the input, context and outcome with it at the moment it happens, and record the reviewer’s resolution alongside. A worked example of that record is on the receipts page.