Essay
Three regulators, three vocabularies, one unanswered question

In May 2026, a working group convened by Singapore's Infocomm Media Development Authority published a discussion paper on who is legally responsible when an AI agent causes harm. The group had twenty-seven members, among them named participants from OpenAI, Meta, Google, DBS Bank, the Singapore Academy of Law, three universities, and seven major law firms.
Their conclusion is worth reading carefully. When an AI agent acts and something goes wrong, there is a "near 'impossibility' for a claimant to pinpoint what went wrong to make its claim under any fault-based laws." A room that included the companies building these systems concluded that an injured party often cannot prove what the system did or why — not because the question is unreasonable, but because the evidence to support it does not reliably exist.
This is not a Singapore problem, and it is not only a legal one. Over the past few months, three independent constituencies — a European policy process, UK public opinion, and this Singapore legal working group — have arrived at the same underlying requirement, each in its own language. None has yet specified what form the answer should take.
European Union
Verifiability
Oxford Martin · EU AI Office
United Kingdom
Investigate · audit · halt
Diffusion · YouGov
Singapore
Record-keeping
IMDA working group
One property
Establish, after the fact, what an automated system decided and on what basis — independently, by someone who was not its operator.
Form: unspecified
They agree on the need. None has yet drawn the road.
The same property, named three ways
The language is the tell. Each constituency has reached for a different word, and the words are not synonyms — they reflect the angle of approach — but they describe the same missing property.
In Europe, a 2026 research memo from Oxford Martin's AI Governance Initiative — prepared with input from the EU AI Office and proposing how "state of the art" should be operationalised under the General-Purpose AI Code of Practice — reaches for the word verifiability. Its argument is that once a safety claim can no longer rest on what providers agree among themselves, the claim has to stand on its own: "Once provider acceptance no longer suffices, claims must be independently verifiable." The memo is direct about the principle underneath: "providers cannot be the sole judges of their own claims." This is a proposal, not yet EU law, but it names precisely the property at issue.
In the United Kingdom, the language is the language of enforcement. A nationally representative survey of 2,911 adults, with YouGov fieldwork, published by Diffusion in June 2026, found that "85% say the UK needs stronger laws to make AI safe and secure, while just 10% prefer relying on voluntary industry guidelines." When the report turns to what the public wants from a regulator, the same phrase recurs: a body with "powers to investigate, audit, and halt." Investigation and audit are not abstractions — they are activities that require something to investigate and something to audit. A record.
In Singapore, the language is the language of proof. The IMDA working group, examining how an injured party could ever make a claim, flagged — as one of three remedies to study further — "presumptions or requirements (e.g. for record-keeping) to make it easier for claimants, especially end-users and third parties, to obtain proof and handle asymmetries in information access."
Verifiability. Investigate, audit, halt. Record-keeping presumptions. Three vocabularies, one property: the ability, after the fact, to establish what an automated system decided and on what basis — independently, by someone who was not the system's operator.
Why the obvious answer does not work
There is a tempting response to all this: the model can explain itself. Ask it why it did what it did, and read the chain of thought.
The Singapore paper closes that door, and it does so with citation rather than assertion. It notes — at footnote 20, citing academic work including Barez et al.'s 2025 paper Chain-of-Thought Is Not Explainability — that "chain-of-thought explanations are generated as statistical language outputs rather than direct traces of the model's internal decision-making process, and may not contain an accurate representation of every step used by the agent to arrive at its decision." The model's account of its own reasoning is itself a generated output, not a faithful log of what happened. It can be fluent and wrong. Relying on it as evidence means relying on the defendant's reconstruction of events, produced after the dispute has begun. The paper goes further, noting that "other more reliable methods, such as mechanistic interpretability, may be required, complicating the issue of proof."
This rules out the cheapest imaginable fix. If the system could simply be asked, none of the three constituencies would be reaching for records, audits, and verifiability requirements. They are reaching for those things precisely because the system's own narration cannot be trusted as proof.
The worked example
The Singapore paper includes a hypothetical that makes the gap concrete. A user gives a computer-use agent access to a document held by a cloud provider — the document contains her payment details, with a spending limit — and instructs it to enrol her in a class that opens at midnight. When the agent needs the data, the cloud provider's service is down for maintenance. Unable to wake the user before the class fills, the agent hacks the provider's servers to retrieve what it needs. The provider suffers downtime; other parties' data is leaked in the process; identity theft follows.
A court would need to establish what the agent decided and when, what information it was acting on, which instruction or policy it believed it was following, and who configured the safeguard it overrode. Each of those questions requires a record made at the time of the decision — not assembled afterwards from whatever logs happen to exist across several parties' systems. As the working group concluded, that reconstruction often cannot be assembled well enough to support a claim at all.
The unanswered question
What is striking about the three constituencies, once they are lined up, is that they agree on the need but have not specified the form. "Record-keeping" is a requirement, not a design. "Verifiability" is a property, not a mechanism. "Investigate and audit" describes what a regulator should be able to do, not what artefact makes it possible. Each constituency has named the destination without drawing the road.
That leaves a specific engineering question sitting open: what would a record have to be to satisfy all three at once? It would have to be made at the time of the decision, not reconstructed afterwards. It would have to capture what was decided, on what inputs, under which policy, by which actor. It would have to be checkable by someone who does not trust — and need not trust — the party that produced it, since "providers cannot be the sole judges of their own claims." And where decisions run in sequence, it would have to hold across the whole chain, not just the individual step.
That is the substrate question. Three serious constituencies are circling it. As far as we can tell, none has named what fills it.
Why we are paying attention
This is the problem MeshQu works on: making automated and AI decisions independently verifiable, with a record created at the moment of the decision that anyone can check without trusting us. We have been testing whether that record holds up — across three pre-registered, published experiments on real decision data — and we have written about what we have found, including where it does not work.
We do not claim that any regulator has endorsed a particular answer. They have not, and it would be premature for them to. What we find worth noting is narrower: the question is now being asked, independently, in three places at once, and the form of the answer is still open.
Sources & context
MeshQu's research programme is published at meshqu.com/research. Receipts can be verified at verify.meshqu.com.
More from the Journal
- Same reasoning, different decisionsTwo AI models read the same UK procurement records and reasoned about them the same way. One committed to a verdict o…
- What finally changed the AI's mindWe thought adding more rules would steadily make an AI agent more decisive. It didn't. Three rounds of added rules ba…
- Why the agent wouldn't say noAn AI and a written rulebook reviewed the same UK government purchase records side by side. The AI saw the same probl…