Skip to main content

Opening the black box of autonomous work

We connected MeshQu to selected decisions across our coding task lifecycle. The records helped us identify bottlenecks, question our measurements and begin improving both the work and the rules around it.

Sam Carter10 Oct 20267 min read
A researcher examines three gold decision tokens in a glass case beside a working brass mechanical arm and an open reference book in a sunlit courtyard.

An agent completes a task, hands work to another agent and moves on. The workflow keeps running. When someone asks why a particular step went ahead, the team needs more than a plausible explanation. It needs something it can examine.

Which policy applied? What information was used? What answer came back? Was there something a person should investigate?

We call this Decision Assurance: making the policy, evidence and decision boundaries around consequential work explicit, so they can be examined while the process is running and later.

Those are the questions we are building MeshQu to help teams answer. At chosen points in a workflow, it checks supplied information against approved policy and returns an assessment. Recording that check gives people evidence they can examine, discuss and learn from while the work continues.

On 8 October 2026, we began using MeshQu in our own agentic coding workflow, connecting it to the AI agents that help us build the product. We started in staging, in shadow mode: MeshQu advises and records while the work continues. Its assessments do not stop an agent or prevent a change.

The first benefit was a clearer view of where work was getting held up, supported by actual timing data. But the useful questions were not all about the agents. Some were about our own rules, measurements and assumptions.

Decision Assurance across the AI coding task lifecycle

A task passes through several stages: work begins, changes are built, reviewers examine them, and a decision is made about whether they can proceed. Different agents may handle different stages, with people involved where judgement or approval is required.

This is one practical form of AI agent governance: making the decisions around autonomous work inspectable rather than relying only on an agent’s explanation afterwards.

We connected MeshQu to selected decision points across that task lifecycle. At each checkpoint, our coordinating system supplies the relevant information. MeshQu assesses it against the policy in force and records the answer in a signed Decision Receipt.

A Decision Receipt is the signed, versioned record of that assessment: the policy version used, the information supplied and the answer returned.

Linking those receipts lets us follow the task rather than examine each agent's activity in isolation. Alongside timing and model information, we can investigate how the work progressed, where it slowed down and what supported the decisions along the way.

A single assessment may look reasonable on its own. Seeing it alongside earlier checks, handovers and later results can reveal a question that would otherwise be difficult to notice.

Raw logs may contain many of the underlying details. The value here is connecting them to the requirements and assessments around a decision. This is not access to a model's internal reasoning or a complete recording of everything an agent does. It is a connected view of the checkpoints we have chosen to assess.

The record gave us things to improve

Our early use led to changes in how we supplied information and connected the workflow. It also raised questions about what our controls were measuring.

Was the information keeping up with the work?

An early assessment prompted us to compare the resource-availability information MeshQu had received with the current state of the workflow. The reading was not keeping pace with the work.

We changed when that information was updated. With the information refreshed, we could distinguish an outdated session count from current capacity pressure. The record showed what the assessment had used; examining the surrounding work helped us identify what needed changing.

Did the connection reflect how work actually moved?

We compared the recorded checks with the moments we expected to see across the task lifecycle. That review led us to improve the connection.

It made coverage a concrete question. Saying the workflow was connected to MeshQu was not enough. Were the relevant checks being recorded at the right moments? A receipt cannot show a checkpoint that was never sent. Comparing the record with the wider workflow remains essential.

Did our measurements mean what we thought?

Reviewer sessions and recorded review rounds did not always correspond. We clarified that several reviewer passes could occur within one recorded round. That left a policy question: was the round count the quantity we intended to assess?

Capacity alerts raised a similar distinction. Running above a standing limit after a deliberate temporary change is different from exceeding even the temporary setting. Understanding an alert required examining the operating context, not just its label.

These were not all failures by an agent. Some concerned the way we had defined and measured the work.

We were improving two things at once: the workflow, and the rules and measurements we used to assess it.

Let AI agents help read the decision record

The same records can support an AI assistant helping the team investigate.

Through an OAuth connection we approved and can revoke, Claude read our staging decisions and receipts using MeshQu's MCP server. It helped us identify recurring rule results, connect related records and raise questions for investigation. That analysis helped shape this article.

Its explanations remained interpretations, not conclusions we accepted automatically. We could return to the receipts and compare them with the surrounding work. The record made the explanation open to challenge.

A teammate can inspect a particular decision. An assistant with approved read access can help examine a larger set and suggest where to look more closely. People decide which findings matter and what to investigate or change.

Some useful work also happens before there is a commit or pull request. A connected agent working locally can record an assessment that another authorised person or agent can inspect through MeshQu. Source control still records the code changes; the receipt contributes evidence around the work.

That means a handover need not depend entirely on the original agent's summary or the engineer's memory. The evidence can remain available to the next reader. It is information to assess, not automatic permission to act.

Improve the process, not just the pass rate

Different findings call for different responses. Refreshing stale information is not the same as changing a threshold. Repairing a connection is not the same as deciding that an action needs human approval.

The record helps people distinguish those questions rather than treating every alert as a reason to stop, or every passing check as reassurance that the process is sound.

A signed receipt does not make the supplied information true. It does not establish that the decision was correct, lawful or fair, or that an action actually occurred. People use the assessment alongside other evidence to understand what it means.

Keeping earlier records and policy versions matters for improvement too. A process producing fewer alerts might have improved, or we might simply have made its rules easier to pass. Those are different outcomes. The history gives us a basis for investigating the difference, though a before-and-after comparison alone does not establish cause.

For us, this is Decision Assurance becoming useful in everyday work: evidence we can question, act on and return to.

What comes next

Next week, we will keep MeshQu in shadow mode while we tune the rules and learn from the results. We want to understand which findings deserve attention, which reflect intentional operating choices and where our measurements need to be clearer.

That is preparation for taking selected checks out of shadow mode, when their assessments can begin to influence what proceeds, what waits and what needs a person's decision. Before that happens, we need confidence in the rules, the information they use and how the surrounding workflow responds. Changes to those controls remain deliberate and human-approved.

Part of the preparation is making receipts useful to agents and team members throughout the task lifecycle, so the next decision can draw on what has already been assessed rather than starting again.

We are also preparing controls for our coordinating system to scale build agents up or down according to available capacity and the requirements of the work. We can tune thresholds and how readily particular checks call for attention. The question will be whether those changes help, not simply whether we can make the alerts disappear.

This is an early account of internal use, not a completed study of productivity or control effectiveness. We have identified places to improve and made practical changes; their longer-term impact remains to be measured.

We are starting with a small team of coding agents, but the broader vision extends across people, conventional systems and AI. They may contribute in different ways, yet the questions remain recognisable: what criteria applied, what information was available, who was permitted to decide and what happened next?

MeshQu's purpose is to make that evidence useful wherever decisions matter. Not only for a future audit, but for the people and agents trying to understand and improve the work today.

We are using evidence from the task lifecycle to improve how the work runs, then keeping the record needed to examine what happens next.

Working with AI agents?

See how MeshQu adds Decision Assurance around consequential agent decisions, or bring one decision to Qu and map where policy, evidence and human judgement sit today.

Cited in this piece

Sources and further reading.

  1. 01
  2. 02

Read next

Continue exploring this question.