Questo articolo in italiano: Come verificare un agente AI in produzione: episodi, metodo e limiti
How to verify an AI agent in production: episodes, method, and limits
In brief. Verifying an AI agent means checking that "done" corresponds to done: I reread the state instead of the message, I ask the verifier what it doesn't measure, and I reproduce outside the agent's sandbox. Eight dated episodes from my systems, each with its source, and the limits of what I measure.
This guide exists because the question comes back every time an agent declares a result: how do you verify that "done" is done? I cover leading, lagging, and proxy metrics elsewhere; here I write about how to verify an AI agent's work when the agent itself declares the result.
"Done," without having done anything
On May 7, 2026, in trading-agents, the agent answers "IMPORT COMPLETATO – 51 posizioni salvate" ("import complete – 51 positions saved") without a single write (commit 1940df7). The message described work that didn't exist: no rows written, just the sentence.
Three months later, on August 28, 2026, in quant-advisor I find its twin: two agents with no data had been sitting at "IN ATTESA" ("waiting") for over two months, and I mark them "CIECO" ("blind") (commit b7865a4). Here too the declared state — waiting — hid the real one: with no data, there was nothing to wait for.
The rule I took from it: reread the state, not the message. If the agent says "saved," I look at the table; if it says "waiting," I check since when and for what. The three cases in full — the import that was never written, the wait with no data, the wrong measurement — are in The AI agent said "done," and it hadn't.
The number that looked like a measurement
On July 14, 2026, in social-quest, the cards report specificity 0.83 but "fuori-verticale 67%" ("off-vertical 67%"), later brought to 0% (STATUS.md line 458). A precise number — 0.83 — lived alongside two thirds of off-topic work: the measurement was true, it measured the wrong thing.
Since then, before trusting an agent's metric, I look for the second measure that could disprove it. A metric with no measure that contradicts it is a story with numbers in it.
The verifier that measured nothing
On October 7, 2026, the content factory issues 5 product "failed" verdicts over a verification script that didn't exist in the release: npm exited 1 with no output. The verifier worked fine; it was verifying nothing, and failing it with authority.
The next day, October 8, 2026, the subtler case: 55 releases rebound to the live release a3ed8f8 and 56 independent verifications passed, while the contracts declared coverage "SUFFICIENT" — but about 119 of 212 criteria are content criteria and the command doesn't measure them, a keyword-based estimate. A true green, on a narrower perimeter than declared.
It's the twin of the case I tell in When the verifier fails your green tests: there the green depended on the environment, here it depends on the perimeter. The question to ask every verifier is the same, and I record it in the traces that count as evidence: what don't you measure? These two cases, plus the bug that only existed in the sandbox, are told in full in False red and partial green.
The bug that only existed in the agent's sandbox
On October 9, 2026, a worker confirms a cause — "npm ci: Missing @emnapi/core" — that exists only in its sandbox: outside it, npm ci passes in 4 seconds. The red was real inside the sandbox and unreal outside it: the agent's sandbox is not the world.
The fix is easy to say and easy to skip: reproduce outside the agent's sandbox before believing the verdict. It's the same lesson as the red test that depended on the shell, where an environment variable decided the verdict — except nobody can see the agent's shell.
The review that actually counted
On October 9, 2026, an independent review with Muse rereads 51 published items against 173 criteria: 153 satisfied with file:line citations, 20 turned into follow-ups (independent review of October 9, 2026; figures reread in the approved sources at writing time). The point I care about is the mechanism: the check that every citation really exists in the file was done by the reviewer's script, external to the factory — not a framework feature.
Independent hands find what the author can no longer see: true for articles as it is for code.
Measuring yourself against your own spec
On September 26, 2026, Northstar measures CURE parity from 55.1 to 76.3 against a threshold of 90 (commit d386a02): the initial 55.1 came from an independent audit noting "the worker re-measured its own cure" (docs/cure-conformance-*.md files). Anyone who grades their own homework hands themselves free points; the second, independent measurement is the one that counts.
What I do now, in three checks
First: I reread the state instead of the message — tables, files, lines, not the agent's sentences.
Second: I ask the verifier what it doesn't measure — environment, perimeter, excluded criteria — before believing it.
Third: I reproduce outside the agent's sandbox, with hands independent of its own.
On AI-written reports I apply the same principle differently: every number must be in the fact sheet.
And when the measurement itself shifts from one run to the next, I measure the classifier inside the noise: same code, 30 and 27 errors across two runs, and the three moves that separate signal from noise.
Three sentences, eight episodes behind them. Not a statistic: a verifier's diary.
Facts and limits
FACT. On May 7, 2026, trading-agents declares "IMPORT COMPLETATO – 51 posizioni salvate" with no writes (commit 1940df7).
FACT. On August 28, 2026, quant-advisor marks "CIECO" two agents sitting at "IN ATTESA" for over two months with no data (commit b7865a4).
FACT. On July 14, 2026, social-quest reports specificity 0.83 with "fuori-verticale 67%", later 0% (STATUS.md line 458).
FACT. On October 7, 2026, the content factory issues 5 "failed" verdicts over a verification script missing from the release, npm exited 1 with no output.
FACT. On October 8, 2026, 55 releases are rebound to live a3ed8f8 with 56 independent verifications passed, but about 119 of 212 criteria are content criteria the command doesn't measure, a keyword-based estimate.
FACT. On October 9, 2026, a "npm ci: Missing @emnapi/core" cause exists only in the worker's sandbox; outside it npm ci passes in 4 seconds.
FACT. On October 9, 2026, the independent review with Muse over 51 items and 173 criteria yields 153 criteria satisfied with file:line citations and 20 follow-ups; the mechanical citation check was done by the reviewer's script, external to the factory.
FACT. On September 26, 2026, Northstar CURE parity goes from 55.1 to 76.3 against a threshold of 90; the 55.1 from an independent audit (commit d386a02, docs/cure-conformance-*.md files).
INTERPRETATION. Declaring is not doing; a verifier should be judged on what it doesn't measure too.
INTERPRETATION. I reread the state, I ask for the perimeter, I reproduce outside the sandbox: that's my method after these episodes, not a measured result.
INTERPRETATION. The question on the metrics page signals interest in measurement; my reading is that what's needed is a guide to verifying agents, not that classic metrics suffice.
Limits: eight episodes from my systems, not an industry survey. The review figures are a snapshot, not a time series. No claims beyond the cited sources: where the sources are silent, so am I.
Frequently asked questions
How do you verify that an agent actually finished?
I reread the state, not the message: tables, files, lines. If the agent declares a save, I look where it should have written; if it declares a wait, I check since when and for what. That's the rule from the first episode, above.
What should a good verifier measure?
What it declares it measures — and it should declare the rest too: environment, perimeter, excluded criteria. A green without a perimeter is a story, even when the numbers are big. I cover that above, in the verifier episodes.
Who should verify an agent's work?
Hands independent of its own: another run, another environment, a reviewer's script that doesn't belong to the framework being verified. Anyone who grades their own homework hands themselves free points; the independent review is the one that counts.
This article is also available in Italian.