Questo articolo in italiano: Agent factory: il loop che osserva, seleziona, esegue, verifica e impara

Agent factory: the loop that observes, selects, executes, verifies, and learns

In brief. The loop I use in my factories has five phases — observe, select, execute, verify, learn — and it holds for one reason only: each phase has its own judge, different from whoever executed it. Whoever observes does not select, whoever executes does not verify, whoever learns does not remember but writes tests. It is the thesis I derived from two factories and a handful of dated episodes: two points that coincide, not a sample. Published October 5, 2026.

Why five phases instead of one

The first loop I built had a single phase: "fix things." An agent looked at the project in the morning, adjusted something, wrote that everything was fine. It worked as long as I believed it, and I always believed it: it was the only voice in the loop. Then I discovered the data routine had been stalled for 33 days under its eyes, and I understood I had not built a loop. I had built a monologue with an audience of one.

The full story is the Italian piece on the factory that repairs software while we sleep. What matters here is the design that came out of it: splitting the monologue into five phases, and giving each an external judge. The separation is the mechanism, not the text's organization. Every phase that judges itself sooner or later becomes the point where the loop lies to itself, and lies in the most dangerous format: a verdict, not an error.

The cycle in a table, before the detail:

Phase What it decides Who judges it Episode that shaped it
Observe What is really happening A measure a machine can read 33 days of silence
Select What to do now A deterministic function The 21-row backlog
Execute How to do it One driver's seat at a time Four fixes closed
Verify Whether it worked An actor outside the session 14 green here, 7 red there
Learn What remains A test that replays the case The second missing accent

Observe: reality before opinions

Observing means one thing only: recording what happens without interpreting it. My loop always starts here, and every system I built started from a level I call observe, where it watches and nothing more. It looks like wasted time, and it is the opposite: it is the only phase that cannot fail from excess caution, because it touches nothing.

The episode that taught me to see it this way is the silent failure of July 24: routine stuck, no error, I notice on August 26. When I looked inside, one hour of audit over 115 files showed me everything. The strategy with the best numbers was fake: it read a cache that had not refreshed in two months. In the log I found 18 days never lived, and a 71-stock portfolio pinned at 0.00% while the market rose. Two more strategies had seen no fresh data in 78 days. All the detail is in the factory story: here the rule I drew from it counts. If a system cannot notice on its own that it stopped, calling it autonomous is generous: it is unattended.

The second lesson on observing came later, and it is subtler: observing the fact is not enough, you need the fact's scale. The governor watching my runs mistook a global human block for a local wait, and the driver's seat read free where it had to read occupied. The tests were green — 27 governor tests plus a 2,959-strong suite — because none asked the scope question: does it hold here or everywhere? A peer review found it, with CONCERN gov1:c5d56f63, and the fix made the block's scope machine-readable. I told it in the piece on the governor that misread: information with no scale written into a field is understood only from context, and context is understood by whoever was there; whoever arrives later — almost always a machine — finds only the fields.

Third lesson, the most uncomfortable: sometimes the observer generates on its own the noise it observes. My gap probe had a hand-compiled list of expected articles, with slugs invented months earlier, and produced tidy backlog for files never meant to exist under those names. I noticed only because the counts did not match what I saw. I discuss it in Backlog Zero: homemade noise is the most insidious, because it carries the system's signature.

Select: one job at a time, by written rule

With reality recorded, the loop must choose what to do now. This is the phase I took away from language models first: selection is made by a deterministic function, which always decides the same way under the same conditions and motivates every choice. Strategy stays with the model; the work queue does not.

The selector works in layers: only the first layer with something to propose advances — verified measures that unblock work, signals observed today, expired evidence waits, and last the backlog campaign. The backlog comes last by design, not modesty: a weeks-old row weighs less than this morning's measure. It is epistemic hierarchy. I described it from the rhythm side in "I no longer write prompts" and from the rows side in Backlog Zero, where I tell the day the count said 21 and the executable rows were few. The rest waited on dependencies, audits, or evidence I did not have.

There is one selector rule I care about in particular: if today's reading arrives malformed, selection stops instead of silently sliding onto the backlog. No invented work may be born from a broken observation. It is the same principle that governs escalation to the human: my factory used to stall to route the owner questions it could decide on its own — integration queues, internal mechanics, selector layers — and every stop looked like prudence while it was a defect. Two framework PRs closed the hole, first #294 then #309, with a single rule: the question to the owner fires only when something that changes for the product's user is at stake. For everything else the worker decides, inside the guardrails, and records it as fact. I told it in the false owner gates, with the diary row proving it: a row that left the human block with no owner answering.

Execute: one boundary, one driver's seat

Execution has a single rule, and the ledger enforces it, not the prompt: while a change is in flight on one point, that point accepts no others. One causal boundary at a time. Written in the prompt, a rule lasts as long as the model's memory; written in the ledger, it lasts forever.

The full cycle — observe, select, execute, verify, record — completed four fixes on a real project, PRs number 4, 5, 6, and 7, from the first to the last step with no intervention from me. Four closures with a receipt for every step: the evening the first cycle closed, I looked at the receipts and they seemed like bureaucracy. Then I counted. The story is the factory's, and the number stays small on purpose: it proves the cycle exists and works, not how far it scales.

On execution I also learned the second half of the craft, the one about executors when there are many of them. In Arvo, the coaching app where my agents work, I first put the specialists side by side: one computes progression, one picks exercises, one generates workouts. In my case the strain showed between three and four agents: queues waiting on each other, the same work done twice, with me in the middle sorting by hand. The answer was not a better agent, it was a small organization: roles with clear inputs and outputs, formal handoffs, a small model underneath (GPT-5-mini). When I had to change one exercise's selection criterion, I redesigned a handoff instead of rewriting an agent. I told it in "I stopped building agents": roles survive, the engine gets swapped.

Verify: green counts only where the verifier verifies

This is the phase holding up everything else, and the rule fits on one line: the verdict on a job arrives from outside whoever did its session, on the code actually released. Whoever executes does not certify. It is not distrust, it is the same separation that in a company keeps whoever spends distinct from whoever approves the budget. The full principle, with the other three controls I designed instead of removing, is in "Autonomy is not born from removing control".

The episode that taught me to take it seriously is the one of the failed green tests: a suite 14-out-of-14 green locally, run in the verifier's environment, scored 7 PASS against 7 FAIL. Same code, same tests: only the environment changed the verdict. The fault was not fragile tests but the implicit environment: working directory taken as default, loaded review never re-read, criteria tied to never-declared components. The fix, described in PR #323's body, attached each criterion to the components measuring it, and after the fix the suite passes with 3,216 greens in both environments. I told it in the piece on the verifier that fails: with no declared environment, the proof stays a story.

That case's twin concerns the clock instead of the environment. My runs' supervisor computed budget from the createdAt declared by the agent itself: the hour was written by whoever the budget was supposed to limit. With PR #339 the kernel-observed recordedAt dictates the budget, on both backends, declaring the legacy fallback where observation is impossible. The rule I drew from it, which I now apply to every measure governing something, is: no control may rest on data the controlled party can change. I told it in the piece on the agent's clock, keeping the PR's facts separate from my interpretation.

Learn: errors leave tests, not memories

The last phase is what makes each cycle different from the previous one: turning failures into lasting rules with their regression test. In the document governing the project, rules 17–26 were all born this way: each from a defect actually seen in execution, each with its test. They do not rain from the sky: each is a scar with its report.

The mechanism has a threshold written in advance: when the same defect shows up twice, a general gate fires. The freshest example is almost embarrassingly small: first a "così" with no accent, then a "più" with no accent. Spelling was covered by no gate, so on the second case a deterministic one was born, with its tests. And the rule I hold tightest is the threshold's reverse: from a single weak observation no lasting rule is born. Without it, coincidences would become laws, and the loop would learn to chase noise. I told both mechanisms in the machines' marginal gains: for machines, compound gain is accumulation of invariants, not a sum of micro-wins.

I keep the two writings separate: on one side what happened, on the other what we learned. Excluded roads count as much as confirmed causes: knowing what does not work is half the wealth. It is the loop's piece I described from the rhythm side: every well-made cycle leaves lasting state, and the next one starts higher.

Where what I saw ends

I speak for one person and two factories: the five-judges thesis is my interpretation, not a measured result, and I mark it as such because honesty requires it. I have no evidence the design holds at organizational scales I have not seen, and I do not claim it does. So far my loop has fixed code failures, without making product decisions: between "fixing alone" and "deciding alone" lies another boundary, with different rules, which I have not crossed.

The rest of the picture — what building the company around these loops means, with the limits of what I know — is in Building the Autonomous Company. This page is the operational section: the five phases, the judges, the episodes. If you start from zero, begin with observe and a measure: seven days of recorded observation teach you about your work more than any course.

Frequently asked questions

Which phase do you build first?

Observe, always. A system that watches and records without touching anything cannot break anything, and produces the diary from which you design everything else. The machine-readable measure comes right after: without it, the other four phases spin on emptiness.

Why can't the selector be a language model?

It can, but then the work order changes with the model's mood and nobody can say why one thing goes first today and another tomorrow. A written rule always decides the same way under the same conditions, and the choice can be re-read. Strategic decisions stay with the model; the queue does not.

Does the loop also decide what to do, or only how to do it?

In my case only how, and not even all of it: I decide the outcomes that matter, the system explicitly parks the cases it cannot solve, and I judge those. The line between "fixing alone" and "deciding alone" is the boundary I have not crossed, and it is declared as such.

What sets observe apart from closed loop?

How many phases run with no intervention. In observe the loop only watches; it climbs one level at a time — selection, execution, verification — and each level must be earned with machine-readable measures. Starting from total autonomy and adding controls after the first incident is the sequence I see botched most often.

This article is also available in Italian.