Questo articolo in italiano: Agenti AI in azienda: architettura, costi e organizzazione
AI agents in business: architecture, costs, and organization
In brief. An AI agent is a program built on a language model that does more than answer: it receives a goal, decides the steps, uses tools and data, and delivers a result. Bringing agents into a company raises three practical problems: how to make them collaborate, what they really cost, and how to keep the system manageable. In this guide I collect what I learned building multi-agent products, and point to the in-depth articles.
What sets an AI agent apart from a chatbot
A chatbot answers one question at a time. An agent works toward goals: it breaks the task down, calls external tools (databases, APIs, other agents), checks the result, and decides whether another step is needed. For the basics of how language models and prompts work, I always start from the Italian article on how ChatGPT works. When the steps become a cycle that observes, executes, and verifies, the work changes nature, as I describe in the Italian piece I no longer write prompts, I orchestrate loops.
Organizing agents like a company
With a single agent everything looks simple. With ten, the problems are the same as in a human organization: who does what, how work passes along, who checks quality. In Stop Managing AI Agents, Start Building Organizations I describe the approach I use: an org chart of agents with fixed and dynamic roles, formal handoffs, and metrics per role. Working with three agents side by side taught me you need organizations, not geniuses: roles and handoffs. But autonomy only comes from designing control, grounded in observed data — never in data from the party being controlled, as in the case of the supervisor that trusted the agent's clock. When loops run on their own, the autonomous company is born.
What they really cost
The cost of an agent system is not just the model's per-token price: the context resent on every call, retries, quality checks, and tests all weigh in. For the basics you find the formula in the Italian article on what using an LLM costs; for the costs that emerge as the system grows, read the Italian guide on the hidden technical debt of AI agents, where I explain why the "tokens times cost times requests" estimate understates real spend and how an abstraction layer over providers helps keep it under control. Switching models is a change of quality contract, not a cost cut: I did it with a paired A/B and rollback written first. To estimate spend I use the AI agent cost estimator (in Italian).
Data comes before the model
An agent decides well only with access to the right data, fresh and governed. In B2B SaaS this means rethinking data architecture and the role of data teams, as I discuss in the Italian article on LLMs in B2B SaaS.
Two real cases: Arvo and Zeno
- Arvo is a personal training app built on a multi-agent architecture: each agent owns a specific task (planning, choosing exercises, validating changes, analyzing progress) and the system coordinates them.
- Zeno is a personal growth app offering five-minute daily micro-actions; it uses a RAG (retrieval-augmented generation) system to personalize suggestions.
I build both in the Aetha AI lab: if you want to see them live, you find Arvo and Zeno on their sites. About the lab I wrote what it means to build Aetha as a full-stack AI company instead of a software house with AI (in Italian).
The decisions to make before starting
| Decision | Question to ask | Where to go deeper |
|---|---|---|
| One agent or many | Can the task split into roles with clear inputs and outputs? | Arvo |
| Coordination | Who hands work to whom, and who checks quality? | Agent organizations |
| Costs | How much context do I send per call, and how many calls do I make? | LLM costs |
| Maintenance | Can I switch model or provider without rewriting everything? | Tech debt |
| Data | Does the agent have access to reliable, fresh data? | LLMs and data |
From the lab: field episodes
- The factory that repairs software while we sleep: a routine stuck for 33 days with nobody noticing, and the loop born to find failures on its own.
- Marginal gains for machines: turning every failure into a lasting rule with a test, so systems improve each time they err.
- Backlog Zero: 21 open issues, few executable — the problem is knowing which are real.
- When AI becomes the interface: if a chat were enough, what is SaaS for? In Arvo, screens, guards, and physical context remain.
- The governor that misread a global block: a global wait mistaken for a local one, found by peer review and not by tests.
- When the verifier fails green tests: suite 14/14 locally but 7 out of 14 from the verifier — declared green counts only once reproduced.
- The false owner gates: technical choices parked as owner decisions — over-escalation is a defect.
- The retry that kept receiving the dead row: resending the same idempotency key returned the terminal failure — idempotency needs failure semantics.
- The plan denied by execution: the plan said one thing, execution another — I kept Predicted next to Observed.
- The red test that depended on the shell: it blocked the release only in my shell because it read the production catalog.
Frequently asked questions
What is an AI agent?
A system built on a language model that pursues a goal over several steps: it plans, uses tools and data, verifies the result, and decides how to proceed.
What is the difference between an AI agent and a chatbot?
The chatbot answers one request at a time; the agent completes a whole task, including using other tools or other agents, until it delivers a result.
What does an AI agent system cost?
It depends on the model, the amount of context sent per call, and the number of calls. Beyond token cost you must count retries, quality checks, and maintenance, which often weigh more than a single answer.
Where do you start with AI agents in a company?
With a repetitive, well-defined process, with clear inputs and outputs and a simple way to verify result quality. One agent that truly works beats ten agents with no collaboration rules.
This article is also available in Italian.