An agent with ten tools and a long instruction works while the task is narrow. Then rules for three roles land in one prompt, the tool list grows to thirty, and the model starts mixing them up. The usual engineering answer is to split: several narrow agents and an explicit way to hand work between them. There are only a few ways to do that, and each has a name.
Why one agent stops being enough
An agent picks its next step from everything in its context. Each new duty adds text, and the model chooses worse because it chooses from more. An agent with five tools for one role picks more accurately than one with thirty tools for five roles. Splitting also lets you test each agent separately and take dangerous rights away from agents that do not need them. The price: every hop is another model call, and every boundary loses context.
Pipeline: agents one after another
Agents form a chain and the order is fixed in code: parse the email, find the order, write the reply. Each gets only the previous output. Predictable and testable per link; breaks when the route should depend on content.
Coordinator: one decides who gets the task
A router agent does not solve the task, it picks a specialist. Specialists look like tools with descriptions, and descriptions must be written as boundaries, including what the agent does not do. Overlapping descriptions make the choice random; log the coordinator's choice and test it first.
{
"name": "refund_agent",
"description": "Refunds and cancellations of paid orders. Does not answer delivery questions.",
"input": { "order_id": "string", "reason": "string" }
}
Parallel fan-out and shared state
Independent subtasks go to several agents at once, and a final step reads a shared state map. The rule: each agent writes only its own key. Two agents editing one field is a race, just slower and more expensive.
Hierarchy: a task cut into subtasks
A top agent splits a big task and runs sub-agents with their own context; each returns a short summary. Context stays clean, but the top agent only sees the summary. Trees deeper than two levels are rare because errors multiply.
Generator and critic
One agent drafts, another checks against criteria and returns remarks, the loop repeats. It works only with checkable criteria and an iteration limit; parts of the check are better done by code.
Human in the loop
A human sits where an error is expensive: the agent drafts a refund and waits for approval. The point is chosen by cost of error, not by model confidence, and the pending action is stored in a database so it survives a restart.
When one agent is enough
Start with one agent and good tools; add a coordinator when roles get mixed; add parallelism and hierarchy when latency or context size hurts.
In short
- Split when one agent mixes roles and tools.
- Pipeline: order fixed in code. Coordinator: descriptions as boundaries, choice logged.
- Fan-out: each agent owns one key of shared state.
- Hierarchy: no deeper than two levels.
- Generator and critic: checkable criteria and an iteration limit.
- Human in the loop where money or data move; the draft lives in a database.
What to read next
- Agent evaluation — checking that the system picks the right path.
- AI agent security — rights per agent and where to put a human.