2.10 THE AGENT

Tradeoffs

The five dials every agent design turns, and why a coding agent and a voice agent end up on opposite sides of all of them.

AUTHOR LEA DANG WORDS 1,660 READ 7 MIN SECTIONS 5

There is no settled blueprint for an agent. Ask five teams how to build one and you will get five architectures, all of them defensible, because each team was quietly optimizing for something different.

Underneath the disagreements, though, most of the choices are the same trade wearing different clothes. You spend model work to buy capability, and you pay for it in latency, cost, and how predictable the result is. A deeper loop solves harder problems and runs longer. More tools reach more of the world and add more round trips. Memory makes an agent feel personal and slowly fills up with stale junk. More agents work in parallel and then lose the plot between them.

This chapter is the map: five dials, what each one costs, and which way two very different systems turn them. The running example is the coding agent, which turns almost everything toward capability. Its foil is a real-time voice agent, which cannot afford to.

The five dials

Five decisions account for most of what separates one agent from another. Here they are on a single axis, running from the fast and cheap end on the left to the capable and expensive end on the right.

FIG_001[ THE FIVE DIALS ]
FAST · CHEAPCAPABLE · FLEXIBLE1 CONTROL FLOWsingle-shotdeep loop2 INFORMATIONcontexttools3 MEMORYstatelesspersistent4 TOPOLOGYone agentmany agents5 OVERSIGHTautonomousoversightCODING AGENTVOICE AGENT
  • 1 CONTROL FLOW · Voice · Single-shot

    A voice agent cannot afford a deep loop. With 300 to 500 milliseconds to start talking, there is no time to feed a result back and go again, so it runs close to a single shot and keeps whatever looping it does shallow and bounded.

  • 1 CONTROL FLOW · Coding · Deep loop

    Coding agents loop, and they loop hard. One request can run for dozens of steps: read a file, run the tests, watch them fail, edit, run them again. The loop is the whole product, and the only thing between it and infinity is a step limit.

  • 2 INFORMATION · Voice · Context

    The budget eats the tool loop. One round trip can cost more than the entire reply window, so a voice agent front-loads everything it can into context and skips the calls it would otherwise love. When a fetch is unavoidable, a small fast model covers the gap with a human “let me check that” while the real work happens out of earshot.

  • 2 INFORMATION · Coding · Tools

    Coding agents live on tools. No repository fits in a context window, so the agent greps, lists, and reads on demand, pulling in only the handful of files a task actually touches.

  • 3 MEMORY · Voice · Stateless

    A voice agent keeps memory thin. Every recall is one more lookup the latency budget cannot spare, so it leans on what already fits in the window and reaches for stored history only when it truly has to.

  • 3 MEMORY · Coding · Thin memory layer

    Coding agents stay mostly stateless per task, with a thin layer of memory on top: a project file like CLAUDE.md read in at the start, plus whatever the session has scribbled down. Enough to know your conventions, not so much that runs stop being reproducible.

  • 4 TOPOLOGY · Voice · One agent

    One voice, one loop. Every handoff between agents costs a round trip the conversation would hear as a pause, so a voice agent stays a single coherent thread instead of a committee.

  • 4 TOPOLOGY · Coding · Many agents

    Coding agents are going multi-agent, but carefully. A main loop spawns short-lived subagents for searches or self-contained chores, and each reports back a summary instead of its whole transcript. The default stays single-agent, because one shared context is far easier to keep honest.

  • 5 OVERSIGHT · Voice · Autonomous by force

    Oversight flips into guardrails. A voice agent cannot freeze mid-sentence to ask permission, so it has to act on its own. Safety moves somewhere the conversation will not feel it: a narrow, mostly reversible set of actions with hard limits, instead of a confirmation prompt.

  • 5 OVERSIGHT · Coding · Oversight as a setting

    Coding agents make oversight a setting, because their actions run from harmless to catastrophic. Reading is automatic, edits might apply on their own or wait for review, and shell commands usually ask first. Plan mode is the cleanest version: the agent proposes, you approve, then it runs.

Hover a point to read where that system lands, and why. Filled is the coding agent, open is the voice agent.

The rest of this section takes each dial in principle. Where the two examples actually land on it, and why, is a hover away on the figure above. Oversight is the odd one out, and we will get to why.

Single-shot or loop

The first decision is whether the model runs once or keeps going. A single-shot call takes a prompt and hands back an answer. It is fast, bounded, easy to test, and completely unable to react to anything it did not already know. A loop feeds the model its own results and lets it pick the next move, over and over, until the job is done. That is what makes open-ended work possible, and it is also why an agent’s runtime and bill have no natural ceiling.

Sitting underneath is a bigger question: how much of the control flow lives in your code, where it is predictable, and how much lives in the model’s head, where it is flexible.

Context or tools

The second decision is how the agent learns what it needs. You can pack everything into the context window up front, every instruction and example and document, where it is instant but frozen at request time and capped by the size of the window. Or you hand the agent tools and let it fetch what it needs, which scales past any window and stays current, at the price of a round trip per call and the occasional confidently wrong choice of tool.

Tools are also the reason loops exist. A tool result is the thing the next turn reacts to, so the moment you add tools you have basically signed up for a loop.

Stateless or memory

The third decision is whether anything survives the turn. A stateless agent wakes up new every time, which is wonderful: it scales, it caches, and the same input always produces the same run, so you can actually test it. A persistent agent remembers across turns and sessions, which feels like magic right up until the memory goes stale, the window bloats, and the behavior starts depending on history you cannot see in the request.

The quiet truth is that memory is usually just a tool. A recall call, a vector search, something that pulls the past back into the window. So this dial mostly rides on the previous one. It is context or tools, stretched across time.

One agent or many

The fourth decision is how many agents are in the room. One agent has one context and one loop, so it stays coherent and you can actually follow what it did. Many agents split the work, a planner here, specialists there, a synthesizer at the end, which buys parallelism and a lot more total context. The bill comes due as coordination: every handoff drops information, the token cost multiplies, and failures compound across the seams.

Squint and a subagent is mostly a way to rent a fresh context window and a second loop. So this dial sits on top of the first two. You reach for it when one window or one sequential loop has run out of room.

Autonomy or oversight

The last decision is how much the agent does without asking. Full autonomy is fast and scales to a volume no human could babysit, but nobody is checking the work. Oversight, whether that means confirm-before-acting or plan-then-approve, catches mistakes before they land, at the cost of throughput and a human who has to be present.

This is the odd dial out, the one whose expensive side buys safety rather than capability. Set it by blast radius. Reading a file can happen on its own. Dropping a production table should not.

The dials are coupled

It is tempting to treat these as five independent knobs. They are not. Turn one and the others move.

  • Memory rides on tools. Persistent memory is a retrieval tool, so dial three is really dial two with a clock attached.
  • Many agents ride on loops and context. A subagent is a fresh window plus a parallel loop, so dial four is dials one and two at the level of the whole system.
  • Autonomy rides on loop depth. The longer an agent runs unattended, the more it does without you, so turning autonomy up quietly raises the stakes on dial one going wrong.

You are not picking five numbers off a shelf. You are choosing a shape, and the shape has its own logic.

Two that aren’t dials

Two more tradeoffs come up often enough that leaving them off the list needs a reason. In both cases the reason is the same: neither is a dial of its own.

  • Streaming or batch? Wait for the whole answer, or start emitting the instant the first token lands? A real choice, but not a new shape. It is a lever on the latency budget, the master dial, and the only reason a voice agent can hit its number at all: it starts talking before the sentence it is saying exists. Less a sixth dial than the trick behind the first one.
  • Structured or free-form? Pin the model to a schema and get back something you can parse, or let it write freely and keep the range? A genuine fork, reliability against expressiveness, and big enough to earn a chapter of its own. It gets one. Here you only need to know the fork is there.
FIG_002[ NOT A DIAL ]
STREAMING / BATCHLATENCY BUDGET · THE MASTER DIALSTREAMBATCHRIDES ON THE LATENCY DIALSTRUCTURED / FREE-FORMCONTROL FLOWthe five dialsONE LAYER DOWNOUTPUT FORMATSTRUCTURED ↔ FREE-FORMHAS ITS OWN CHAPTER
Two real tradeoffs, neither a dial: streaming is a knob on the latency dial, structured output lives a layer down in its own chapter.

Choosing a side

None of this is really about taste. Your constraints make most of the calls for you, and the job is mostly being honest about which constraint is actually binding. Four questions get you most of the way.

  • What is the latency budget? The master dial. A hard real-time ceiling, like voice or autocomplete or anything in the hot path of a request, forces the cheap side everywhere. Minutes of headroom, like a coding task or an overnight research run, let you spend freely.
  • How open-ended is the task? If it is bounded and well specified, a single shot or a fixed pipeline will do. If it has to be discovered as you go, you need a loop.
  • Where is the cost ceiling? It caps loop depth and agent count directly, because every extra step and every extra agent is more tokens on the bill.
  • How big is the blast radius? Reversible and read-only can run on its own. Destructive, expensive, or irreversible wants a human in the loop.

The worked example: real-time voice

To watch the four questions actually decide something, run them on the case that breaks a coding agent’s playbook: an agent that has to talk back in real time. Everything a coding agent does, a human-like voice agent does backwards, and the very first question explains the whole inversion. To feel like a conversation instead of a transaction, the agent has somewhere around 300 to 500 milliseconds to start talking. Miss that window and the person on the other end assumes the call dropped. That budget is not a preference. It is closer to physics, and it forces the cheap side of every dial.

FIG_003[ CODING vs VOICE ]
CODING AGENTVOICE AGENTBUDGET: MINUTESBUDGET: ~400 MScontrol flowDEEP LOOPSINGLE-SHOTinformationTOOLSCONTEXTmemoryPERSISTENTROLLINGtopologyMANY AGENTSONE AGENToversightTUNABLEAUTONOMOUS
One constraint, opposite builds. The latency budget sets everything downstream.

Answer the first question with a hard real-time ceiling and it makes the rest of the calls for you. An open-ended loop would blow the budget, so the loop stays shallow. Extra tool round trips and a second agent both cost time it does not have. And there is no room to stop and ask a human. That single number puts the voice agent on the opposite side of every dial from the coding agent. Hover the open points on the first figure for the reason on each one.