Ask what an agent is and you get either a definition that fits on a sticker or one that fills a whiteboard. The sticker version is close enough to build on: an agent is a language model plus a harness.
The model is the part you don’t build. It takes text and returns text, it remembers nothing between calls, and on its own it cannot read a file, send a message, or check whether it was right. The previous section is about what it can and can’t do. Everything else people mean when they say “agent” lives in the harness.
The harness is the code around the model. It decides what the model sees on each turn, what it’s allowed to do, when it runs again, when it stops, what it leaves behind, and who’s allowed to interrupt. Swap the model and you get a better or worse agent. Swap the harness and you get a different one. That’s why two products calling the same API can feel nothing alike.
The harness, part by part
The harness isn’t one thing, and the useful move is to stop treating it as one. It has parts, seven of them, and each is a chapter in this section.
The order runs from the inside out. The first three are the engine, the machinery of a single turn. The rest widen: how a human gets in, what comes out, where the whole thing lives, and how many of them there are.
Loop
The loop is what turns a model into an agent. Call the model, let it act, feed the result back, call it again. Take the loop away and you have a very good autocomplete. The interesting part isn’t that it repeats, it’s how it decides to stop, which is the half most write-ups skip.
Tools
Tools are how the model reaches anything outside its own output. You hand it a function and a description of when to use it, and that’s where the agent gets its arms and legs. It’s also where most of the failure modes people blame on the model actually live.
Context
Context is everything the model can see on a given turn: your instructions, whatever you retrieved, the results of the tools it just called, and whatever you carried over from last time. Memory belongs here too. Persistence is mostly recall with a clock on it, so it rides on context rather than standing beside it.
Steering
Steering is how a human gets into a loop that’s already running. Three moments: the prompt that starts it, the correction you throw in halfway, and the gates where the agent stops and asks. Turn steering all the way down and you get speed with nobody checking the work. Turn it all the way up and you have an expensive way to do the task yourself.
Artifacts
Artifacts are what the agent leaves behind. Files, diffs, a pull request, a document, a note to itself for next time. Easy to treat as an afterthought, and it isn’t one: an agent that has to produce a reviewable diff needs a different loop than one that answers in a sentence, so the shape of the output quietly decides a lot of what sits upstream of it.
Surfaces
The surface is where the agent operates. A terminal, an editor, a chat window, a phone call, a queue with nobody watching. This is the part of the definition people leave out, and it does more work than any other. Pick the surface and it settles most of the harness for you: the latency you can afford, whether a human is even present, what an artifact is allowed to look like, how far the agent gets before anyone notices. A harness isn’t generic. It’s fitted to a surface.
Orchestration
The last question is how many agents are in the room. Several buy parallelism and more total context, and charge you in coordination, tokens, and information dropped at every handoff. Two shapes actually work, the swarm and the handoff, and the honest default is still one.
The parts lean on each other
Seven parts, but not seven independent knobs. Pull one and the others move.
- Context and tools are the same question stretched over time. A retrieval tool is how context gets filled, and memory is that tool with a clock attached.
- Orchestration rides on the loop and context. A subagent is mostly a way to rent a fresh window and a second loop. You reach for it when one of those ran out of room.
- Steering rides on loop depth. The longer an agent runs unattended, the more a bad turn costs, so turning autonomy up raises the stakes on everything else.
- The surface constrains all of it. A 400 millisecond budget rules out a deep loop, extra tool round trips, and a second agent, without you deciding any of that on purpose.
Worth holding onto while you read the rest. When a decision feels stuck, it’s often because something upstream already settled it.
What isn’t a part
Plenty of things get called an agent design decision and then go looking for a knob that isn’t there. Each one misses in one of three ways.
- A sub-choice. Whether to retrieve, stuff, or summarize is a real decision, but you make it inside context, not beside it. It’s how a tool fills the window.
- A lever. Streaming or batch doesn’t change what the agent does, only when you watch it happen. It buys perceived latency and rides on the budget.
- A layer down. Structured or free-form output is a genuine fork, but it’s the shape of a single response, one level below the loop, and it has its own chapter back in the model section.
None of these are unimportant. They just aren’t parts of this machine, and treating them as if they were is how architecture arguments run in circles.
One entry is an honest near miss. Which model runs each turn, and whether you route between a cheap one and an expensive one, is genuinely independent, and it’s the single biggest lever on cost. By rights it could be an eighth part. It sits with the production chapters instead, because it lives in the infrastructure rather than in the shape of the harness. If you ever redraw this, that’s the first line to move.
Using this as a router
Treat the seven parts as an index. When a design decision has you stuck, find which part it belongs to. If it’s one of the seven, you’re at a real fork and that chapter is where the argument lives. If it isn’t, follow it back to the part it rides on, the layer it sits below, or the section it belongs to.
The last chapter here closes the loop. Once you know the parts, the five dials in Tradeoffs are the summary: which way to turn each one, and why a coding agent and a voice agent end up on opposite sides of all of them.