← Back to insights
AI Product Development•12 min read•

Agent harnesses: how we build the runtime around an LLM

Underline

An agent harness is the runtime around a language model: tools, memory, a loop, permissions, evals and a budget. This is how we build ours.

An agent harness is the runtime around a language model once a demo has to survive contact with a real task. The model proposes the next step. The harness decides which tools exist, what may sit in the window, whether a person has to approve, and what is written down when the turn ends. We build that runtime ourselves, and we run it on our own engineering work. A completion is a call. A harness is the system you can operate.

A turn has a shape

Each turn moves through named phases: understand the ask, ask when it is unclear, choose a way of working, explore with read-only tools, plan, execute, check the result, hold for a person, answer, and then finish, stop, or fail. A move that does not belong on that path is refused. An answer that has already been given stays an answer.

A clean run reads: understand, route, plan, execute, check, answer, done.

One commit lands at the end of a step, before anything the next step can see. The allow-list, the working directory and the isolation choice live on the call and are read again every step. A later step stays tied to what the caller actually sent.

Tool calling

A tool is a capability with a name and a schema. The catalog stays off the model. We expand an allow-list, globs included, against the live catalog, and we cap the native list at 128 tools so the window stays about the task. Everything else stays searchable. The prompt carries a one-line description of each tool that is actually on the list. The longer instructions for a capability stay out of the window until something asks for them.

Control-plane tools stay off the model: the turn runner, the queue, the session store, and the raw completion call. An orchestrator may spawn a child, ask for status, and walk the tree. A leaf stays a leaf, unless it is marked as an orchestrator itself.

When a turn pauses in the middle of a tool loop, we store the assistant's tool calls and the matching results. The next generation is still a valid tool conversation, which is what lets a stop or an approval land between tools instead of inside a batch the runtime can no longer see.

Context is a budget. Memory is a store.

Context assembly is stateless. The caller passes the messages and the model. The assembler opens neither the session nor the memory store. It has one job: return a window that fits. Limits come from the model record when we have one, and from a conservative fallback when we do not. If the count is still over the usable budget after compaction, the provider is not called. Overflow stops the turn. Calling the model on a window that has already dropped the goal is the failure we refused.

The caller writes the compaction back as its own transcript entry, marked so the next assemble can see that earlier turns were summarised. A hook that runs after the window is built may replace the system prompt, or append lines, for that request only. Those lines are not written into the transcript.

The transcript is short-term memory. Long-term memory is a separate store: working notes, observations and embeddings, under a scope string that store does not interpret. After a turn we may add, rewrite or drop a fact, or link two subjects. Before the next generation a hook searches that index with the latest user message. An exact subject ranks first, then a linked subject, then the text. A text-only hit has to clear a similarity bar before it is injected.

Recalled lines are added to the model request only. They are not appended to the transcript, so a bad recall cannot become the history of the session. The hook fails open: a memory miss degrades the turn, and the turn still runs. We keep harness state for the last 20 turns of a session and delete the older turn records.

Planning, and the loop around it

The production scheduler writes down what it understood, and what done means, before it picks a hop. An unclear understanding asks the person and stops there. A clear one names the hop: a direct answer, a single tool action, a frozen checklist, a reactive tool loop, or a read-only investigation. Investigation gathers, then we classify again. A read-only phase, and a child marked read-only, deny writes and command execution. That phase cannot promote itself into a change.

If the classifier is skipped, unconfigured, or returns an error, production mode falls back to the reactive loop, so a missing classifier still does work. A caller that forced decide-then-act does not get that fallback. The turn stops.

A plan the caller forces is an overlay on that scheduler, for when a person wants a checklist opened or a gate in front of it. The model may open a plan and edit the steps. The runtime writes the plan body. Execution then takes one checklist step per turn, so a stop or an approval can land between tools. Each step carries a success check on the tool result: a token, a contains or absent test, an exit code, a JSON check, a combination of those, or a proof command. A sentence the model writes about its own work is ignored. A miss replans that step. A second miss drops the unfinished suffix and keeps the steps that already passed.

Parallel work is a spawn, with a parent and child sessions. A child inherits the parent's loop choice unless it sets its own, and it does not inherit a forced plan. Its remaining turns cannot exceed what the parent has left. Depth defaults to 3, and a parent may start at most five new children in one turn. The model cannot raise either cap. A later turn may spawn again, up to the same cap. A parent with children still running yields instead of declaring itself done. When every child has finished, one line is stitched back and the parent continues. By default the parent waits for every child. It can also wait for the first, for a quorum, or supervise that cohort.

A session that is still waiting on a one-shot wake stays open instead of finishing, until that wake fires or the session is stopped. In that state, done means the wake has not happened yet.

Permissions, gates and isolation

Tools are deny-by-default. An allow-list names what may run, and a deny rule wins after that, including when the allow-list is a wildcard. A standing rule still hides spend and control-plane surfaces under a wildcard: secrets, the approval gate itself, the raw model call, a scoped hold that can freeze work, and the ability to rewrite stored state. Approvals and that hold sit outside the agent's tool list.

Dispatch is three outcomes. Allow runs the tool. Deny blocks it. Anything else holds for a person, and that same call is replayed after the decision. A hold does not enqueue the next step. The approval is bound to the same scope as the session, so a decision from another scope cannot release the hold.

A small set of irreversible actions asks a person even when the loop would otherwise allow them: merging, rebasing, landing a change, and applying, patching or deleting a cluster resource.

Filesystem access is a grant. A classifier scores whether the path is in scope, and the middle of that range opens a gate. Shell and git inside an isolated run are stamped with the sandbox and the granted root. The caller's working directory is not reused as the shell's current directory inside that sandbox.

Hooks sit on the turn: before the turn, before and after generation, before and after a tool, and after the turn. The pre-phases and the post-turn hook fail closed unless the caller says otherwise. A hook that denies a turn about to finish sends it back, up to two retries, and then the turn stops. Structured output is checked the same way. Invalid JSON is sent back. It is not treated as a finished turn.

Isolation is one sandbox for the run, and one sandbox per spawned child, stopped when the turn ends unless we explicitly keep it. A dedicated work copy keeps the checkout stable across a pause and across children.

Evals, proof and a goal

We keep three checks apart, because they fail in different places.

An eval watch is armed on a session with criteria and a turn budget. When a turn ends, an assert runs. A pass stops the watch. A fail continues the session, with the same tools, model and isolation, until the budget is spent, and then the session is stopped. The default assert is a substring of the turn result. A caller can supply its own assert when a substring is the wrong test. The watch is how a session gets another turn. It is also how a session is forced to stop.

Proof is a verdict bound to a revision. We run a command and store the result against that revision, and the verdict can block a push, a merge or a land. If that check is absent, the land proceeds. If it is present and the verdict says no, the land is blocked. The person who must approve is still the gate. Proof is the machine check beside that person.

A goal is the durable outcome across several waves of work. The watch covers one session.

What we record

Every turn publishes a start and a completion. The completion carries a ledger: which loop was requested, which handler actually ran, what the front door decided, which tools fired, which were blocked, token counts split by layer (the classifier, the planner, the reactive loop, the executor), an outcome of success, partial, abstain or handed to a person, and the phase path. The ledger is append-only, so the dispatcher and the loop write their own figures and leave each other's alone.

The same payload is written as a receipt when that component is present. Traces correlate on a run id. Logs, traces and metrics can stay in process or leave through OpenTelemetry. That layer records the run, and it stays clear of the product's own vocabulary.

Retries, timeouts and cost

A provider call retries only while nothing real has been forwarded to the caller. The default is two retries, with backoff, and an abort cancels the in-flight call. A stream that goes idle is not retried. The default idle bound is two minutes, and the whole stream is capped at five minutes. Retrying a quiet stream risks paying twice for tokens that may already have been generated.

Output length is clamped to the tighter of the request, the model's own ceiling and a soft cap. The default soft cap is 32,000 output tokens. A request cannot raise the model's maximum. A thinking level of low, medium or high is sent on every generation. Medium is the default. Children inherit the level, a later message on the same turn can change it, and a hold keeps the level it had.

Token use is recorded per layer on the receipt. When the provider omits a cost, we attach one from the model catalog, so a turn can be costed afterwards from the record of the turn itself.

The other budgets are structural. A child cannot outlive the parent's remaining turns. Fan-out and depth are capped where the model cannot edit them. An eval watch stops at its turn budget. A bad output or a denied completion is retried twice, then the turn stops. A plan step is replanned once, then the unfinished tail is dropped. A window that does not fit never reaches the provider.

Where this meets a delivery

We run this on our own engineering work. The same ideas show up, at a smaller grain, when a team is shipping one agent or putting coding agents in front of a group.

A single agent that has to run in production is a Hephaestus Sprint: one system, a golden eval set, monitoring hooks, a runbook, and a decision on whether to scale. A team that already uses coding agents, or is about to, is an AI Coding Rollout: repo context, skills, a security boundary, and a gate in CI before a change lands.

Frequently asked questions

What is an agent harness? It is the runtime around a language model: the loop that calls tools, the window of context, the permissions, and the record of what happened. The model proposes the next step. The harness decides whether that step runs.

Why keep tools behind an allow-list? A long tool list crowds the window and widens what a bad call can touch. We allow a bounded set, hide spend and control-plane tools even from a wildcard, and hold anything unmatched until a person decides.

How does a loop actually stop? A turn has a budget. A parent cannot spawn past the depth and the per-turn fan-out, and the model cannot raise those caps. A failed check is replanned once, then the unfinished tail is dropped. An eval watch continues only until its turn budget, and then the session stops.

When is a coding assistant enough? A coding assistant still needs a boundary in the repo: what it may touch, how a change is proved, and who approves the landing. A product agent that runs with no person on every turn needs the rest of the harness as well: isolation, a gate, a receipt, and an eval that can abort it.

← Back to insights

Let's talk about what you're building

The Aidoni team together outdoors in Gothenburg