How Coding Agents Actually Work: Inside the Agent Harness
A technical look inside coding agent harnesses, covering context management, repository search, tools, sandboxes, subagents, durable sessions, recovery, concurrency, completion and evaluation.
On this page
Suppose I give a coding agent this task:
Investigate why payment retries occasionally create duplicate charges and prepare a fix.
The obvious thing to focus on is the model. Which model is running, how good it is at code, how much reasoning it can do, and whether it has seen enough Java during training.
That matters, but it explains only part of what happens next.
One coding system may give the model a repository description and ask it to produce a patch. Another may let the same model search the repository, inspect symbols, run the failing test, read the payment client, modify two files, run the tests again, delegate a separate investigation to another agent, recover after the sandbox restarts, compact old context when the task gets long, and refuse to finish until the duplicate charge reproduction stops failing.
Those systems can behave very differently even with the same model.
The difference is the harness around it.
OpenAI now uses this separation directly in the Agents API. The harness runs the model and tool loop and maintains the agent session, while the execution environment is where commands run and files change. Anthropic has reached a similar conclusion through its work on Claude Code and long running coding agents, where context management, sandboxing, structured handoffs, planning, evaluation, and session continuity materially affect what the model can accomplish.
When I think about a coding agent now, I do not picture a model with a shell attached to it. I picture a runtime that continuously decides what the model should see, what it is allowed to do, where that work should execute, what state needs to survive, and whether the task has actually reached a valid end state.
That runtime is the agent harness.
What the harness is doing while the model is coding#
For the payment retry bug, the first model call probably should not contain the entire repository. The harness may start with the task, repository instructions, available tools, current branch, a small repository map, relevant skills, and anything already known about the incident.
The model might respond that it wants to search for retry handling.
The harness executes the search and sends the useful result back.
The model then reads PaymentClient.java, follows a call into RetryPolicy.java, searches for idempotency key generation, and eventually asks to run a test that reproduces the duplicate charge.
At that point the model has not simply answered a prompt. It has participated in a loop.
A useful harness does much more than forward model output into a function call. It interprets the requested action, checks whether that action is currently allowed, routes it to the right environment, records what happened, decides what part of the result belongs in the next model context, and updates the durable state of the run.
This is why OpenAI describes the harness as the control plane around model directed work, while the sandbox provides the execution plane with files, commands, packages, processes, and stateful workspace. Keeping those boundaries separate means the runtime can retain session state, approvals, audit information, and recovery state even if the sandbox itself is restarted or replaced.
For a small coding assistant this may sound like unnecessary machinery. Once the task can run for an hour, call external tools, modify a repository, spawn subagents, pause for review, and recover from failures, the machinery stops being optional.
A coding agent is an iterative runtime around a model. The harness builds context, interprets model intent, enforces policy, coordinates tools and subagents, persists state, recovers from failures and decides when the task is actually complete.
I would make the run state explicit#
One improvement I would make early in a serious harness is to stop treating the agent as either running or finished.
There are several states where the system may be waiting for completely different things.
CREATED
|
v
PREPARING_CONTEXT
|
v
WAITING_FOR_MODEL
|
v
EVALUATING_INTENT
|
+------------------------------+
| |
v v
WAITING_FOR_TOOL WAITING_FOR_APPROVAL
| |
v |
RUNNING_TOOL |
| |
+---------------+--------------+
|
v
EVALUATING_PROGRESS
|
+----------+-----------+
| | |
v v v
WAITING_FOR COMPACTING RECOVERING
MODEL
\ | /
\ | /
+--------+---------+
|
+----+----+
| |
v v
COMPLETED FAILED
This is not a protocol I would force every coding agent to copy. The useful part is making lifecycle state observable.
If a user asks why nothing is happening, I want to know whether the agent is waiting for the model, waiting for a shell command, waiting for approval, compacting context, or recovering from a broken executor connection.
If a process restarts, I want the durable session to tell me what work was in flight rather than reconstructing it from the last chat message.
OpenAI's current Agents API makes a similar distinction between sessions, turns, saved items, events, and execution environments. Sessions can continue over time while turns execute asynchronously, and a session can outlive the environment connected to it.
That is closer to a workflow runtime than a chat loop, which is how I would design it.
A production coding agent moves through model calls, tool execution, approvals, compaction, subagent waits and recovery. Persisting those states separately from the live model context makes long running work resumable and observable.
Context is being rebuilt all the time#
The model can only reason over what the harness puts into its current context, so context construction becomes one of the most important parts of the coding system.
At the beginning of the payment retry task, discovery matters. The model may need repository instructions, a package map, search tools, and the issue description.
Twenty minutes later, after it has found the bug, that context is different. Now the useful information may be:
task objective
PaymentClient.java
RetryPolicy.java
idempotency implementation
relevant tests
current patch
latest test failure
repository coding rules
The model probably does not need every search result from the earlier investigation, the full repository tree, several old test logs, or all the hypotheses it has already disproved.
Anthropic describes context engineering in almost exactly this way. Context is finite, and the goal is not to maximize the amount of information available to the model. The goal is to keep the smallest useful set of high signal information for the current inference. Their guidance for longer agent tasks combines just in time retrieval, compaction, structured notes, and isolated subagent contexts rather than continuously appending everything the system has observed.
I would make every important context item carry some provenance as well.
ContextItem {
source
source_version
trust_level
created_at
freshness
token_cost
relevance_reason
}
If the model sees a failing test, I want to know whether that result came from the current workspace or from a run before the last code edit.
If it sees repository instructions, I want to know which version was loaded.
If it reads a web page or issue comment, that content has a different trust level from source code in the repository.
If a summary from a previous session becomes stale after another agent changes the workspace, the context builder should have enough information to notice.
This does not require a complicated provenance engine for every token. Even basic source, version, timestamp, and trust information makes context easier to reason about when runs become long.
Repository navigation is part of the agent's effective intelligence#
For the duplicate charge bug, the model cannot reason correctly about PaymentClient if the harness never helps it find PaymentClient.
That sounds obvious, but it changes how I evaluate coding agents. Repository retrieval is not just an optimization around the model. It determines which evidence the model gets to reason over.
I would use the structure already present in code before reaching for generic semantic retrieval everywhere.
Paths matter.
Symbols matter.
References matter.
Imports matter.
Git history sometimes matters.
Compiler errors matter.
Tests matter.
A useful search layer might combine:
symbol lookup
text search
reference search
file tree
git diff
recent commits
language server information
with semantic retrieval where it genuinely improves discovery.
For our payment bug, the harness might first search for retry policies, then follow symbols into payment submission, then inspect where the idempotency key is created. That path contains more useful structure than embedding every file and asking for the nearest vectors to “duplicate charge.”
The retrieval system should help the model narrow the repository progressively rather than making the model pay context tokens for thousands of lines it never needed.
Tools are part of the reasoning interface#
Once the model knows what it wants to do, it still does not directly touch the repository. It sees a set of tools that the harness exposes.
A basic coding harness may have:
read_file(path)
search_code(query)
list_directory(path)
apply_patch(diff)
run_command(command)
git_diff()
The quality of these tools matters because they define the action space the model has to reason over.
A test tool that gives the model 60,000 lines of raw stdout is technically correct and operationally poor. A better interface can retain the full log as an artifact while returning the information the next model turn is likely to need.
TestResult {
command
exit_code
passed
failed
failed_tests[]
failure_summary
artifact_location
duration
}
Repository search benefits from the same treatment.
SearchResult {
symbol
file
line_range
match_type
excerpt
}
Anthropic's tool design guidance makes this point directly. Tools form a contract between deterministic software and a nondeterministic agent, so names, descriptions, response shape, context quality, and token efficiency all affect model performance.
OpenAI is also moving more orchestration into the harness. Programmatic Tool Calling lets a model write a small program that can loop, branch, invoke tools in parallel, and process intermediate results before returning only the useful result to the model conversation. That becomes useful when the model needs to inspect many similar items and does not benefit from another inference after every low level operation.
For example, if the agent needs to inspect forty call sites for one API, I would rather let a tool program search them, filter the relevant ones, and return a concise result than create forty alternating model and file read turns.
The model should spend its expensive reasoning on the part that actually requires judgment.
The repository and workspace become memory outside the context window#
Once the payment task has been running for a while, useful state exists in several places.
Model Context
current working memory
Session State
orchestration and progress
Workspace
task artifacts and executable state
Repository
product state
I would not collapse these into one thing.
The model context is temporary and curated for the next inference.
The session should remember things such as the task, current run state, tool operations, approvals, subagents, and checkpoints.
The workspace contains the files, test artifacts, temporary scripts, generated reports, and processes that the agent is using to do the work.
The repository contains the actual software state we ultimately care about.
OpenAI's sandbox architecture treats the environment as a stateful workspace where agents can manipulate files, run commands, install packages, expose ports, and resume work, while the managed harness keeps session information outside that compute boundary.
Anthropic's long running coding work uses the same broad idea from another direction. Their agents leave structured artifacts, progress information, and working repository state so a later context or later agent can continue without having to reconstruct the whole task from conversation history.
This is why increasing the model context window does not remove the need for persistent task state.
A two hour coding task produces too much state, and much of that state is better represented as files, commits, test artifacts, or structured run metadata anyway.
Long running coding agents spread state across several layers. The context window contains only the working set for the next inference, while session state, workspace artifacts and repository state survive compaction, context reset and process restart.
Compaction is useful, but I would not depend on it for everything#
As a session gets longer, eventually the current conversation cannot keep growing.
OpenAI's managed harness automatically compacts earlier context as sessions approach their context limit. Anthropic also uses compaction as one technique for long horizon work, but its long running coding experiments found that structured handoffs and clean context resets can sometimes work better than carrying an increasingly compressed history forever.
For the payment retry task, a useful handoff might preserve:
Objective
Duplicate charge occurs when retry creates
a new idempotency identifier.
Evidence
PaymentClient.java:184
RetryPolicy.java:72
Current changes
PaymentClient.java
RetryPolicyTest.java
Verification
unit tests pass
integration reproduction still failing
Next step
inspect retry path when gateway timeout occurs
A fresh context can understand that quickly.
It does not need every earlier failed grep query and every speculative hypothesis.
I would use compaction when conversational continuity is still useful, and a clean context plus structured handoff when old history is becoming more distracting than helpful. The harness should be able to do either because the durable state of the task lives outside the context window.
Recovery should be designed before the first failure happens#
Now suppose the agent asks the harness to run the integration tests.
The command starts in the sandbox.
The executor connection drops.
The tests finish anyway.
When the harness reconnects, blindly rerunning the command may only waste five minutes. That is annoying, but probably harmless.
The same recovery behavior around:
create_pull_request
or:
publish_package
can create duplicate external state.
I would therefore treat important tool execution as durable work with its own identity.
ToolOperation {
operation_id
run_id
tool
arguments_hash
state
started_at
completed_at
external_reference
}
and give it lifecycle states that are explicit enough to reconcile after failure.
REQUESTED
|
DISPATCHED
|
RUNNING
|
+---- SUCCEEDED
|
+---- FAILED
|
+---- UNKNOWN
UNKNOWN matters because distributed systems occasionally lose the answer to whether something happened.
If the harness knows a pull request operation was dispatched but lost the result, the recovery path should first check whether the pull request already exists. It should not simply repeat the operation because the model asked again.
OpenAI's self hosted environment design includes executor reconnection and explicitly separates session lifecycle from environment lifecycle. Their documentation also warns applications not to create duplicate environments through repeated or concurrent lifecycle requests.
The general lesson is familiar. Once agent tools can perform durable external work, tool invocation starts looking like distributed job execution, and the harness needs operation identity, reconciliation, and idempotency where the external system supports it.
Once coding agent tools create durable external effects, execution becomes distributed work. Persist operation identity and state so the harness can reconcile uncertain outcomes before retrying anything that may already have happened.
The sandbox should remain an execution boundary#
A coding agent needs enough capability to be useful. It may run arbitrary repository code, install dependencies, start servers, execute tests, and modify files.
That is exactly why I would not put every control plane secret and responsibility inside the same environment.
The harness can keep session identity, authorization, approval state, audit logs, recovery information, and other trusted state outside the sandbox, while the sandbox receives only the files, credentials, mounts, and network access required for the task.
OpenAI explicitly recommends this split in its sandbox architecture. Anthropic's Claude Code sandboxing work makes the same operational argument from a security angle, using filesystem and network boundaries so the agent can work more freely inside an approved area without receiving unrestricted access to the host machine.
For the payment task, allowing the agent to change a branch inside /workspace/payments does not imply it needs the developer's SSH directory, production database credentials, or unrestricted internal network access.
The harness should know that difference even if the model does not.
Permissions also become easier to reason about at the tool boundary. A normal file read may proceed automatically. A destructive git operation, production deployment, or external write may cross a stronger policy boundary.
Anthropic's recent work on approval automation is a useful reminder that asking a human every time is not enough. Their telemetry showed users accepted around 93 percent of Claude Code permission prompts, which creates approval fatigue and pushes the system toward deterministic containment plus more selective review.
I would rather make ordinary work safe by construction than interrupt the user every thirty seconds and assume attention will remain perfect.
Subagents are useful because context and ownership can be separated#
Suppose our duplicate charge investigation splits naturally into three areas.
One branch needs to inspect recent deployment changes.
Another needs to understand the payment retry code.
A third needs to inspect database evidence and gateway responses.
The root agent can do all of that itself, but then one context accumulates every search, dead end, file read, and tool result from all three investigations.
Another option is:
Root Agent
|
+---------------+---------------+
| | |
v v v
Deployment Retry Code Data
Analysis Analysis Analysis
| | |
+---------------+---------------+
|
v
Distilled Findings
Each subagent gets a cleaner working context, can explore deeply, and returns only the useful findings.
Anthropic describes subagents in almost exactly these terms. They are useful when independent work benefits from context isolation and can return condensed results to the lead agent. OpenAI's Agents API similarly gives subagents independent contexts while the root agent coordinates their work.
Where I become much more careful is concurrent modification.
Three agents reading the same repository is easy.
Three agents writing the same repository is a concurrency problem.
The simplest policy is one writer and many readers. The root agent or one designated implementation agent owns mutation while other agents investigate and propose changes.
For more parallel implementation, the harness needs an explicit ownership model. Depending on the repository and task, that could mean separate git branches, worktrees, file ownership, patch proposals, or some other merge protocol.
Subagent A
|
v
Branch A
Subagent B
|
v
Branch B
\ /
\ /
v v
Merge Owner
|
v
Shared Branch
I would not let several agents mutate one workspace concurrently and hope the model figures out the conflicts afterward. At that point the harness is managing shared mutable state, and normal concurrency rules still apply.
Planning, implementation, and evaluation do not always belong in one context#
For a small bug like the duplicate retry issue, one agent can probably investigate, edit, and verify the fix.
For a multi hour implementation, I would consider separating some of these responsibilities.
Anthropic's recent harness work used planner, generator, and evaluator agents for larger application development tasks. One reason was that an agent evaluating its own work can be too forgiving, while a separate evaluator can inspect the result with a different context and grading objective.
A larger coding workflow might therefore look like:
Planner
|
v
Task Graph
|
v
Implementation
|
v
Deterministic Verification
|
v
Evaluator
|
+---- accepted
|
+---- revision required
I would still put deterministic verification before model judgment whenever possible.
For the payment fix, software can verify:
build succeeds
payment unit tests pass
retry integration test passes
expected files changed
no unexpected generated files
A model evaluator becomes useful for questions such as whether the patch actually addresses the intended failure mode, whether the change is unnecessarily broad, or whether the implementation fits the architecture of the repository.
The harness should not ask a model to judge something the compiler already knows.
Completion should be a system condition, not a sentence from the model#
Coding agents often fail in a very ordinary way. They make a plausible edit and declare the task complete.
For our payment retry issue, that is not enough.
The harness may know that the task requires:
CompletionContract {
required_checks[]
required_artifacts[]
prohibited_states[]
human_review_required
}
and the specific contract could require:
PaymentClient compiles
payment unit tests pass
duplicate charge reproduction passes
retry integration tests pass
no unrelated files modified
diff exists
The model can still tell the harness that it believes the task is ready.
The harness then checks whether the conditions that define ready have actually been satisfied.
Anthropic observed this failure mode in its long running coding experiments. Agents would often make code changes and perform partial verification but mark features complete before checking the actual end to end behavior. Explicit testing instructions and access to realistic verification tools improved the result.
I would make completion contracts task specific where the product knows what success means, rather than relying only on a generic prompt that says “run tests before finishing.”
Skills, MCP, hooks, repository instructions, and subagents all modify the harness#
Modern coding systems expose many extension mechanisms, and it is easy to treat each one as a separate architectural idea.
I find it simpler to ask what part of the harness the extension changes.
Repository instructions change the context.
Skills change context and procedure.
MCP adds capabilities and external data.
Hooks add deterministic control around lifecycle events.
Subagents create additional contexts and execution branches.
Sandbox configuration changes the execution boundary.
The practical implication is that harness configuration needs versioning just like model configuration.
If the same model suddenly behaves differently, the cause may be:
prompt version
repository instructions
skill version
tool schema
context builder
reasoning effort
compaction logic
sandbox image
permission policy
model version
Anthropic's April 2026 Claude Code postmortem is a good example of why this matters. Reports that looked like model quality degradation were traced to multiple product changes while Anthropic said the underlying API and inference layer were unaffected, including a change in the default reasoning effort.
When I evaluate a coding agent release, I therefore want a harness_version beside model_version.
Observability should let me reconstruct the work#
For a normal API request, latency and status may be enough.
For a coding agent, I want enough information to reconstruct why the run behaved the way it did.
A useful trace might include:
run_id
turn_id
model_version
harness_version
context_builder_version
context_tokens
compaction_version
tool_call_id
tool
sandbox_id
workspace_snapshot
files_read[]
files_changed[]
subagent_id
approval_id
completion_checks[]
latency
cost
I would also record the important transitions in the run state machine.
That makes it possible to answer questions such as:
Why did this task spend twelve minutes waiting?
Why did the agent inspect eighty files?
Why did it rerun the same test?
Did the context reset before the wrong edit?
Was the change made by the root agent or a subagent?
Did the harness recover an existing operation or execute it twice?
OpenAI's current observability surface exposes session history, turns, tool calls, subagent activity, usage, and exportable traces, which reflects the same general need.
The trace is also the raw material for evaluation.
I would evaluate the harness along with the patch#
Imagine two coding agents both eventually produce the same correct payment fix.
The first does this:
search retry symbol
read 4 files
run reproduction
change 2 files
run targeted tests
run full tests
finish
The second does this:
read 70 files
change 8 files
revert 6 files
run the same test 5 times
lose context
rediscover the bug
produce the same final patch
If the benchmark checks only whether the patch passes, both may receive the same score.
From a production perspective they are not equivalent.
I would measure at least:
task success
unsafe actions
tool calls
model turns
context tokens
files read
files changed
test behavior
recovery behavior
subagent fan out
latency
cost
Then I would look at behavior and trajectory alongside the final outcome.
A harness release that preserves task success while doubling model turns, context use, tool retries, and touched files has changed something important.
OpenAI's current guidance for programmatic tool calling recommends evaluating model turns, tool calls, retries, recovery, latency, cost, safety, and final correctness together. The general lesson is that a correct final answer can hide an inefficient or fragile trajectory.
That is also why harness improvements can materially move coding benchmarks without changing the base model. Context construction, tool design, compaction, permissions, subagents, and verification all influence how much of the model's underlying capability becomes useful behavior.
If I were building the harness, I would keep these pieces separate#
I would not literally require one class per component, but conceptually I want the boundaries to exist.
CodingHarness {
session_store
run_state_machine
context_builder
context_provenance
instruction_registry
skill_loader
retrieval_layer
tool_registry
tool_router
policy_engine
operation_store
sandbox_manager
workspace_manager
subagent_manager
concurrency_controller
compaction_manager
recovery_manager
completion_evaluator
trace_store
}
The loop then becomes easier to reason about.
load durable run state
build context for the next decision
invoke model
interpret requested action
check policy
execute or delegate
record operation and observation
update durable state
evaluate progress
compact or checkpoint when needed
continue until completion contract passes
When something goes wrong, the location of the failure becomes much clearer.
If the model cannot find the relevant payment code, I inspect repository retrieval and context.
If it forgets that the duplicate charge only happens after gateway timeouts, I inspect context provenance, handoff state, and compaction.
If it executes an unsafe command, I inspect policy and sandbox boundaries.
If it opens two pull requests after reconnect, I inspect operation identity and recovery.
If three subagents overwrite each other, I inspect concurrency ownership.
If it says the fix is complete before reproducing the incident, I inspect the completion contract.
Calling all of these “the model made a mistake” does not give the engineering team much to work with.
The model is still central, but the harness decides how much of that capability becomes useful#
A better model can reason about harder code, understand larger changes, recover from ambiguity, and make better decisions. None of the harness design removes that.
The point is that the model is operating inside a system.
OpenAI's current Agents API makes the harness a first class product because long running agents need managed context, tools, subagents, session continuity, and execution environments. Anthropic's coding work has repeatedly shown that decomposition, structured state, context management, sandboxing, and separate evaluation can substantially change what the same model accomplishes.
So when I compare coding agents, I would look beyond which model name appears in the settings.
I would want to know how the system finds code, how it decides what reaches context, what happens when the context fills up, how tools return information, where commands execute, how writes are coordinated, what survives a restart, how side effects are reconciled, when subagents are useful, how completion is verified, and what the evaluation actually measures.
That is much closer to understanding how the coding agent works than looking only at the model underneath it.
Keep reading
Continue with a guided sequence of free production engineering essays.
Find your next reading pathRelated reading
Continue this path
Designing Containment for Production AI Agents
Production AI agent containment needs more than a sandbox. It needs separate information, execution and effect boundaries, scoped capabilities, egress controls, quarantine, and independent runtime enforcement.
Read essayDesigning Continuous Evaluation for Production AI Agents
Production AI agent evaluation needs more than task success. It needs behavioral checks, trajectory analysis, outcome verification, failure lineage, repeated trials, statistical release gates, and production failures feeding new evaluation evidence.
Read essayA Skill Is Not a File. It Is a Deploy: Designing Agent Skills Infrastructure
Production Agent Skills are deployments, not files. They need immutable versions, controlled rollout, revocation, trust boundaries, progressive resolution, and a runtime path designed for scale.
Read essay