Suppose I give a coding agent this task:

Investigate why payment retries occasionally create duplicate charges and prepare a fix.

The obvious thing to focus on is the model. Which model is running, how good it is at code, how much reasoning it can do, and whether it has seen enough Java during training.

That matters, but it explains only part of what happens next.

One coding system may give the model a repository description and ask it to produce a patch. Another may let the same model search the repository, inspect symbols, run the failing test, read the payment client, modify two files, run the tests again, delegate a separate investigation to another agent, recover after the sandbox restarts, compact old context when the task gets long, and refuse to finish until the duplicate charge reproduction stops failing.

Those systems can behave very differently even with the same model.

The difference is the harness around it.

OpenAI now uses this separation directly in the Agents API. The harness runs the model and tool loop and maintains the agent session, while the execution environment is where commands run and files change. Anthropic has reached a similar conclusion through its work on Claude Code and long running coding agents, where context management, sandboxing, structured handoffs, planning, evaluation, and session continuity materially affect what the model can accomplish.

When I think about a coding agent now, I do not picture a model with a shell attached to it. I picture a runtime that continuously decides what the model should see, what it is allowed to do, where that work should execute, what state needs to survive, and whether the task has actually reached a valid end state.

That runtime is the agent harness.

What the harness is doing while the model is coding#

For the payment retry bug, the first model call probably should not contain the entire repository. The harness may start with the task, repository instructions, available tools, current branch, a small repository map, relevant skills, and anything already known about the incident.

The model might respond that it wants to search for retry handling.

The harness executes the search and sends the useful result back.

The model then reads PaymentClient.java, follows a call into RetryPolicy.java, searches for idempotency key generation, and eventually asks to run a test that reproduces the duplicate charge.

At that point the model has not simply answered a prompt. It has participated in a loop.

A useful harness does much more than forward model output into a function call. It interprets the requested action, checks whether that action is currently allowed, routes it to the right environment, records what happened, decides what part of the result belongs in the next model context, and updates the durable state of the run.

This is why OpenAI describes the harness as the control plane around model directed work, while the sandbox provides the execution plane with files, commands, packages, processes, and stateful workspace. Keeping those boundaries separate means the runtime can retain session state, approvals, audit information, and recovery state even if the sandbox itself is restarted or replaced.

For a small coding assistant this may sound like unnecessary machinery. Once the task can run for an hour, call external tools, modify a repository, spawn subagents, pause for review, and recover from failures, the machinery stops being optional.

Architecture diagram showing a coding agent harness around the model. Durable agent state is used by a context builder, which invokes the model. Model intent passes through policy and tool routing into a sandbox containing the repository, shell, build tools and tests. Observations return to durable state and the loop continues until completion, recovery, approval or another lifecycle transition.

A coding agent is an iterative runtime around a model. The harness builds context, interprets model intent, enforces policy, coordinates tools and subagents, persists state, recovers from failures and decides when the task is actually complete.

I would make the run state explicit#

One improvement I would make early in a serious harness is to stop treating the agent as either running or finished.

There are several states where the system may be waiting for completely different things.

CREATED
   |
   v
PREPARING_CONTEXT
   |
   v
WAITING_FOR_MODEL
   |
   v
EVALUATING_INTENT
   |
   +------------------------------+
   |                              |
   v                              v
WAITING_FOR_TOOL            WAITING_FOR_APPROVAL
   |                              |
   v                              |
RUNNING_TOOL                      |
   |                              |
   +---------------+--------------+
                   |
                   v
          EVALUATING_PROGRESS
                   |
        +----------+-----------+
        |          |           |
        v          v           v
WAITING_FOR   COMPACTING   RECOVERING
MODEL
        \          |           /
         \         |          /
          +--------+---------+
                   |
              +----+----+
              |         |
              v         v
          COMPLETED   FAILED

This is not a protocol I would force every coding agent to copy. The useful part is making lifecycle state observable.

If a user asks why nothing is happening, I want to know whether the agent is waiting for the model, waiting for a shell command, waiting for approval, compacting context, or recovering from a broken executor connection.

If a process restarts, I want the durable session to tell me what work was in flight rather than reconstructing it from the last chat message.

OpenAI's current Agents API makes a similar distinction between sessions, turns, saved items, events, and execution environments. Sessions can continue over time while turns execute asynchronously, and a session can outlive the environment connected to it.

That is closer to a workflow runtime than a chat loop, which is how I would design it.

State machine for a long running coding agent showing context preparation, model inference, intent evaluation, tool execution, approvals, progress evaluation, compaction, recovery, completion and failure. Durable session state records meaningful transitions so the coding task can survive process restarts and disconnected clients.

A production coding agent moves through model calls, tool execution, approvals, compaction, subagent waits and recovery. Persisting those states separately from the live model context makes long running work resumable and observable.

Context is being rebuilt all the time#

The model can only reason over what the harness puts into its current context, so context construction becomes one of the most important parts of the coding system.

At the beginning of the payment retry task, discovery matters. The model may need repository instructions, a package map, search tools, and the issue description.

Twenty minutes later, after it has found the bug, that context is different. Now the useful information may be:

task objective

PaymentClient.java

RetryPolicy.java

idempotency implementation

relevant tests

current patch

latest test failure

repository coding rules

The model probably does not need every search result from the earlier investigation, the full repository tree, several old test logs, or all the hypotheses it has already disproved.

Anthropic describes context engineering in almost exactly this way. Context is finite, and the goal is not to maximize the amount of information available to the model. The goal is to keep the smallest useful set of high signal information for the current inference. Their guidance for longer agent tasks combines just in time retrieval, compaction, structured notes, and isolated subagent contexts rather than continuously appending everything the system has observed.

I would make every important context item carry some provenance as well.

ContextItem {
    source
    source_version

    trust_level

    created_at
    freshness

    token_cost

    relevance_reason
}

If the model sees a failing test, I want to know whether that result came from the current workspace or from a run before the last code edit.

If it sees repository instructions, I want to know which version was loaded.

If it reads a web page or issue comment, that content has a different trust level from source code in the repository.

If a summary from a previous session becomes stale after another agent changes the workspace, the context builder should have enough information to notice.

This does not require a complicated provenance engine for every token. Even basic source, version, timestamp, and trust information makes context easier to reason about when runs become long.

Repository navigation is part of the agent's effective intelligence#

For the duplicate charge bug, the model cannot reason correctly about PaymentClient if the harness never helps it find PaymentClient.

That sounds obvious, but it changes how I evaluate coding agents. Repository retrieval is not just an optimization around the model. It determines which evidence the model gets to reason over.

I would use the structure already present in code before reaching for generic semantic retrieval everywhere.

Paths matter.

Symbols matter.

References matter.

Imports matter.

Git history sometimes matters.

Compiler errors matter.

Tests matter.

A useful search layer might combine:

symbol lookup

text search

reference search

file tree

git diff

recent commits

language server information

with semantic retrieval where it genuinely improves discovery.

For our payment bug, the harness might first search for retry policies, then follow symbols into payment submission, then inspect where the idempotency key is created. That path contains more useful structure than embedding every file and asking for the nearest vectors to “duplicate charge.”

The retrieval system should help the model narrow the repository progressively rather than making the model pay context tokens for thousands of lines it never needed.

Tools are part of the reasoning interface#

Once the model knows what it wants to do, it still does not directly touch the repository. It sees a set of tools that the harness exposes.

A basic coding harness may have:

read_file(path)

search_code(query)

list_directory(path)

apply_patch(diff)

run_command(command)

git_diff()

The quality of these tools matters because they define the action space the model has to reason over.

A test tool that gives the model 60,000 lines of raw stdout is technically correct and operationally poor. A better interface can retain the full log as an artifact while returning the information the next model turn is likely to need.

TestResult {
    command

    exit_code

    passed
    failed

    failed_tests[]

    failure_summary

    artifact_location

    duration
}

Repository search benefits from the same treatment.

SearchResult {
    symbol

    file

    line_range

    match_type

    excerpt
}

Anthropic's tool design guidance makes this point directly. Tools form a contract between deterministic software and a nondeterministic agent, so names, descriptions, response shape, context quality, and token efficiency all affect model performance.

OpenAI is also moving more orchestration into the harness. Programmatic Tool Calling lets a model write a small program that can loop, branch, invoke tools in parallel, and process intermediate results before returning only the useful result to the model conversation. That becomes useful when the model needs to inspect many similar items and does not benefit from another inference after every low level operation.

For example, if the agent needs to inspect forty call sites for one API, I would rather let a tool program search them, filter the relevant ones, and return a concise result than create forty alternating model and file read turns.

The model should spend its expensive reasoning on the part that actually requires judgment.

The repository and workspace become memory outside the context window#

Once the payment task has been running for a while, useful state exists in several places.

Model Context
    current working memory

Session State
    orchestration and progress

Workspace
    task artifacts and executable state

Repository
    product state

I would not collapse these into one thing.

The model context is temporary and curated for the next inference.

The session should remember things such as the task, current run state, tool operations, approvals, subagents, and checkpoints.

The workspace contains the files, test artifacts, temporary scripts, generated reports, and processes that the agent is using to do the work.

The repository contains the actual software state we ultimately care about.

OpenAI's sandbox architecture treats the environment as a stateful workspace where agents can manipulate files, run commands, install packages, expose ports, and resume work, while the managed harness keeps session information outside that compute boundary.

Anthropic's long running coding work uses the same broad idea from another direction. Their agents leave structured artifacts, progress information, and working repository state so a later context or later agent can continue without having to reconstruct the whole task from conversation history.

This is why increasing the model context window does not remove the need for persistent task state.

A two hour coding task produces too much state, and much of that state is better represented as files, commits, test artifacts, or structured run metadata anyway.

Diagram showing where state lives in a long running coding agent. The model context contains short lived working memory, durable session state stores orchestration and progress, the workspace holds task artifacts and executable state, and the repository contains durable product state. A context builder selects relevant information with provenance for each model invocation while compaction or context reset leaves durable task state intact.

Long running coding agents spread state across several layers. The context window contains only the working set for the next inference, while session state, workspace artifacts and repository state survive compaction, context reset and process restart.

Compaction is useful, but I would not depend on it for everything#

As a session gets longer, eventually the current conversation cannot keep growing.

OpenAI's managed harness automatically compacts earlier context as sessions approach their context limit. Anthropic also uses compaction as one technique for long horizon work, but its long running coding experiments found that structured handoffs and clean context resets can sometimes work better than carrying an increasingly compressed history forever.

For the payment retry task, a useful handoff might preserve:

Objective

Duplicate charge occurs when retry creates
a new idempotency identifier.

Evidence

PaymentClient.java:184
RetryPolicy.java:72

Current changes

PaymentClient.java
RetryPolicyTest.java

Verification

unit tests pass
integration reproduction still failing

Next step

inspect retry path when gateway timeout occurs

A fresh context can understand that quickly.

It does not need every earlier failed grep query and every speculative hypothesis.

I would use compaction when conversational continuity is still useful, and a clean context plus structured handoff when old history is becoming more distracting than helpful. The harness should be able to do either because the durable state of the task lives outside the context window.

Recovery should be designed before the first failure happens#

Now suppose the agent asks the harness to run the integration tests.

The command starts in the sandbox.

The executor connection drops.

The tests finish anyway.

When the harness reconnects, blindly rerunning the command may only waste five minutes. That is annoying, but probably harmless.

The same recovery behavior around:

create_pull_request

or:

publish_package

can create duplicate external state.

I would therefore treat important tool execution as durable work with its own identity.

ToolOperation {
    operation_id

    run_id

    tool

    arguments_hash

    state

    started_at
    completed_at

    external_reference
}

and give it lifecycle states that are explicit enough to reconcile after failure.

REQUESTED
   |
DISPATCHED
   |
RUNNING
   |
   +---- SUCCEEDED
   |
   +---- FAILED
   |
   +---- UNKNOWN

UNKNOWN matters because distributed systems occasionally lose the answer to whether something happened.

If the harness knows a pull request operation was dispatched but lost the result, the recovery path should first check whether the pull request already exists. It should not simply repeat the operation because the model asked again.

OpenAI's self hosted environment design includes executor reconnection and explicitly separates session lifecycle from environment lifecycle. Their documentation also warns applications not to create duplicate environments through repeated or concurrent lifecycle requests.

The general lesson is familiar. Once agent tools can perform durable external work, tool invocation starts looking like distributed job execution, and the harness needs operation identity, reconciliation, and idempotency where the external system supports it.

Diagram showing durable tool execution for a coding agent. A tool operation moves through requested, dispatched and running states before succeeding, failing, or becoming unknown. In an example, a pull request is created successfully but the harness loses the connection before receiving the result. Recovery reconciles external state and marks the operation successful instead of blindly retrying and creating a duplicate side effect.

Once coding agent tools create durable external effects, execution becomes distributed work. Persist operation identity and state so the harness can reconcile uncertain outcomes before retrying anything that may already have happened.

The sandbox should remain an execution boundary#

A coding agent needs enough capability to be useful. It may run arbitrary repository code, install dependencies, start servers, execute tests, and modify files.

That is exactly why I would not put every control plane secret and responsibility inside the same environment.

The harness can keep session identity, authorization, approval state, audit logs, recovery information, and other trusted state outside the sandbox, while the sandbox receives only the files, credentials, mounts, and network access required for the task.

OpenAI explicitly recommends this split in its sandbox architecture. Anthropic's Claude Code sandboxing work makes the same operational argument from a security angle, using filesystem and network boundaries so the agent can work more freely inside an approved area without receiving unrestricted access to the host machine.

For the payment task, allowing the agent to change a branch inside /workspace/payments does not imply it needs the developer's SSH directory, production database credentials, or unrestricted internal network access.

The harness should know that difference even if the model does not.

Permissions also become easier to reason about at the tool boundary. A normal file read may proceed automatically. A destructive git operation, production deployment, or external write may cross a stronger policy boundary.

Anthropic's recent work on approval automation is a useful reminder that asking a human every time is not enough. Their telemetry showed users accepted around 93 percent of Claude Code permission prompts, which creates approval fatigue and pushes the system toward deterministic containment plus more selective review.

I would rather make ordinary work safe by construction than interrupt the user every thirty seconds and assume attention will remain perfect.

Subagents are useful because context and ownership can be separated#

Suppose our duplicate charge investigation splits naturally into three areas.

One branch needs to inspect recent deployment changes.

Another needs to understand the payment retry code.

A third needs to inspect database evidence and gateway responses.

The root agent can do all of that itself, but then one context accumulates every search, dead end, file read, and tool result from all three investigations.

Another option is:

                    Root Agent
                        |
        +---------------+---------------+
        |               |               |
        v               v               v
   Deployment        Retry Code       Data
    Analysis          Analysis       Analysis
        |               |               |
        +---------------+---------------+
                        |
                        v
                  Distilled Findings

Each subagent gets a cleaner working context, can explore deeply, and returns only the useful findings.

Anthropic describes subagents in almost exactly these terms. They are useful when independent work benefits from context isolation and can return condensed results to the lead agent. OpenAI's Agents API similarly gives subagents independent contexts while the root agent coordinates their work.

Where I become much more careful is concurrent modification.

Three agents reading the same repository is easy.

Three agents writing the same repository is a concurrency problem.

The simplest policy is one writer and many readers. The root agent or one designated implementation agent owns mutation while other agents investigate and propose changes.

For more parallel implementation, the harness needs an explicit ownership model. Depending on the repository and task, that could mean separate git branches, worktrees, file ownership, patch proposals, or some other merge protocol.

Subagent A
    |
    v
Branch A

Subagent B
    |
    v
Branch B

       \       /
        \     /
         v   v
      Merge Owner
          |
          v
     Shared Branch

I would not let several agents mutate one workspace concurrently and hope the model figures out the conflicts afterward. At that point the harness is managing shared mutable state, and normal concurrency rules still apply.

Planning, implementation, and evaluation do not always belong in one context#

For a small bug like the duplicate retry issue, one agent can probably investigate, edit, and verify the fix.

For a multi hour implementation, I would consider separating some of these responsibilities.

Anthropic's recent harness work used planner, generator, and evaluator agents for larger application development tasks. One reason was that an agent evaluating its own work can be too forgiving, while a separate evaluator can inspect the result with a different context and grading objective.

A larger coding workflow might therefore look like:

Planner
   |
   v
Task Graph
   |
   v
Implementation
   |
   v
Deterministic Verification
   |
   v
Evaluator
   |
   +---- accepted
   |
   +---- revision required

I would still put deterministic verification before model judgment whenever possible.

For the payment fix, software can verify:

build succeeds

payment unit tests pass

retry integration test passes

expected files changed

no unexpected generated files

A model evaluator becomes useful for questions such as whether the patch actually addresses the intended failure mode, whether the change is unnecessarily broad, or whether the implementation fits the architecture of the repository.

The harness should not ask a model to judge something the compiler already knows.

Completion should be a system condition, not a sentence from the model#

Coding agents often fail in a very ordinary way. They make a plausible edit and declare the task complete.

For our payment retry issue, that is not enough.

The harness may know that the task requires:

CompletionContract {
    required_checks[]

    required_artifacts[]

    prohibited_states[]

    human_review_required
}

and the specific contract could require:

PaymentClient compiles

payment unit tests pass

duplicate charge reproduction passes

retry integration tests pass

no unrelated files modified

diff exists

The model can still tell the harness that it believes the task is ready.

The harness then checks whether the conditions that define ready have actually been satisfied.

Anthropic observed this failure mode in its long running coding experiments. Agents would often make code changes and perform partial verification but mark features complete before checking the actual end to end behavior. Explicit testing instructions and access to realistic verification tools improved the result.

I would make completion contracts task specific where the product knows what success means, rather than relying only on a generic prompt that says “run tests before finishing.”

Skills, MCP, hooks, repository instructions, and subagents all modify the harness#

Modern coding systems expose many extension mechanisms, and it is easy to treat each one as a separate architectural idea.

I find it simpler to ask what part of the harness the extension changes.

Repository instructions change the context.

Skills change context and procedure.

MCP adds capabilities and external data.

Hooks add deterministic control around lifecycle events.

Subagents create additional contexts and execution branches.

Sandbox configuration changes the execution boundary.

The practical implication is that harness configuration needs versioning just like model configuration.

If the same model suddenly behaves differently, the cause may be:

prompt version

repository instructions

skill version

tool schema

context builder

reasoning effort

compaction logic

sandbox image

permission policy

model version

Anthropic's April 2026 Claude Code postmortem is a good example of why this matters. Reports that looked like model quality degradation were traced to multiple product changes while Anthropic said the underlying API and inference layer were unaffected, including a change in the default reasoning effort.

When I evaluate a coding agent release, I therefore want a harness_version beside model_version.

Observability should let me reconstruct the work#

For a normal API request, latency and status may be enough.

For a coding agent, I want enough information to reconstruct why the run behaved the way it did.

A useful trace might include:

run_id
turn_id

model_version
harness_version

context_builder_version
context_tokens

compaction_version

tool_call_id
tool

sandbox_id
workspace_snapshot

files_read[]
files_changed[]

subagent_id

approval_id

completion_checks[]

latency
cost

I would also record the important transitions in the run state machine.

That makes it possible to answer questions such as:

Why did this task spend twelve minutes waiting?

Why did the agent inspect eighty files?

Why did it rerun the same test?

Did the context reset before the wrong edit?

Was the change made by the root agent or a subagent?

Did the harness recover an existing operation or execute it twice?

OpenAI's current observability surface exposes session history, turns, tool calls, subagent activity, usage, and exportable traces, which reflects the same general need.

The trace is also the raw material for evaluation.

I would evaluate the harness along with the patch#

Imagine two coding agents both eventually produce the same correct payment fix.

The first does this:

search retry symbol

read 4 files

run reproduction

change 2 files

run targeted tests

run full tests

finish

The second does this:

read 70 files

change 8 files

revert 6 files

run the same test 5 times

lose context

rediscover the bug

produce the same final patch

If the benchmark checks only whether the patch passes, both may receive the same score.

From a production perspective they are not equivalent.

I would measure at least:

task success

unsafe actions

tool calls

model turns

context tokens

files read

files changed

test behavior

recovery behavior

subagent fan out

latency

cost

Then I would look at behavior and trajectory alongside the final outcome.

A harness release that preserves task success while doubling model turns, context use, tool retries, and touched files has changed something important.

OpenAI's current guidance for programmatic tool calling recommends evaluating model turns, tool calls, retries, recovery, latency, cost, safety, and final correctness together. The general lesson is that a correct final answer can hide an inefficient or fragile trajectory.

That is also why harness improvements can materially move coding benchmarks without changing the base model. Context construction, tool design, compaction, permissions, subagents, and verification all influence how much of the model's underlying capability becomes useful behavior.

If I were building the harness, I would keep these pieces separate#

I would not literally require one class per component, but conceptually I want the boundaries to exist.

CodingHarness {
    session_store

    run_state_machine

    context_builder

    context_provenance

    instruction_registry

    skill_loader

    retrieval_layer

    tool_registry

    tool_router

    policy_engine

    operation_store

    sandbox_manager

    workspace_manager

    subagent_manager

    concurrency_controller

    compaction_manager

    recovery_manager

    completion_evaluator

    trace_store
}

The loop then becomes easier to reason about.

load durable run state

build context for the next decision

invoke model

interpret requested action

check policy

execute or delegate

record operation and observation

update durable state

evaluate progress

compact or checkpoint when needed

continue until completion contract passes

When something goes wrong, the location of the failure becomes much clearer.

If the model cannot find the relevant payment code, I inspect repository retrieval and context.

If it forgets that the duplicate charge only happens after gateway timeouts, I inspect context provenance, handoff state, and compaction.

If it executes an unsafe command, I inspect policy and sandbox boundaries.

If it opens two pull requests after reconnect, I inspect operation identity and recovery.

If three subagents overwrite each other, I inspect concurrency ownership.

If it says the fix is complete before reproducing the incident, I inspect the completion contract.

Calling all of these “the model made a mistake” does not give the engineering team much to work with.

The model is still central, but the harness decides how much of that capability becomes useful#

A better model can reason about harder code, understand larger changes, recover from ambiguity, and make better decisions. None of the harness design removes that.

The point is that the model is operating inside a system.

OpenAI's current Agents API makes the harness a first class product because long running agents need managed context, tools, subagents, session continuity, and execution environments. Anthropic's coding work has repeatedly shown that decomposition, structured state, context management, sandboxing, and separate evaluation can substantially change what the same model accomplishes.

So when I compare coding agents, I would look beyond which model name appears in the settings.

I would want to know how the system finds code, how it decides what reaches context, what happens when the context fills up, how tools return information, where commands execute, how writes are coordinated, what survives a restart, how side effects are reconciled, when subagents are useful, how completion is verified, and what the evaluation actually measures.

That is much closer to understanding how the coding agent works than looking only at the model underneath it.