Evaluation and Quality · Principal
The coding agent passes each task alone. Why does the benchmark fail when tasks run together?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Start by asking what was reused. If two eval cases share a checkout, database, background process, container volume or result cache, task B may start from a state task A changed. A successful run might leave a patch, generated fixture or compiled artifact that makes the next task easier. A failing run might leave a lock or broken dependency that makes it harder. The benchmark is then measuring the task order and harness behavior along with the agent. The SWE-bench evaluation harness uses containerized instance environments as a reproducibility boundary. That is an example of a controlled harness, not proof that merely choosing Docker isolates every external service.
I would replay two cases A and B in both orders, alone and in parallel. Log the base commit and patch hash before each run, container and volume identity, network fixtures, test-process IDs, random seed, environment version and any cache key. If B fails only after A, compare all state B could observe. Separate model nondeterminism from environment interference by holding model inputs and generation settings fixed where possible, then checking which bytes and tool results actually differed.
An isolation plan needs to include mutable external dependencies. A fresh git worktree does not isolate a shared test database or a hosted API with rate limits. Use per-task namespaces, deterministic fixtures and scoped credentials, or record and replay allowed tool observations. Clean up processes and resources after failures, then verify cleanup rather than assuming it. Content-addressed caches can be shared for immutable inputs, but a cache keyed only by task name or run ID can accidentally return another model's result. Evaluate cache behavior separately from agent skill.
There is a release question here. If production agents also share a workspace, some interference may be a real product failure. Keep a clean isolated benchmark for capability measurement and a second concurrent-workload test for the shared product design. Both numbers matter, but they answer different questions.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →