Model and Inference Engineering · Principal
The coding policy passes every training test. Why does it fail hidden cases?
The question
Interview question
A code-generation policy gets binary reward for passing a small, fixed set of unit tests. After enough updates it produces short solutions that pass nearly all of them. Independent hidden tests still fail on empty inputs, overflow and boundary cases. Is the policy improving at programming?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
It is improving on the reward it can observe. A small test suite is evidence about a program's behavior on those inputs, not a proof of the intended function. If the same tasks and tests are reused for many training steps, the policy can learn shortcuts that exploit their coverage even without literally reading or editing the tests. Research on fuzzing RLVR verifiers studies how verifier bugs can become optimization targets. That supports a risk, not a claim that every failing hidden case was consciously gamed by the model. Some may be ordinary generalization failures or a bad task specification.
I would classify failures first. Is the prompt ambiguous, the solution invalid under the stated constraints, the verifier wrong, or the hidden suite checking a requirement never expressed? For a sample of tasks, compare public-test pass rate with held-out edge and property tests, mutation testing, and differential checks against trusted reference implementations. Keep test generation and labels outside the policy's writable sandbox. Randomized tests can raise coverage, but fixed seeds used forever can become another memorized target. Version and audit test suites so reward changes are visible in the training trace.
I would not secretly move all hidden cases into the reward and declare success. That can lift training reward while consuming the independent evidence of generalization. Keep a genuinely held-out task set, preferably with fresh problems and tests from a separate process. For training, strengthen verifier coverage and problem specifications, and use reward signal only from trusted tests with bounded compute. Measure pass on new tasks under the actual sampling budget, not just the number of public tests passed on seen prompts.
What if the policy cannot learn from a binary pass/fail signal when the hard test suite is broad? Every sampled solution gets zero reward. Why is the GRPO training run doing almost no learning? explains how all-zero groups remove relative task advantage. A more informative reward can help, but partial-credit tests can make the policy optimize easy cases while ignoring essential rare ones. A coding policy earns full reward by changing the tests. What was the verifier allowed to trust? covers changing the test files themselves. Here the tests are immutable and the reward is still an incomplete proxy for correct code.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →