Model and Inference Engineering · Principal
Annotators split 50/50 on refusal. What should the preference label teach the model?
The question
Interview question
An assistant is asked for security guidance on a dual-use topic. Half the annotators prefer a refusal. Half prefer a bounded, defensive answer. The data pipeline breaks the tie at random and emits one chosen response for preference training. The model becomes inconsistent on nearby prompts. Is more labeling enough, or is the training target itself unclear?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would read the prompt, both responses and the labeling instructions before adding another hundred votes. Some disagreement is noise. A rater missed a detail, a response quietly includes a harmful step, or the pair was presented with missing context. Other disagreement reflects a real policy or value choice. A binary chosen/rejected row cannot tell those apart once we collapse the votes. Training hard on arbitrary tie breaks asks the model to learn a decision that the organization has not made.
Separate the questions we are mixing. Is one response outside a non-negotiable safety boundary? If so, policy and expert review should define that boundary and the preference pair should not overturn it because a narrow crowd vote favored fluency. If both are allowed, what is the product default and how should it handle a user who is clearly doing authorized defensive work? A refusal can be safe but unhelpful. A detailed answer can be useful but exceed the intended boundary. We need concrete criteria for safe completion, specificity, uncertainty and escalation. OpenAI's work on collective alignment discusses disagreements that reveal real trade-offs and notes that no single default satisfies everyone. That does not tell us which response to choose in this hypothetical case.
Keep the original votes, rater instructions, rationales and relevant cohorts, rather than only the majority label. Re-label a sample with clarified guidance and experts who can judge the technical risk. Look at disagreement by prompt subtype. Maybe the split disappears when the task includes the operator's authority, or maybe the boundary remains genuinely contested. A more expressive training set can mark a pair as ambiguous, use graded preferences or separate policy categories rather than forcing every example into one winner. Whatever algorithm we choose, hold out disputed cases and assess behavior on both sides of the boundary. Do not report one overall win rate that rewards whichever response the current evaluator happens to prefer.
Would I personalize it? Only inside the allowed policy. A user's preference for a terse refusal or a detailed defensive explanation may be respected where both are permitted. It cannot grant a prohibited instruction. Research on learning from diverse preferences shows why one averaged reward can miss subgroup differences, but a production policy still needs an explicit safety boundary and testable behavior.
The interviewer may ask for a single training row because the trainer expects it. I would hold the ambiguous pair out until the policy owner resolves it, or encode a deliberately chosen default with its rationale and confidence. A random coin toss may make a dataset shape valid. It does not create ground truth. The useful artifact from disagreement is often a better rule and a better evaluation slice, not another scalar label.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →