How do you design a judge agent that picks between two agents' answers?
answer
- ground it in evidence, not arguments
- a rubric it did not invent per call
- cite which candidate each claim came from
- blind and shuffle the candidates
- let it abstain instead of always deciding
basics
~20 sGive the judge the rubric and the underlying evidence, not just the two arguments, so it checks claims instead of rating rhetoric. Constrain its output to a verdict plus cited support, blind and shuffle the candidates, and route low-confidence cases to a human.
solid answer
~50 sTake a content-moderation appeal argued by a proposer and a critic over two rounds, with a third agent issuing the verdict. The judge's context should carry the policy text and a decision rubric alongside the two positions — without the source of truth it can only score persuasiveness. I constrain its output to a schema: verdict, the specific policy clause relied on, and a citation of which candidate each supporting claim came from, so the decision is auditable rather than a paragraph of prose. Candidates get shuffled and stripped of speaker identity, because judges drift toward the last or the longer argument. The judge's own failure modes matter: it cannot catch an error both debaters shared, and it will produce a verdict even when neither position is defensible. So I add an abstain option with escalation to a human, and where the question is actually checkable I run a verifier instead of a judge.
go deeper
Know what a judge agent does — it reads the competing answers and issues one decision — and that it should be given the evidence and the rules, not just the arguments. Being able to say the judge is another model with its own mistakes is the point to land.
Explain the inputs (rubric, source of truth, blinded and shuffled candidates) and the constrained output (verdict from a closed set, cited support). Be ready to name why a judge holding only the transcripts ends up rewarding the more fluent argument.
Demonstrate you treat the judge as a component with failure modes: it cannot catch an error both sides shared, it will decide when it should abstain, and it needs lookup tools for checkable premises. Expect to describe the abstain-plus-human-escalation path and how you audit verdicts against human adjudication.
Own the accountability story — provenance in every verdict, a rubric that is versioned and reviewable, and a stated policy for which decisions a model may finalise at all. Be ready to argue where a deterministic verifier must replace the judge, and how judge drift is detected after a policy or model change.
## Where the judge sits In a debate or proposer/critic topology, someone has to convert an argument transcript into a single decision. That is the judge — or, when the job is to fuse several partial answers rather than pick one, the aggregator. Both are agents in the graph with their own prompt, their own context and their own failure modes, and treating them as a neutral piece of plumbing is the mistake the question is testing for. Run the concrete case throughout: a content-moderation appeal. A user disputes a takedown. A proposer agent argues to reinstate, a critic argues to uphold, they exchange two rounds, and a third agent issues the verdict. ## What the judge must be given **The ground truth, not just the arguments.** The judge needs the policy text, the original content, and the enforcement history. A judge holding only two arguments is scoring rhetoric, and the more fluent agent wins regardless of who is right. This is the single most common design defect. **A rubric.** Spell out the decision criteria and their precedence: which policy clauses apply, what evidence is required to overturn, what happens on genuine ambiguity. Without a rubric the judge invents one per call, and the system becomes unreproducible. **The candidates, blinded and shuffled.** Strip speaker identity and randomise presentation order. Judge models show order and length effects, and a judge that quietly favours the last speaker will look like a consistent policy for weeks before anyone notices. **Verifiers where they exist.** If part of the dispute is checkable — the account's strike count, whether the content matches a hash in the known-violations set — let the judge look it up rather than believe an agent's assertion about it. ## What the judge must return Prose verdicts are close to useless operationally. Constrain the output: - The verdict itself, from a closed set — uphold, reinstate, escalate. - The specific rubric or policy clause the verdict rests on. - Provenance: for each supporting claim, which candidate argument it came from, and which piece of evidence backs it. - A confidence or abstain signal. Provenance is the requirement that turns the judge from a black box into an auditable component. When a verdict is later challenged, you can show which argument the decision leaned on and check whether that argument was itself grounded. It also creates a cheap automated check: a supporting claim that cites nothing, or cites a candidate that never made it, is a fabrication you can detect without a human reading the transcript. ## The aggregator variant When the job is to fuse rather than choose — several agents each contributed part of an answer — the same discipline applies with one addition: every sentence in the fused output must be traceable to a source candidate. Aggregators left unconstrained produce smooth synthesis containing claims that no input actually made, which is the worst possible failure because the fusion step is exactly where a reader stops checking. ## The judge's own failure modes **Shared blind spots.** A judge cannot catch an error both debaters made. If both agents misread the policy the same way, the judge sees agreement and ratifies it. This is why grounding the judge in the policy text rather than the transcript matters so much. **Compulsory verdicts.** Given two positions, a model will pick one, even when neither is defensible and the correct output is "insufficient evidence". Always provide an abstain path, and treat a rising abstain rate as a signal about the upstream agents rather than a defect in the judge. **Surface bias.** Length, confidence of phrasing and position in the prompt all move judge decisions. Blinding, shuffling and a low sampling temperature reduce it; a periodic audit against human labels tells you how much remains. **Cost and criticality.** The judge is a single point of failure at the end of an expensive pipeline. It usually deserves the strongest model in the system, which inverts the instinct to save money on the small final step. ## When not to use a judge at all If the disagreement is decidable — a policy lookup, a number that must reconcile, a test that passes or fails — run the check. A deterministic verifier is cheaper, faster, reproducible and actually correct. Reserve the judge for genuinely contested calls where no oracle exists. In the moderation example, "does this account have three prior strikes" is a lookup; "is this satire or harassment" is a judgement, and only the second needs an agent. ## Keeping a human in the path For consequential decisions, wire the abstain signal and low confidence to human review rather than to a default. Then sample decided cases too — a judge that only gets audited when it abstains gives you no measurement of the cases it got confidently wrong. Comparing a sampled slice against human adjudication on a schedule is the only way to know the judge has not drifted.
- What does a judge miss that neither debater can surface?An error both sides share. If the proposer and the critic both misread the same policy clause, the disagreement never touches it, so the judge sees consensus on a false premise and ratifies it. The mitigation is to ground the judge in the source material directly rather than in the transcript, and to give it lookup tools so it can check factual premises that neither agent contested.
- Why require the judge to cite which candidate each supporting claim came from?It makes the verdict auditable and catches fabrication cheaply. A claim citing nothing, or citing an argument that was never made, is detectable programmatically without a human reading the transcript. It also lets you replay disputed decisions and see which argument the outcome actually rested on, which matters when the decision is later challenged.
- Should the judge run on the same model as the debaters?Not necessarily, and there are reasons not to. A different model breaks shared blind spots, so a premise both debaters got wrong has some chance of being caught. The judge is also the last step before a decision ships, so it usually justifies the strongest model in the system even when the debaters run on cheaper ones — the opposite of the instinct to economise on the final call.
- How do you know the judge is still calibrated six months in?Audit against human adjudication on a sampled slice, and sample decided cases rather than only abstentions — otherwise you never measure the confidently wrong ones. Track agreement rate over time by decision type. Drift shows up as a slow shift in the verdict mix or in agreement on a particular category, usually after a policy change or a model upgrade that nobody connected to the judge.
saying these in an interview costs you the question
- Gives the judge only the two arguments and no source evidence
- Lets the judge return free prose with no verdict schema or citations
- Assumes the judge is neutral because it did not produce either answer
- Forces a verdict with no abstain or escalation path
- Uses a judge where a deterministic lookup or test would settle the question