skip to content

How do you catch semantic misalignment when two agents both report success?

level: seniorimportance: must knowfreq 58%

answer

  1. same words, different meanings
  2. both succeed locally, the whole is wrong
  3. silent — nothing throws
  4. check the artifact, not the claim
  5. enums and IDs instead of prose handoffs

basics

~20 s

Self-reports cannot catch it, because each agent met its own reading of the task. Catch it by verifying the produced artifact against the original request with an independent checker, and by making handoffs typed rather than prose.

solid answer

~50 s

Semantic misalignment is agents sharing words but not meanings: a planner writes "deploy to stage" meaning an environment, the executor reads it as a pipeline stage, both finish cleanly, and the combined result is wrong. It is the dominant multi-agent failure precisely because it is silent — every local success criterion is satisfied, so no retry, timeout or exception fires. Detection has to be independent of self-report: verify the *artifact* against the original request, using a checker that sees the request and the output but not the producing agent's reasoning, so it does not inherit the misreading. Cheaper structural checks help — ask each worker to restate its subtask and diff the restatements, and assert invariants on the result. Prevention is better: a frozen plan as the single referent, handoffs carrying enumerated values and stable IDs rather than free prose, and orchestrator-written acceptance criteria per subtask.

go deeper

for a junior

Recall what the term means: agents interpreting the same instruction differently and both finishing successfully. Be able to give an example of an ambiguous word in a handoff and say why the run shows no error.

for a middle

Explain why it is silent — each agent satisfies its own criterion, so no retry or exception fires — and name the structural cause: prose handoffs plus context isolation that strips shared grounding.

for a senior

Demonstrate a detection strategy that ignores self-reports: programmatic checks on the artifact, a verifier given a clean context, invariants across the joint result, and pre-work restatement diffs. Tie each to a concrete symptom you would have seen.

for a principal

Argue where the verification budget should go, since cross-checking can cost as much as the work. Own the design position that typed contracts and a single frozen referent are cheaper than any amount of downstream checking, and say when ambiguity should be escalated instead of resolved by an agent.

## What semantic misalignment is Semantic misalignment is the situation where two agents in a fleet operate on the same words while holding different meanings for them. The planner writes "stage"; it means the pre-production environment. The executor reads "stage"; it means a step in the build pipeline. Both agents do competent, internally coherent work. Both report success. The composed result is wrong, and nothing in the run signals it. It is the coordination failure that matters most in practice — more than deadlock, more than races. Deadlock is loud: work stops and a deadline fires. A race corrupts state in ways that verification and diffs can surface. Misalignment produces a clean run with a wrong answer, which is the most expensive kind of failure because it is discovered downstream, often by a user. ## Why multi-agent designs are structurally prone to it The handoff medium is natural language. Prose is expressive and lossy at the same time: it carries intent well enough for humans who share context, and terribly for a receiver that has none. Compounding this, the isolation that makes multi-agent designs pay — giving each subagent its own clean context rather than the orchestrator's full history — is exactly what removes shared grounding. The subagent gets a paragraph where the orchestrator had a conversation. Every ambiguity in that paragraph is resolved by the subagent's priors, not by the task's actual intent. Three recurring sources: - **Underspecification.** The subtask omits a constraint the orchestrator considered obvious. The worker fills it in, plausibly and wrongly. - **Overloaded vocabulary.** Domain words with multiple readings — stage, environment, customer, order, run, ticket, job — resolved differently on each side. - **Long-horizon drift.** Over many steps, the working definition an agent holds slides away from the original, especially after summarization or compaction, and no single step looks wrong. ## Detection The governing rule: never treat an agent's claim of success as evidence of success. A misaligned agent is not lying; it genuinely satisfied its own criterion. **Verify the artifact, not the report.** Take the produced output and check it against the original request. Where a programmatic check exists — a test, a schema, a dry-run plan, a type check, a query that counts rows — use it, because it is deterministic and cheap. Where the check must be judgemental, use a separate reviewer agent, and give it the original request plus the artifact, deliberately *not* the producing agent's trajectory. A reviewer that reads the producer's reasoning tends to be persuaded by it and inherits the same misreading; one with a clean context is the whole point of the generator-verifier split. **Surface disagreement explicitly.** Before work starts, have each worker restate its subtask, the key terms and its acceptance criteria in its own words. Diff those restatements against the plan. This is a small number of tokens and it catches the overloaded-vocabulary class before any work is wasted. **Assert invariants.** Properties that must hold of the joint result regardless of interpretation — the environment touched, the resource count, the schema version, referential integrity across the artifacts different agents produced. An invariant violation is a hard signal where two success reports are not. **Watch for the fingerprint.** Two agents describing the same object with different nouns, mismatched identifiers across handoffs, or a result that is internally consistent but answers a question nobody asked, all read as misalignment rather than tool failure. Tool failures produce errors; misalignment produces coherent wrongness. ## Prevention, which is where the leverage is **One frozen referent.** A plan artifact accepted before fan-out, which every worker cites. Disagreement then happens once, in front of the orchestrator, rather than N times in private. **Typed handoffs.** Replace prose parameters with enumerated values and stable identifiers: an environment field constrained to a known set, a resource ID rather than a name, a version rather than "latest". This converts an entire class of ambiguity into a validation error at the boundary — the receiving side either matches the contract or fails loudly. **Orchestrator-written acceptance criteria.** Each dispatched subtask carries a checkable definition of done, written by the component that holds the whole picture. "Done" defined by the worker is defined by the worker's interpretation. **Return artifacts, not claims.** A subagent that hands back a diff, a file, or a structured record gives the orchestrator something to inspect. One that hands back "completed successfully" gives it nothing. **Single-writer discipline as containment.** When workers can only propose, a misaligned worker produces a rejected proposal instead of corrupted shared state. ## Honest limits Verification costs tokens, and thorough cross-checking can approach the cost of the work itself — one reason single-agent designs sometimes match multi-agent ones at equal budget. Judgemental verification is itself fallible, and correlated: a verifier sharing the producer's model, prompt framing and context can share its blind spot. The mitigations are to prefer deterministic checks wherever one exists, to vary the verifier's context, and to accept that for genuinely ambiguous requests the right move is to escalate the ambiguity rather than let an agent resolve it silently.

  • What distinguishes semantic misalignment from an ordinary tool failure in a trace?
    A tool failure produces an error, a retry, or an abandoned step — visible discontinuity. Misalignment produces a clean trajectory with no anomalies: every call succeeds, every step looks reasonable, and the wrongness is only visible when the output is compared to the original request. If a run looks healthy but the result answers a slightly different question, suspect misalignment.
  • How do you make a handoff contract machine-checkable rather than merely clearer prose?
    Give the handoff a schema: enumerated fields for anything with a fixed domain, stable identifiers instead of human names, explicit units and versions, and required acceptance criteria. Validate at the boundary and reject on mismatch. Ambiguity then fails as a validation error at dispatch time instead of surviving as a plausible misreading through the whole run.
  • Why give the verifier a clean context rather than the producer's full trajectory?
    A verifier that reads the producer's reasoning tends to be persuaded by it and adopts the same interpretation, so it confirms rather than checks. Handing it only the original request and the artifact forces an independent reading. It is the same reason a code reviewer who watched you write the code catches fewer bugs than one who did not.

It is the two-builders problem: one builds the staircase to metric, one to imperial, both finish on schedule, and neither notices until the pieces have to meet.

saying these in an interview costs you the question

  • Trusting each agent's success report as evidence the task was done
  • Assuming deadlock and races are the main multi-agent failure modes
  • Believing longer, more polite prose instructions eliminate ambiguity
  • Using a verifier that shares the producer's context and reasoning
  • Treating a coherent-looking trace as proof the output is correct

context