When would you build a DeepEval DAGMetric instead of a single GEval metric?
answer
- one number vs a traced decision path
- hard gates that nothing can compensate for
- small decisions are judged more reliably
- leaves can hand off to a nested judged metric
- every node on the path is a model call
basics
~20 sUse DAGMetric when the judgement is a sequence of decisions rather than one holistic score — a decision tree of judgement nodes ending in verdicts. It makes the reasoning path explicit and auditable, where GEval collapses everything into one opaque judged number.
solid answer
~50 s`GEval` asks a judge for one number covering the whole rubric. That works while the rubric is one idea. Once it becomes "first check the format is right; if it is not, this fails outright; if it is, then judge whether the summary is faithful; if it is, check the tone", a single judged score hides which check failed and lets a strong showing on one dimension paper over a hard failure on another. `DAGMetric` in DeepEval expresses that as a decision tree. You build a `DeepAcyclicGraph` from nodes — `TaskNode` to extract something, `BinaryJudgementNode` for a yes/no decision, `NonBinaryJudgementNode` for a multi-way one, and `VerdictNode` to terminate a branch with a fixed score or delegate to a nested `GEval`. Each node is one small decision the judge can make reliably. The price is authoring effort and more model calls down a path, so I reach for it when a criterion has genuine hard gates or when someone needs to audit why a case scored what it did.
go deeper
Know that DeepEval offers DAGMetric as an alternative to GEval, and that it evaluates a case by walking a decision tree of judgement nodes instead of asking for one holistic score.
Name the pieces — DeepAcyclicGraph with root nodes, TaskNode, BinaryJudgementNode, NonBinaryJudgementNode and VerdictNode — and explain that a verdict either fixes a score or hands off to a nested GEval.
Argue the tradeoff concretely: hard gates that block compensation, per-node reliability, failure classes you can aggregate, against authoring cost and one model call per node on the traversed path. Say that cheap gates go first to short-circuit.
Decide when the org invests in structured metrics at all: which criteria carry enough regulatory or product weight to justify an auditable decision procedure, who maintains the trees as outputs change, and what the judged-call budget for evaluation is.
## The limitation being worked around A single `GEval` call asks a judge to hold an entire rubric in mind and emit one number. When the rubric contains several independent conditions, three things go wrong: 1. **Compensation.** Excellent prose can pull up a score that a hard format violation should have floored. 2. **Opacity.** A score of 0.55 with a paragraph of reasoning does not tell you which condition failed, so you cannot aggregate failures by cause across a run. 3. **Reliability.** Judges are more consistent on small, decidable questions than on compound holistic ones. ## What DAGMetric is `DAGMetric` (from `deepeval.metrics`) takes a `DeepAcyclicGraph` and walks it per test case. The graph is assembled from node types in `deepeval.metrics.dag`: - **`TaskNode`** — an extraction or transformation step. You give it `instructions`, the `evaluation_params` it may look at, an `output_label` naming what it produced, and `children`. Use it to pull out the part of the output later decisions are about. - **`BinaryJudgementNode`** — a yes/no judgement over `criteria`, with children for each branch. - **`NonBinaryJudgementNode`** — the same idea with more than two outcomes, for classifying a case into one of several buckets. - **`VerdictNode`** — a leaf. It matches a parent's verdict and either assigns a fixed `score` or hands off to a nested `GEval` metric via `g_eval`, which is how you get graded scoring at the end of a branch that has already passed its gates. The graph's `root_nodes` are where evaluation begins, and `DAGMetric` reports a 0-1 score against a `threshold` just like any other DeepEval metric, so it drops into the same suite alongside built-in metrics. ## Why the structure earns its keep **Hard gates stop compensation.** If the first node asks "is the output valid JSON matching the required keys?" and the `False` branch terminates in a `VerdictNode` with `score=0`, no amount of downstream quality can rescue it. That is often exactly the semantics you want and exactly what a holistic judge will not respect. **Each decision is small.** "Does this contain a date in ISO format?" is a question a judge answers the same way nearly every time. "Rate this output's overall quality from 0 to 10 considering format, faithfulness and tone" is not. Decomposing raises per-decision reliability, which is worth more than it sounds when you are trying to keep a suite stable. **The path is the explanation.** Because evaluation follows a traced route through named nodes, a failing case tells you which gate it hit. Over a run, that turns into a distribution — "63% of failures are format, 21% are unsupported claims" — which is a prioritisable bug list rather than a bag of low scores. **Mixed determinism.** Some branches end in fixed scores; some delegate to a nested `GEval` for the genuinely fuzzy part. You spend judged, variable scoring only where the question is actually a matter of judgement. ## The costs **Authoring.** A DAG is code, and a tree with a dozen nodes is a small program to design, review and keep in step with the product. A one-line criterion is not. **Model calls.** Each judgement node on the traversed path is a call. A five-deep path costs roughly five times a single `GEval` on that case. Across thousands of samples that is a real bill, so you want the cheap deterministic gates early to short-circuit before the expensive nodes run. **Brittleness to product change.** A tree encodes assumptions about output structure. When the product's output shape changes, the tree needs editing in more places than a prose rubric does. ## Choosing between them Reach for `GEval` when the criterion is one coherent idea, when you want a gradient, and when you are still exploring what "good" means. Reach for `DAGMetric` when there are hard disqualifiers that must not be compensated for, when different failure classes need distinguishing in the results, or when a reviewer outside the team needs to audit the decision procedure. A common and healthy end state is a DAG whose early nodes are cheap mechanical gates and whose surviving branch delegates the interesting judgement to a nested `GEval` — deterministic where you can be, judged only where you must be.
- How do you keep a DAG metric from becoming expensive on a large suite?Order the tree so the cheapest, most decisive gates run first. A format or presence check that terminates a branch in a fixed-score verdict costs one small call and stops traversal, so the expensive nested judged scoring only runs on cases that got that far. Deep, judgement-heavy paths at the top of the tree are what turn a DAG into a bill; short-circuiting early is the whole optimisation.
- What does a VerdictNode do at the end of a branch?It terminates that branch. A VerdictNode matches the parent node's verdict and either assigns a fixed score — which is how you encode a hard fail or a hard pass — or delegates to a nested GEval metric so the branch is graded rather than binary. That mix is the point: mechanical outcomes get constant scores, genuinely fuzzy ones get a judged score, in the same metric.
- Why are small per-node judgements more reliable than one holistic score?A narrow question like "is there an ISO-format date present?" has a decidable answer the judge reproduces consistently. A compound one asks the model to weigh several incommensurable dimensions and compress them into a number, which it has no principled way to do, so the answer wanders between runs. Decomposition trades one unstable judgement for several stable ones plus an explicit combination rule you wrote.
- When is a DAG the wrong choice?When the criterion really is one idea — helpfulness, tone, relevance — a tree adds nodes without adding information, and you have paid authoring cost for structure that mirrors nothing real. It is also wrong while you are still discovering what good means, because every discovery means re-plumbing the graph. Start with GEval, and promote to a DAG once the failure classes have stabilised.
saying these in an interview costs you the question
- Thinks DAGMetric removes the judge model from scoring
- Uses a DAG for a single coherent criterion like tone
- Ignores that each judgement node on the path is another model call
- Cannot name any node type beyond the metric class itself
- Assumes a DAG makes evaluation deterministic end to end