skip to content

Why is entity resolution the hardest part of building a graph RAG index?

level: seniorimportance: should knowfreq 46%

answer

  1. mentions are not identities
  2. the extractor sees one chunk only
  3. fragmentation thins every neighbourhood
  4. one bad merge fabricates a relation
  5. block, score, then adjudicate the middle

basics

~20 s

Extraction yields surface forms, not identities. Under-merging shatters one real entity into many weakly-connected nodes so traversal finds nothing; over-merging fuses two distinct entities and manufactures relations that were never in the source. Both corrupt every query that follows.

solid answer

~50 s

An extraction pass over 40k discovery documents in an M&A review will emit "Acme Holdings Ltd", "Acme Holdings", "AHL" and "the Company" as separate entities, because the LLM sees each chunk in isolation. If you do not merge them, the subsidiary's node degree is split four ways, its neighbourhood is thin, community detection splits it across clusters, and a multi-hop query that should connect an indemnity clause to a signing subsidiary finds no path. If you merge too aggressively, two genuinely different subsidiaries collapse into one node and the graph now asserts that one signed a clause the other signed — a fabricated relation that reads as sourced fact. In practice you block candidate pairs cheaply, score them on name similarity plus contextual embedding plus shared neighbours, and send the ambiguous middle band to an LLM or a human adjudicator. Keep provenance on every merge so it can be undone, and measure merge precision and recall on a labelled pair set rather than eyeballing the graph.

go deeper

for a junior

Know that the same real-world thing appears under many names across documents, and that a graph is only useful once those mentions are merged into one node.

for a middle

Explain both failure directions concretely — fragmentation thinning neighbourhoods and breaking traversal, versus over-merging creating edges no source supports — and describe blocking plus similarity scoring as the practical pipeline.

for a senior

Demonstrate operational judgement: precision-biased thresholds in high-stakes corpora, LLM or human adjudication of the ambiguous band, provenance so merges are reversible, and a labelled pair set so quality is measured rather than assumed.

for a principal

Frame it as record linkage, a long-standing problem LLMs bootstrap but do not solve, and own the consequences — resolution quality caps the ceiling of every graph query, so budget for adjudication effort and re-resolution on ingest before committing to the architecture.

## The gap between extraction and identity Graph construction has two distinct steps that are easy to conflate. Extraction reads a chunk and emits entities and relations found in that chunk. Resolution decides which of those extracted mentions refer to the same real-world thing and merges them into one node. Extraction is the step people demo; resolution is the step that determines whether the graph is usable. The reason is scope. An extraction pass sees one chunk at a time. It cannot know that the "AHL" in a 2019 side letter and the "Acme Holdings Ltd" on a 2021 share purchase agreement are the same legal entity, or that "the Company" in a schedule refers to whichever party the parent contract defined. Every corpus produces this: people with initials and married names, products with SKUs and marketing names, drugs with generic and brand names, subsidiaries with pre- and post-rename identities. ## Two failure modes, asymmetric in cost **Under-merging** fragments one entity into several nodes. The consequences are graph-wide, not local: node degree is divided across the fragments, so each looks peripheral; two-hop traversal from one fragment cannot reach edges attached to another; community detection places the fragments in different clusters, so the global summaries describe the same entity twice, each partially. The visible symptom is a graph that looks rich but answers poorly — retrieval returns thin neighbourhoods and multi-hop questions come back empty. **Over-merging** is rarer and worse. Fusing two distinct subsidiaries creates edges between things that were never connected in any source document. The system will then answer, with a citation, that entity A signed an indemnity clause that entity B actually signed. This is not a retrieval miss you can detect by low confidence; it is a confident, sourced, wrong answer. In a due-diligence or pharmacovigilance setting that is the failure that ends the project. The asymmetry matters for threshold choice. In high-stakes domains you tune towards precision — merge only when confident, accept some fragmentation — and treat the remaining fragmentation as a known, measurable gap rather than an invisible fabrication. ## How resolution is actually done All-pairs comparison is quadratic and infeasible past a few thousand entities, so the pipeline is staged. **Blocking** cheaply groups candidates that could plausibly match — shared normalised tokens, phonetic keys, same entity type, approximate-nearest-neighbour search over name embeddings. Only within-block pairs are scored, which turns the quadratic problem into something tractable. **Scoring** combines several weak signals: string similarity on canonicalised names (case, punctuation, corporate suffixes like Ltd/Limited/LLC stripped), embedding similarity over the entity's generated description rather than just its name, and structural evidence — do the two candidates share neighbours, appear in the same documents, have compatible types and date ranges? **Adjudication** handles the ambiguous middle band. High-confidence pairs auto-merge, low-confidence pairs stay separate, and the middle is sent to an LLM given both entities' descriptions and a few source snippets, or to a human when the domain warrants it. Explicit alias tables and domain identifiers — company registration numbers, drug codes, employee IDs — should short-circuit all of this whenever the corpus contains them; a deterministic key beats any similarity score. **Relations need the same treatment.** After nodes merge, duplicate edges expressing the same relation in different words ("acquired", "purchased a controlling stake in") should be canonicalised to a relation vocabulary, or traversal filters by relation type become unreliable. ## Making it operable Three practices separate a resolution step that can be maintained from one that cannot. Record **provenance**: every merged node keeps the list of mentions and source chunks it absorbed, so a bad merge can be traced and reversed without rebuilding the index. Merges should be reversible data, not a destructive transformation. **Measure it.** Hand-label a few hundred candidate pairs as match/non-match and report precision and recall on merges, tracked over time. Without this, resolution quality is a matter of opinion and silently degrades as new documents introduce new naming conventions. **Re-resolve on ingest.** New documents introduce new surface forms for existing entities. If ingest only adds nodes and never revisits merge decisions, fragmentation grows monotonically and the graph degrades even though nothing visibly broke. Finally, be honest in an interview that entity resolution is a decades-old record-linkage problem that LLMs have made easier to bootstrap but have not solved. Treating it as a solved preprocessing detail is the single clearest sign someone has demoed graph RAG rather than run it.

  • Which direction would you bias the merge threshold for a legal due-diligence corpus, and why?
    Towards precision — merge only on strong evidence. Under-merging leaves fragmented entities and missing answers, which is a visible, measurable gap. Over-merging invents relations, so the system states that one subsidiary signed a clause another signed, with a citation attached. A confidently wrong sourced claim is far more damaging in that setting than a recall miss, and it is much harder to detect.
  • How would you measure whether entity resolution is working at all?
    Hand-label a few hundred candidate pairs as match or non-match, sampled from the blocking output so the hard cases are represented, and report merge precision and recall against that set on every index build. Complement it with graph-level signals — duplicate-name clusters remaining, degree distribution, and the share of multi-hop eval queries that find a path.
  • What happens to resolution when new documents arrive after the graph is built?
    New documents introduce new surface forms for entities that already exist, so an ingest path that only appends nodes lets fragmentation grow monotonically while nothing visibly breaks. Incremental ingest has to re-run blocking and scoring for new mentions against existing nodes, and periodically revisit borderline pairs, otherwise graph quality decays quietly between full rebuilds.

saying these in an interview costs you the question

  • Treats entity resolution as a solved preprocessing detail
  • Assumes exact string matching is sufficient for merging
  • Thinks over-merging is safer than under-merging
  • Compares all entity pairs without blocking
  • Merges destructively with no provenance or reversal path

context