Which Ragas metrics need an embedding model, not just a judge LLM?
answer
- not every metric calls the same kind of model
- two dependencies, wired separately
- similarity means vectors, vectors mean embeddings
- ResponseRelevancy embeds generated questions
- judge-only metrics need no embeddings at all
basics
~20 sMost ragas metrics are pure LLM-as-judge and need only an evaluator LLM. Similarity-based ones — ResponseRelevancy above all — additionally embed text and compare vectors, so a run containing them fails or misbehaves if you wired an LLM but no embeddings.
solid answer
~50 sRagas has two kinds of model dependency and they are not interchangeable. Metrics such as `Faithfulness`, `LLMContextRecall`, `LLMContextPrecisionWithReference`, `FactualCorrectness`, `NoiseSensitivity`, `AspectCritic` and `RubricsScore` only ever prompt a judge model, so an evaluator LLM is all they need. `ResponseRelevancy` is different: it asks the judge to reverse-engineer questions the response would answer, then embeds those questions and the original user input and compares them by cosine similarity — that second half needs an embedding model. Any similarity-scored metric has the same requirement. So when you assemble a metric list, check whether it contains a similarity-based metric; if it does, `evaluate()` needs `embeddings=` as well as `llm=`, wired with `LangchainEmbeddingsWrapper`. Forgetting it is the single most common first-run failure, and the error surfaces on that metric only, which is why people misread it as a bug in the metric.
go deeper
Know that ragas can need two models, not one: a judge LLM for most metrics and an embedding model for similarity-based ones like ResponseRelevancy. Recognise a missing embeddings= as the cause of a single metric failing.
Explain the mechanism: ResponseRelevancy generates candidate questions with the judge, then embeds them against the user input and averages cosine similarity — the embedding call is the second dependency. Contrast that with prompt-and-parse metrics like Faithfulness.
Treat the embedding model as part of the instrument. Pin it, keep it distinct from the retriever's embedding model, and be able to explain why a relevancy score moved after someone 'just upgraded embeddings'.
Set the policy that measurement models — judge and embedding alike — are versioned artefacts of the evaluation platform, changed deliberately with a re-baseline, never inherited from whatever the application happens to use this quarter.
## Two different model dependencies Everything in ragas is scored by a model, but not by the same *kind* of model. It is worth holding two categories in your head: 1. **Judge-only metrics.** The metric builds one or more prompts, sends them to the evaluator LLM, and turns the answers into a number. `Faithfulness` decomposes the response into claims and asks whether each is supported by the retrieved contexts. `LLMContextRecall` and `LLMContextPrecisionWithReference` ask the judge to compare retrieved context against a reference. `FactualCorrectness`, `NoiseSensitivity`, `AspectCritic` and `RubricsScore` are all prompt-and-parse. None of them ever produces a vector. 2. **Metrics with an embedding step.** `ResponseRelevancy` is the one you will meet first. Its procedure is two-phase: the judge LLM generates a set of questions that the response *would* be a good answer to, and then those generated questions and the user's actual question are embedded and compared by cosine similarity. The score is a similarity average, not a judge verdict. That embedding call is a second, independent model dependency. Similarity-style metrics in general — anything whose definition is "how close is this text to that text" rather than "does the judge think this is supported" — sit in this category. ## Why the distinction bites Because the dependency is per-metric, a partially-wired run does not fail cleanly at startup. You pass `llm=` to `evaluate()`, the judge-only metrics score happily, and then the one similarity metric in your list blows up or produces nothing while everything around it looks fine. Read as "ResponseRelevancy is broken", it sends people hunting in the wrong place. Read as "this metric has a second model dependency I never satisfied", it is a one-line fix: wrap an embeddings object in `LangchainEmbeddingsWrapper` and pass it as `embeddings=`. The same asymmetry runs the other way. If your metric list happens to contain no similarity metric, you genuinely do not need an embedding model at all, and wiring one is dead configuration — one more deployment to provision on Azure, one more thing to get wrong. Knowing which half of the catalogue you are using lets you provision exactly what the run needs. ## The embedding model is a scoring input, not decoration A subtle consequence: the embedding model you wire is part of the measurement instrument. Swap `OpenAIEmbeddings` for a different embedding model and every `ResponseRelevancy` score shifts, because cosine similarity is computed in a different space with different geometry. Absolute values become incomparable across the swap in the same way they become incomparable when you change judge models. If you track relevancy over time, pin the embedding model and treat a change to it as a re-baseline, not a config tweak. It does not have to be — and often should not be — the same embedding model your RAG pipeline uses for retrieval. The application's embedding model is chosen for retrieval quality at your corpus and latency budget; the evaluator's is chosen to compare two short pieces of text sensibly. Coupling them means every retrieval experiment silently moves your scoreboard. ## Cost profile differs too Embedding calls are typically orders of magnitude cheaper than judge calls, so the embedding half of a run is rarely what drives the bill. But it is another provider dependency in the request path: another endpoint that can rate-limit you, another key that can expire, another regional deployment to stand up. On Azure in particular, teams deploy only a chat model, wire it successfully, and are then surprised by the first embedding-dependent metric, because the embeddings deployment was never created. ## What to say in an interview The compact answer is: "Nearly all of them need only the judge LLM; `ResponseRelevancy` also needs an embedding model because its second phase is a cosine-similarity comparison, and any similarity-based metric is the same. So before a run I check the metric list for a similarity metric and wire `embeddings=` if there is one." That answer shows you have actually assembled a metric list rather than pasted one.
- Should the evaluator's embedding model be the same one your retriever uses?Usually not. The retriever's embedding model is chosen for retrieval quality on your corpus; the evaluator's is a measuring instrument for comparing two short texts. Sharing them means every retrieval experiment quietly moves the scoreboard you are using to judge that experiment. Keep them separate and pin the evaluator's.
- If you change the evaluator embedding model, are old ResponseRelevancy scores still comparable?No. Cosine similarity is computed in whatever space the embedding model defines, so a different model shifts the numbers even with identical data and identical judge outputs. Treat an embedding-model change like a judge change: re-baseline the suite and annotate the point in your history where the instrument changed.
- How do you tell, without running anything, whether a metric list needs embeddings?Read each metric's definition for a similarity step. If the score comes from comparing vectors — as ResponseRelevancy's second phase does — an embedding model is required. If the score comes purely from parsing a judge model's verdict, as Faithfulness and the context metrics do, it is not.
saying these in an interview costs you the question
- Thinks every ragas metric needs an embedding model
- Thinks no ragas metric needs one, only a judge
- Assumes Faithfulness is computed by vector similarity
- Reuses the retriever's embedding model as the evaluator's
- Treats an embedding-model swap as score-neutral