skip to content

What are personas in Ragas testset generation, and why set them?

level: middleimportance: nice to knowfreq 28%

answer

  1. Controls whose voice asks
  2. Name plus a role description
  3. Corpus vocabulary is not user vocabulary
  4. Auto-derived from the graph otherwise
  5. Ground them in real query logs

basics

~20 s

A Persona has a name and a role_description, and it tells the query synthesizers whose voice to write in. Passing persona_list to TestsetGenerator makes generated questions read like your real user segments; leave it out and Ragas invents personas from the knowledge graph.

solid answer

~50 s

Personas control *who* is asking. A `Persona` object carries a `name` and a `role_description`, and the query synthesizers use it when phrasing a question, so the same underlying node can yield "how do I turn on two-factor auth?" for a new customer and "where is the MFA enforcement policy documented?" for a compliance reviewer. You supply them as `persona_list` on `TestsetGenerator`. If you do not, ragas derives personas from the knowledge graph itself with an LLM, via `generate_personas_from_kg`. The reason to set them by hand is realism: auto-derived personas reflect what the *corpus* is about, not who your users are, and a testset whose phrasing does not resemble real queries measures retrieval against language your users never use. If you have segments — trial user, admin, support agent, auditor — encoding them costs two lines and makes the generated set noticeably closer to production traffic.

code

python · 25 lines
python
from ragas.testset import TestsetGenerator
from ragas.testset.persona import Persona

personas = [
    Persona(
        name="new customer",
        role_description=(
            "Signed up this week, has configured nothing yet, and uses "
            "everyday words rather than product terminology."
        ),
    ),
    Persona(
        name="compliance reviewer",
        role_description=(
            "Checks whether data-handling claims are documented and cites "
            "policy names precisely."
        ),
    ),
]

generator = TestsetGenerator(
    llm=wrapped_llm,
    embedding_model=wrapped_embeddings,
    persona_list=personas,
)

go deeper

for a junior

Know that a Persona has a name and a role_description and that supplying a persona_list makes the generated questions sound like specific kinds of users rather than one generic voice.

for a middle

Explain that personas affect phrasing and assumed knowledge, not difficulty, and that ragas derives personas from the knowledge graph when you do not supply any.

for a senior

Argue the realism case: generated questions inherit corpus vocabulary, which inflates retrieval scores relative to real traffic, and hand-written personas grounded in query logs are the cheapest correction available.

for a principal

Own whether persona definitions are a shared, versioned artefact across the organisation's eval sets, and how they stay aligned with the user segments product and support actually track.

## The gap personas close Generated questions have a characteristic problem: they read like a documentation author quizzing themselves. They use the corpus' own vocabulary, they are grammatically complete, and they assume the reader already knows what the product calls things. Real queries are shorter, use different words for the same concepts, and come from people with different amounts of context. If your eval set is written in corpus vocabulary and your users are not, retrieval scores on the eval set overstate what production retrieval achieves — the eval queries are lexically closer to the documents than real queries ever are. Personas are ragas' lever on that. They do not change *which* nodes a synthesizer draws or how many hops it takes; they change the framing and vocabulary of the question written from those nodes. ## The object `Persona` lives in the testset package and carries two fields: `name`, a short label, and `role_description`, a sentence or two describing who this person is, what they know, and what they are trying to do. The description is the part that does the work — it goes into the synthesizer's prompt, so specificity pays. "A new customer" produces less differentiation than "Signed up this week, has not configured anything yet, uses everyday words rather than product terminology." You attach a list of them at construction: `TestsetGenerator(llm=..., embedding_model=..., persona_list=[...])`. Personas are then distributed across the generated samples, so a testset of thirty questions over four personas gives you roughly seven or eight questions in each voice. ## The automatic path If `persona_list` is not supplied, ragas generates personas for you from the knowledge graph using the LLM — `generate_personas_from_kg` inspects the graph and proposes a small set of plausible users of that content. This is a reasonable default and it means the feature is never a blocker, but understand what it is doing: it infers users from documents. A corpus of Kubernetes runbooks will produce personas like "platform engineer" and "on-call responder" whether or not those are the people who actually query your system. If your real traffic is 80% first-line support staff who have never touched a cluster, the auto-derived personas will systematically miss the register your production queries are written in. ## When it is worth doing by hand The cost-benefit is clear-cut when you know your segments. Two or four hand-written personas take minutes and make the testset materially more representative. It is worth the effort when: - Your users span very different expertise levels over the same content — the classic novice/expert split, where the novice does not know the vocabulary the documents use. - Different segments ask about different slices of the corpus, and you want coverage weighted the way traffic is. - You have production query logs. This is the strongest case: read fifty real queries, and write personas that describe the people who plausibly wrote them. That grounds the synthetic set in observed reality rather than imagination. It is not worth the effort when the corpus serves a single homogeneous audience — an internal API reference read only by the team that wrote it — because in that case one voice is the truth and the auto-derived persona will land close enough. ## What personas do not fix They change phrasing, not difficulty. A persona cannot turn a single-hop question into a multi-hop one; that is the synthesizer's job and the graph's structure. Nor do they fix the deeper limitation of any generated set, which is that the questions are written *from* chunks known to exist in the corpus. Personas make the phrasing more realistic; they do not make the questions more like the ones your corpus cannot answer. ## Verifying they took effect Read the output. Generate a small set with and without your persona list and compare the `user_input` column side by side — if the two sets read the same, your `role_description` strings are too generic to differentiate, and the fix is to make them concrete about vocabulary and prior knowledge rather than about job title.

  • What happens if you never pass persona_list to TestsetGenerator?
    Ragas derives personas from the knowledge graph with the LLM rather than failing or falling back to a single anonymous user. The generated personas describe plausible readers of that content, which is a sensible default — but it infers users from documents, so it can miss the actual register of your traffic when your users are less expert than the corpus assumes.
  • Do personas change how hard the generated questions are?
    No. Difficulty comes from which synthesizer runs and whether the graph has the relationships a multi-hop question needs. A persona changes vocabulary, framing and assumed prior knowledge — the same node pair phrased for a novice and for an expert is the same retrieval problem, just worded differently.
  • How would you write personas if you already have production query logs?
    Read a sample of real queries and cluster them by who plausibly wrote them, then turn each cluster into a role_description that captures its vocabulary and assumed knowledge, not just a job title. That grounds the synthetic set in observed language and is the single highest-value thing you can do to close the gap between generated and real queries.

saying these in an interview costs you the question

  • Thinking personas change question difficulty or hop count
  • Assuming generation fails without a persona list
  • Writing role descriptions that are job titles with no vocabulary detail
  • Believing auto-derived personas reflect real user segments
  • Expecting personas to fix questions the corpus cannot answer

context