In Ragas, what does a metric's single_turn_ascore() return and why await it?
answer
- one metric, one sample, one float
- the leading a means async
- calling is not running
- coroutine object, never awaited
- asyncio.run outside async code
basics
~20 ssingle_turn_ascore() scores one SingleTurnSample with one metric and resolves to a float. It is asynchronous, so calling it without await or asyncio.run() hands back a coroutine object that never executes and never contacts the judge model.
solid answer
~50 s`single_turn_ascore(sample)` is the single-sample entry point on a Ragas single-turn metric. You pass one `SingleTurnSample` and it resolves to a float score for that sample under that metric — no dataset, no results table, just the number. It is defined as a coroutine because scoring makes network calls to the judge model, so it must be awaited inside an async function, or wrapped in `asyncio.run(...)` from synchronous code. Forgetting that is the classic first-run mistake: you get a coroutine object back, printing it shows `<coroutine object ...>` rather than a number, Python warns that it was never awaited, and no judge call was ever made. Use it when you are debugging one sample or writing a quick sanity check on a newly configured metric; for a whole dataset you drive the metrics through the library's dataset-level evaluation instead, which handles concurrency for you.
code
python · 18 linesimport asyncio
from langchain_openai import ChatOpenAI
from ragas import SingleTurnSample
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import Faithfulness
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))
metric = Faithfulness(llm=evaluator_llm)
sample = SingleTurnSample(
user_input="What is the capital of France?",
response="Paris is the capital of France.",
retrieved_contexts=["Paris has been the capital of France since 987."],
)
score = asyncio.run(metric.single_turn_ascore(sample))
print(round(score, 3))go deeper
Know that single_turn_ascore takes one sample, returns a float score, and is async — so you await it or wrap it in asyncio.run, or nothing runs at all.
Be ready to explain why the API is async in the first place: scoring is a network call to a judge model, and the library is built so a dataset run can have many in flight at once.
Show that you use it as a debugging instrument — isolating a mis-shaped sample from a mis-wired judge from a genuine disagreement — and that dataset-scale scoring goes through the library's evaluation path, not a hand-rolled loop.
Own the convention in your codebase: single-sample scoring belongs in tests and triage, dataset evaluation belongs in the pipeline, and mixing them is how teams end up with slow bespoke harnesses that quietly diverge from the library's behaviour.
## What the method is for Ragas has two granularities. Dataset-level evaluation runs many metrics over many samples and gives you a results object you can turn into a table. `single_turn_ascore` is the other end: one metric, one sample, one number. That granularity matters more than it sounds. When a metric is behaving strangely, running it against one hand-built sample where you know the right answer is the fastest way to separate three different problems — a mis-shaped sample, a mis-wired judge model, or a metric that genuinely disagrees with you. It is also how you sanity-check a newly written AspectCritic definition before spending money running it over a dataset. ## Why it is a coroutine Scoring is not local computation. Every Ragas metric calls out to a judge LLM and, for some metrics, an embedding model. Those are network round-trips, and the library is built async so that a dataset run can have many of them in flight at once rather than serially. The consequence for the single-sample method is the standard Python async rule: calling a coroutine function does not run its body. It constructs a coroutine object and returns immediately. Nothing executes until something awaits it or an event loop drives it. So from an async context: score = await metric.single_turn_ascore(sample) And from ordinary synchronous code — a script, a REPL: score = asyncio.run(metric.single_turn_ascore(sample)) In a Jupyter notebook there is already a running event loop, so `asyncio.run` raises about the loop already running; inside a notebook cell you simply `await` the call directly. ## What going wrong looks like The symptom is unmistakable once you have seen it. `print(metric.single_turn_ascore(sample))` prints something like `<coroutine object ...>` instead of `0.83`. Python emits a `RuntimeWarning` that the coroutine was never awaited. Crucially, your judge model was never called — there is no cost, no latency and no error, which is why people misread it as "the metric returned something weird" instead of "the metric never ran". A related confusion: the returned value is a plain float, not an object with attributes. If you want the judge's reasoning rather than just the number, that is a different surface — the single-sample score gives you the score. ## What must be in place first Two preconditions, and both produce failures that are easy to misattribute. First, the metric needs its judge model. A metric constructed without an LLM has nothing to call, so scoring fails rather than returning a default. Metrics that also need embeddings need both supplied. Second, the sample must carry the fields that metric reads. A `Faithfulness` scored against a sample with no `retrieved_contexts` cannot do its job; a reference-requiring metric against a sample with no `reference` likewise. The failure surfaces at scoring time, not at sample construction time, because the sample object itself accepts any subset of fields. ## When not to use it Do not build your own loop over a dataset calling `single_turn_ascore` sample by sample. You will get serial network calls and none of the concurrency, retry or result-collection behaviour the library's dataset-level evaluation provides. The single-sample method is a debugging and unit-checking tool; the dataset path is the production one. The naming convention is worth internalising because it recurs: the leading `a` marks the async variant, and the `single_turn` prefix marks which sample shape it accepts. A metric that scores multi-turn conversations exposes the corresponding multi-turn method instead, and passing the wrong sample shape to either is a straightforward type error rather than a silently wrong score.
- What exactly do you see if you forget to await it?A coroutine object where you expected a float — printing gives `<coroutine object ...>` — plus a RuntimeWarning that the coroutine was never awaited. No judge call is made, so there is no cost and no error from the provider. That combination is why it reads as a strange return value rather than as code that never ran.
- Why is asyncio.run the wrong call inside a Jupyter notebook?A notebook kernel already has a running event loop, and `asyncio.run` insists on creating and owning one, so it raises about the loop already running. Inside a notebook cell you await the coroutine directly, since modern cells support top-level await. This trips people whose script worked and whose notebook copy of the same two lines does not.
- Why not just loop over your dataset calling single_turn_ascore yourself?Because you would serialise every judge call and reimplement, badly, what the library's dataset-level evaluation already does — running samples concurrently, applying retries and timeouts, and collecting results into a table. The single-sample method exists for debugging one case or checking a freshly written criterion, not as the scaling path.
saying these in an interview costs you the question
- Calling single_turn_ascore without await or asyncio.run
- Expecting a rich result object instead of a float
- Using asyncio.run inside a Jupyter notebook cell
- Looping it over a whole dataset by hand
- Scoring with a metric that has no judge model configured