skip to content

AI Ops, Eval & Observability

Two related problems: scoring whether an LLM's output is good, and seeing what a live LLM or agent application actually did. Interviewers ask because "it worked in the playground" is not something you can ship, and both halves turn up in every production LLM postmortem.

on this pageshow

explore

questions

121 · 6 sections

In DeepEval, what is the difference between a Golden and an LLMTestCase?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A Golden holds only the static half of a test — input, expected_output, context — with no actual_output. An LLMTestCase adds the actual_output your application produced on this run, and that is what metrics score.

open as a page

What does evaluation_params control in a DeepEval GEval metric?

level: juniorimportance: must knowfreq 65%
basics
~20 s

evaluation_params is the list of LLMTestCaseParams members — INPUT, ACTUAL_OUTPUT, EXPECTED_OUTPUT, RETRIEVAL_CONTEXT and so on — naming which fields of the LLMTestCase the judge model is shown. Fields you leave out are invisible to the judge.

open as a page

In DeepEval, how do you score one LLMTestCase with a built-in metric?

level: juniorimportance: must knowfreq 78%
basics
~10 s

Build an LLMTestCase with at least input and actual_output, construct a metric such as AnswerRelevancyMetric(threshold=0.7), then call metric.measure(test_case). DeepEval then fills metric.score, metric.reason and metric.is_successful() for that one case.

open as a page

In DeepEval, what does assert_test() do inside a pytest test function?

level: juniorimportance: must knowfreq 70%
basics
~10 s

assert_test(test_case=..., metrics=[...]) scores the test case with every metric and raises AssertionError if any metric lands below its threshold. That turns an LLM evaluation into an ordinary pytest failure.

open as a page

What does DeepEval's Synthesizer.generate_goldens_from_docs actually do?

level: middleimportance: must knowfreq 62%
basics
~20 s

It reads the documents you point it at, splits and groups them into contexts, then asks an LLM to write synthetic inputs grounded in each context, rewrites them to be harder, and returns Golden objects — optionally with a generated expected_output.

open as a page

In Ragas, how do you run evaluate() over an EvaluationDataset of samples?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Ragas scores a whole dataset in one call. Put each record into a SingleTurnSample, wrap the list in an EvaluationDataset, then call evaluate(dataset=..., metrics=[...], llm=...). It returns an EvaluationResult carrying per-metric averages and a to_pandas() view.

open as a page

Which Ragas calls turn your own documents into a synthetic testset?

level: juniorimportance: must knowfreq 55%
basics
~10 s

Two calls. TestsetGenerator.from_langchain(llm, embedding_model) builds the generator from a LangChain model pair, then generator.generate_with_langchain_docs(docs, testset_size=N) returns a Testset of N synthetic samples with questions, reference contexts and reference answers.

open as a page

In Ragas, what do RunConfig's max_workers, timeout and max_retries control?

level: middleimportance: must knowfreq 58%
basics
~20 s

RunConfig tunes how a Ragas run executes: max_workers caps how many metric calls run concurrently (default 16), timeout bounds a single call (default 180 seconds), and max_retries is how many times a failed call is retried with backoff (default 10). Pass it as evaluate(run_config=RunConfig(...)).

open as a page

In Ragas, how do you give metrics an evaluator LLM and embedding model?

level: middleimportance: must knowfreq 78%
basics
~20 s

Wrap a LangChain chat model in LangchainLLMWrapper and a LangChain embeddings object in LangchainEmbeddingsWrapper, then pass them to evaluate() as llm= and embeddings=, or set them on individual metric objects. Ragas ships no model of its own.

open as a page

Which SingleTurnSample fields does each Ragas metric class require?

level: middleimportance: must knowfreq 68%
basics
~20 s

Each Ragas metric reads only the SingleTurnSample fields it declares. Faithfulness uses user_input, response and retrieved_contexts; LLMContextRecall and FactualCorrectness also need a human-written reference; ResponseRelevancy needs an embedding model in addition to a judge LLM.

open as a page

How do you create a LangSmith dataset and add examples with the Python SDK?

level: juniorimportance: must knowfreq 62%
basics
~10 s

Instantiate langsmith.Client, call client.create_dataset(dataset_name=...) to get a Dataset, then client.create_examples(dataset_id=dataset.id, examples=[...]) where each example is a dict with an inputs dict and, usually, a reference outputs dict.

open as a page

What signature and return value does a custom LangSmith evaluator function need?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A LangSmith evaluator is an ordinary Python function that declares any subset of the arguments inputs, outputs and reference_outputs, and returns a dict such as {"key": "exact_match", "score": 1}. You pass it in the evaluators list of evaluate().

open as a page

How do you attach an end-user thumbs-down to a LangSmith run from your app?

level: juniorimportance: must knowfreq 52%
basics
~20 s

Capture the traced call's run id, then call Client.create_feedback with that run id, a feedback key such as "user_score", and a score. The key names the metric, the score is the number, and comment or correction can carry the user's words or the fixed answer.

open as a page

How do you enable LangSmith tracing for a plain Python function with @traceable?

level: juniorimportance: must knowfreq 78%
basics
~10 s

Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY in the environment, then decorate the function with @traceable from the langsmith package. Each call is logged as a run in the LANGSMITH_PROJECT project, nested under any traced caller.

open as a page

In LangSmith's evaluate(), what must the target function accept and return?

level: middleimportance: must knowfreq 74%
basics
~20 s

The target takes one argument — a single example's inputs dict — and returns a dict of its outputs. LangSmith calls it once per example in the dataset and records each call as a run inside one named experiment.

open as a page

In Langfuse, what does compile() do to a fetched prompt, and what placeholder syntax does it use?

level: juniorimportance: must knowfreq 60%
basics
~20 s

compile() substitutes values into a Langfuse prompt's double-curly-brace placeholders, such as {{question}}. A text prompt compiles to a plain string; a chat prompt compiles to a list of role/content message dicts, ready to pass straight to a model SDK.

open as a page

In Langfuse, how do prompt versions and labels decide what get_prompt() returns?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Every save of a Langfuse prompt creates a new immutable numbered version. Labels such as production and latest are movable pointers to one of those versions. langfuse.get_prompt("name") resolves the version labelled production; pass label= or version= to target a different one.

open as a page

In Langfuse, how do you attach a score to a trace or an observation?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A Langfuse score is a named value hung off a trace or off one observation inside it. From inside an active observation call score_current_span() or score_current_trace(); from anywhere else call score_trace(trace_id=...) with a trace id you stored earlier.

open as a page

How does Langfuse's @observe decorator trace a Python function?

level: juniorimportance: must knowfreq 70%
basics
~20 s

@observe wraps a Python function so each call becomes one Langfuse observation, named after the function, timed, with the arguments captured as input and the return value as output, nested automatically under whatever observation is already active.

open as a page

In Langfuse, what is a trace and what are the observations inside it?

level: juniorimportance: must knowfreq 74%
basics
~20 s

A Langfuse trace is one end-to-end unit of work, typically one request or agent run. Observations are the timed steps inside it: nested and typed, with GENERATION recording an LLM call and SPAN any other step.

open as a page

How do you turn on Helicone's response cache and control how long entries live?

level: juniorimportance: must knowfreq 60%
basics
~10 s

Send Helicone-Cache-Enabled: true on the request, and set the lifetime with a standard Cache-Control: max-age=<seconds> header. Caching only works through the Helicone proxy, because something has to answer in place of the provider.

open as a page

How do you route OpenAI traffic through Helicone's proxy, and which header authenticates it?

level: juniorimportance: must knowfreq 72%
basics
~10 s

Point the OpenAI client's base URL at Helicone's proxy host, https://oai.helicone.ai/v1, and send a Helicone-Auth header holding "Bearer" plus your Helicone API key. Your provider key still travels in the usual Authorization header.

open as a page

When would you choose Helicone's async logging over its proxy integration?

level: middleimportance: must knowfreq 62%
basics
~20 s

Async logging keeps the provider call direct and ships the log separately, so Helicone adds no latency and cannot take your feature down. Choose it when the request path must stay untouched; choose the proxy when you want gateway behaviour, not just logs.

open as a page

What breaks when Helicone's proxy is slow or unreachable, and how do you limit the blast radius?

level: seniorimportance: must knowfreq 58%
basics
~20 s

With the base URL pointed at the proxy, every model call goes through it, so a degraded proxy degrades the feature itself — calls slow down, hang or fail. Limit the damage by making the base URL a flippable runtime setting, bounding timeouts, and moving latency-critical paths to async logging.

open as a page

In Helicone, what does the Helicone-Cache-Seed header change about cache hits?

level: middleimportance: should knowfreq 42%
basics
~20 s

Helicone-Cache-Seed partitions the cache into namespaces. Two identical requests sent with different seed values never share an entry, so you can isolate caches per user or tenant, and changing the seed instantly invalidates everything cached under the old one.

open as a page

What does Traceloop.init() do so existing OpenAI calls emit spans in OpenLLMetry?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Traceloop.init() starts OpenLLMetry's tracing pipeline and monkey-patches the LLM client libraries already installed in the process — OpenAI, Anthropic, LangChain, LlamaIndex, vector-store clients — so their calls emit spans. Your existing call sites stay unchanged.

open as a page

Which OpenTelemetry GenAI span attributes carry the model name and token counts?

level: juniorimportance: must knowfreq 68%
basics
~10 s

gen_ai.request.model records the model you asked for, gen_ai.response.model the one that answered, and gen_ai.usage.input_tokens plus gen_ai.usage.output_tokens the token counts. Standard names let any backend chart cost without reading your code.

open as a page

Why add OpenLLMetry's @workflow and @task decorators when init already traces model calls?

level: middleimportance: must knowfreq 66%
basics
~20 s

Auto-instrumentation only sees individual library calls. The decorators create named parent spans for your own multi-step logic, so a retrieval plus three model calls appears as one named operation you can time, compare and filter on rather than four unrelated spans.

open as a page

What does setting TRACELOOP_TRACE_CONTENT=false change in OpenLLMetry spans?

level: middleimportance: must knowfreq 58%
basics
~20 s

It stops OpenLLMetry recording prompt and completion message text on spans. Metadata still flows: provider, model, token counts, parameters, latency and errors. Content capture is on by default, so this is the switch a regulated team flips.

open as a page

When should you pass disable_batch=True to OpenLLMetry's Traceloop.init(), and what does it cost?

level: middleimportance: should knowfreq 48%
basics
~20 s

Pass disable_batch=True in short-lived processes — serverless handlers, CLI scripts, notebooks, tests — where the process can end before buffered spans are sent. It exports each span as it finishes, which is safe but adds export work on the calling path, so leave it off in long-running services.

open as a page