AI Ops, Eval & Observability
Two related problems: scoring whether an LLM's output is good, and seeing what a live LLM or agent application actually did. Interviewers ask because "it worked in the playground" is not something you can ship, and both halves turn up in every production LLM postmortem.
on this pageshowhide
explore
- DeepEval24 questions
- Built-in Metrics6 questions
- G-Eval & Custom Metrics6 questions
- Pytest Integration & CI6 questions
- Datasets & Synthesizer6 questions
- Ragas24 questions
- Metric Classes6 questions
- Running Evaluations6 questions
- Testset Generation6 questions
- LLM Wiring & Integrations6 questions
- LangSmith24 questions
- Tracing & Runs6 questions
- Datasets & Experiments6 questions
- Evaluators6 questions
- Online Eval & Annotation6 questions
- Langfuse25 questions
- Traces & Observations7 questions
- Scores, Evals & Datasets6 questions
- Prompt Management6 questions
- Self-Hosting & Data Controls6 questions
- Helicone12 questions
- Proxy & Async Logging6 questions
- Caching, Rate Limits & Retries6 questions
- OpenLLMetry12 questions
- SDK & Instrumentation6 questions
- GenAI Semantic Conventions6 questions
questions
121 · 6 sectionsIn DeepEval, what is the difference between a Golden and an LLMTestCase?
basics
~20 sA Golden holds only the static half of a test — input, expected_output, context — with no actual_output. An LLMTestCase adds the actual_output your application produced on this run, and that is what metrics score.
What does evaluation_params control in a DeepEval GEval metric?
basics
~20 sevaluation_params is the list of LLMTestCaseParams members — INPUT, ACTUAL_OUTPUT, EXPECTED_OUTPUT, RETRIEVAL_CONTEXT and so on — naming which fields of the LLMTestCase the judge model is shown. Fields you leave out are invisible to the judge.
In DeepEval, how do you score one LLMTestCase with a built-in metric?
basics
~10 sBuild an LLMTestCase with at least input and actual_output, construct a metric such as AnswerRelevancyMetric(threshold=0.7), then call metric.measure(test_case). DeepEval then fills metric.score, metric.reason and metric.is_successful() for that one case.
In DeepEval, what does assert_test() do inside a pytest test function?
basics
~10 sassert_test(test_case=..., metrics=[...]) scores the test case with every metric and raises AssertionError if any metric lands below its threshold. That turns an LLM evaluation into an ordinary pytest failure.
What does DeepEval's Synthesizer.generate_goldens_from_docs actually do?
basics
~20 sIt reads the documents you point it at, splits and groups them into contexts, then asks an LLM to write synthetic inputs grounded in each context, rewrites them to be harder, and returns Golden objects — optionally with a generated expected_output.
In Ragas, how do you run evaluate() over an EvaluationDataset of samples?
basics
~20 sRagas scores a whole dataset in one call. Put each record into a SingleTurnSample, wrap the list in an EvaluationDataset, then call evaluate(dataset=..., metrics=[...], llm=...). It returns an EvaluationResult carrying per-metric averages and a to_pandas() view.
Which Ragas calls turn your own documents into a synthetic testset?
basics
~10 sTwo calls. TestsetGenerator.from_langchain(llm, embedding_model) builds the generator from a LangChain model pair, then generator.generate_with_langchain_docs(docs, testset_size=N) returns a Testset of N synthetic samples with questions, reference contexts and reference answers.
In Ragas, what do RunConfig's max_workers, timeout and max_retries control?
basics
~20 sRunConfig tunes how a Ragas run executes: max_workers caps how many metric calls run concurrently (default 16), timeout bounds a single call (default 180 seconds), and max_retries is how many times a failed call is retried with backoff (default 10). Pass it as evaluate(run_config=RunConfig(...)).
In Ragas, how do you give metrics an evaluator LLM and embedding model?
basics
~20 sWrap a LangChain chat model in LangchainLLMWrapper and a LangChain embeddings object in LangchainEmbeddingsWrapper, then pass them to evaluate() as llm= and embeddings=, or set them on individual metric objects. Ragas ships no model of its own.
Which SingleTurnSample fields does each Ragas metric class require?
basics
~20 sEach Ragas metric reads only the SingleTurnSample fields it declares. Faithfulness uses user_input, response and retrieved_contexts; LLMContextRecall and FactualCorrectness also need a human-written reference; ResponseRelevancy needs an embedding model in addition to a judge LLM.
How do you create a LangSmith dataset and add examples with the Python SDK?
basics
~10 sInstantiate langsmith.Client, call client.create_dataset(dataset_name=...) to get a Dataset, then client.create_examples(dataset_id=dataset.id, examples=[...]) where each example is a dict with an inputs dict and, usually, a reference outputs dict.
What signature and return value does a custom LangSmith evaluator function need?
basics
~20 sA LangSmith evaluator is an ordinary Python function that declares any subset of the arguments inputs, outputs and reference_outputs, and returns a dict such as {"key": "exact_match", "score": 1}. You pass it in the evaluators list of evaluate().
How do you attach an end-user thumbs-down to a LangSmith run from your app?
basics
~20 sCapture the traced call's run id, then call Client.create_feedback with that run id, a feedback key such as "user_score", and a score. The key names the metric, the score is the number, and comment or correction can carry the user's words or the fixed answer.
How do you enable LangSmith tracing for a plain Python function with @traceable?
basics
~10 sSet LANGSMITH_TRACING=true and LANGSMITH_API_KEY in the environment, then decorate the function with @traceable from the langsmith package. Each call is logged as a run in the LANGSMITH_PROJECT project, nested under any traced caller.
In LangSmith's evaluate(), what must the target function accept and return?
basics
~20 sThe target takes one argument — a single example's inputs dict — and returns a dict of its outputs. LangSmith calls it once per example in the dataset and records each call as a run inside one named experiment.
In Langfuse, what does compile() do to a fetched prompt, and what placeholder syntax does it use?
basics
~20 scompile() substitutes values into a Langfuse prompt's double-curly-brace placeholders, such as {{question}}. A text prompt compiles to a plain string; a chat prompt compiles to a list of role/content message dicts, ready to pass straight to a model SDK.
In Langfuse, how do prompt versions and labels decide what get_prompt() returns?
basics
~20 sEvery save of a Langfuse prompt creates a new immutable numbered version. Labels such as production and latest are movable pointers to one of those versions. langfuse.get_prompt("name") resolves the version labelled production; pass label= or version= to target a different one.
In Langfuse, how do you attach a score to a trace or an observation?
basics
~20 sA Langfuse score is a named value hung off a trace or off one observation inside it. From inside an active observation call score_current_span() or score_current_trace(); from anywhere else call score_trace(trace_id=...) with a trace id you stored earlier.
How does Langfuse's @observe decorator trace a Python function?
basics
~20 s@observe wraps a Python function so each call becomes one Langfuse observation, named after the function, timed, with the arguments captured as input and the return value as output, nested automatically under whatever observation is already active.
In Langfuse, what is a trace and what are the observations inside it?
basics
~20 sA Langfuse trace is one end-to-end unit of work, typically one request or agent run. Observations are the timed steps inside it: nested and typed, with GENERATION recording an LLM call and SPAN any other step.
How do you turn on Helicone's response cache and control how long entries live?
basics
~10 sSend Helicone-Cache-Enabled: true on the request, and set the lifetime with a standard Cache-Control: max-age=<seconds> header. Caching only works through the Helicone proxy, because something has to answer in place of the provider.
How do you route OpenAI traffic through Helicone's proxy, and which header authenticates it?
basics
~10 sPoint the OpenAI client's base URL at Helicone's proxy host, https://oai.helicone.ai/v1, and send a Helicone-Auth header holding "Bearer" plus your Helicone API key. Your provider key still travels in the usual Authorization header.
When would you choose Helicone's async logging over its proxy integration?
basics
~20 sAsync logging keeps the provider call direct and ships the log separately, so Helicone adds no latency and cannot take your feature down. Choose it when the request path must stay untouched; choose the proxy when you want gateway behaviour, not just logs.
What breaks when Helicone's proxy is slow or unreachable, and how do you limit the blast radius?
basics
~20 sWith the base URL pointed at the proxy, every model call goes through it, so a degraded proxy degrades the feature itself — calls slow down, hang or fail. Limit the damage by making the base URL a flippable runtime setting, bounding timeouts, and moving latency-critical paths to async logging.
In Helicone, what does the Helicone-Cache-Seed header change about cache hits?
basics
~20 sHelicone-Cache-Seed partitions the cache into namespaces. Two identical requests sent with different seed values never share an entry, so you can isolate caches per user or tenant, and changing the seed instantly invalidates everything cached under the old one.
What does Traceloop.init() do so existing OpenAI calls emit spans in OpenLLMetry?
basics
~20 sTraceloop.init() starts OpenLLMetry's tracing pipeline and monkey-patches the LLM client libraries already installed in the process — OpenAI, Anthropic, LangChain, LlamaIndex, vector-store clients — so their calls emit spans. Your existing call sites stay unchanged.
Which OpenTelemetry GenAI span attributes carry the model name and token counts?
basics
~10 sgen_ai.request.model records the model you asked for, gen_ai.response.model the one that answered, and gen_ai.usage.input_tokens plus gen_ai.usage.output_tokens the token counts. Standard names let any backend chart cost without reading your code.
Why add OpenLLMetry's @workflow and @task decorators when init already traces model calls?
basics
~20 sAuto-instrumentation only sees individual library calls. The decorators create named parent spans for your own multi-step logic, so a retrieval plus three model calls appears as one named operation you can time, compare and filter on rather than four unrelated spans.
What does setting TRACELOOP_TRACE_CONTENT=false change in OpenLLMetry spans?
basics
~20 sIt stops OpenLLMetry recording prompt and completion message text on spans. Metadata still flows: provider, model, token counts, parameters, latency and errors. Content capture is on by default, so this is the switch a regulated team flips.
When should you pass disable_batch=True to OpenLLMetry's Traceloop.init(), and what does it cost?
basics
~20 sPass disable_batch=True in short-lived processes — serverless handlers, CLI scripts, notebooks, tests — where the process can end before buffered spans are sent. It exports each span as it finishes, which is safe but adds export work on the calling path, so leave it off in long-running services.