skip to content

How does a guardrail on a CrewAI Task work when validation fails?

level: seniorimportance: should knowfreq 45%

answer

  1. a callable that sees the finished output
  2. two-part return: accepted plus value
  3. rejection is not logging, it is a retry
  4. the feedback message is itself a prompt
  5. exhausted retries raise, they do not pass through

basics

~20 s

A Task guardrail is a callable that receives the finished TaskOutput and returns a two-part result: accepted plus a value, or rejected plus an error message. On rejection CrewAI re-runs the task with that message as feedback, up to the task's retry limit, then raises instead of returning bad output.

solid answer

~50 s

`Task(guardrail=fn)` installs a validation hook that runs after the agent produces its final answer. The callable takes the `TaskOutput` and returns a tuple: `(True, value)` accepts, and the returned value becomes the task's result — so a guardrail can transform as well as check. `(False, "message")` rejects, and CrewAI feeds that message back to the agent and re-runs the task, up to `max_retries` on the task (3 by default); when retries are exhausted it raises rather than passing a bad result downstream. That retry-with-feedback loop is the whole point — a plain assertion in your own code after the crew finishes cannot repair anything. Guardrails complement structured output rather than replacing it: `output_pydantic` fixes the shape, the guardrail enforces the semantics (length limits, banned content, a value that must exist in your database). Keep the callable deterministic and fast; recent CrewAI also accepts a plain-language guardrail string, which delegates the judgment to an LLM and costs a call per check.

code

python · 25 lines
python
from typing import Any

from crewai import Task
from crewai.tasks.task_output import TaskOutput


def under_250_words(output: TaskOutput) -> tuple[bool, Any]:
    text = output.raw.strip().removeprefix("```markdown").removesuffix("```").strip()
    words = len(text.split())
    if words > 250:
        return (
            False,
            f"The summary is {words} words but must be under 250. "
            "Cut the background section and keep the recommendation.",
        )
    return True, text  # accepted, and the fence-stripped text becomes the result


summarize = Task(
    description="Summarize the research findings for an executive reader.",
    expected_output="A markdown summary under 250 words ending with one recommendation.",
    agent=writer,
    guardrail=under_250_words,
    max_retries=2,
)

go deeper

for a junior

Know that a guardrail is a function attached to a task that inspects the finished output and can accept or reject it, and that rejection causes the task to be attempted again.

for a middle

State the tuple contract precisely — accepted plus a replacement value, or rejected plus a message — and explain that the message is fed back to the agent for a bounded number of retries before the run fails.

for a senior

Show the division of labour against expected_output and output schemas, write actionable rejection messages, and reason about cost: every rejection is another full agent loop, so a frequently firing guardrail is a prompt defect.

for a principal

Decide the policy — which tasks must fail loudly rather than return unchecked output, where LLM-judged criteria are worth their non-determinism, and how rejection rates are monitored as a quality signal across a fleet of crews.

## The gap guardrails fill An agent can produce output that is well-shaped and still unacceptable: a summary three times the requested length, a product code that does not exist, a recommendation citing a competitor you may not name, a JSON object whose fields are filled with `"N/A"`. Schema validation catches none of that. Checking it after the crew finishes catches it too late — you have paid for the whole run and have no mechanism to fix the offending task. A task guardrail closes that gap by putting the check *inside* the task's lifecycle, where a failure can still be turned into another attempt. ## The contract The callable receives the task's `TaskOutput` and returns a two-element tuple: - `(True, value)` — accepted. The value you return becomes the task's result, so the guardrail doubles as a normalization point: strip a markdown fence, coerce a string to a parsed object, trim whitespace, canonicalize units. - `(False, message)` — rejected. The message is not just logged; it is fed back to the agent as the reason its attempt was not accepted. Because the guardrail sees `TaskOutput`, it can inspect `raw` for the text form or `pydantic`/`json_dict` when the task also declares a schema. Validate against whichever form your rule is actually about. ## The retry loop On rejection CrewAI re-runs the task with the failure message included, giving the agent a concrete correction rather than a blind second attempt. This repeats up to the task's `max_retries` (default 3). If the guardrail still rejects after the last attempt, CrewAI raises — deliberately. The alternative, returning output that failed its own acceptance check, would poison every downstream task through context and land in your database looking legitimate. That failure behaviour has a design consequence: guardrails belong on tasks where being wrong is worse than failing. A crew that must always return something needs the guardrail paired with a caller-side fallback, not a guardrail that always passes. ## Writing good guardrails - **Deterministic and cheap.** The function runs on every attempt. Regexes, length checks, schema cross-field rules, and lookups against a fast local source are ideal. - **Actionable messages.** `"too long"` produces a random shortening; `"430 words, must be under 250; cut the background section"` produces a targeted rewrite. The message is a prompt, so write it like one. - **One concern per rule, several rules in one function.** Return on the first failure with a specific message rather than a list of everything wrong — the agent fixes what it is told most clearly. - **Do not re-check what the schema already guarantees.** If `output_pydantic` requires five findings, the model instance already has five; check that they are *distinct* and *sourced* instead. - **Beware external calls.** A guardrail that hits a network service adds its latency and failure modes to every attempt; a network error inside the guardrail is not a rejection of the model's work but it will look like one. ## LLM-judged guardrails Recent CrewAI also accepts a guardrail expressed as plain language rather than a callable, in which case a model judges whether the output satisfies the stated rule. This is attractive for criteria that resist code — tone, whether an answer actually addresses the question, absence of speculation. The trade is real: an extra model call per attempt, non-deterministic verdicts, and a judge that can be wrong in the same direction as the worker. Use code where a rule can be coded, and reserve the language form for genuinely subjective criteria on high-value tasks. ## Guardrail versus expected_output versus schema A useful way to hold the three apart: - `expected_output` — prose the agent reads. Steers effort and shape. Never enforced. - `output_pydantic` / `output_json` — structural contract. Enforced by conversion, cannot judge content. - `guardrail` — executable acceptance test with the power to retry. Judges content, and can transform. A well-built task uses all three, saying the same thing at three levels of enforcement. ## Operational notes Instrument the rejections. A guardrail that never fires is either perfect or dead, and the two look identical in logs; counting attempts per task tells you which. A guardrail that fires on most runs is a signal that the prompt or the model is wrong, not that the retry budget should go up — three failed attempts cost three full agent loops, which is the most expensive way to discover that `expected_output` was vague. ## What interviewers are checking That you know the tuple contract and that the accepted value replaces the result; that rejection means retry-with-feedback, bounded, then raise; that guardrails are semantics while schemas are shape; and that you have thought about the cost of a rule that fails often.

  • Your guardrail rejects on nearly every run and the task usually fails after its retries. Do you raise the retry limit?
    No. Three rejections mean three full agent loops paid for before the failure, so raising the budget multiplies the cost of the same wrong answer. Treat a high rejection rate as a prompt or model defect: tighten `expected_output` to state the rule the guardrail enforces, check the model is capable of it, and only then consider whether the rule itself is unreasonable.
  • Why put the check in a guardrail rather than validating after the crew finishes?
    Because the guardrail runs inside the task's lifecycle, where a rejection becomes another attempt with the failure message as feedback. Post-run validation happens after every downstream task has already consumed the bad output through context, and leaves you with nothing to do but discard the entire run or write your own repair loop.
  • What does returning (True, value) let you do beyond accepting?
    It replaces the task's result with whatever you return, so the guardrail is also a normalization point — strip a stray markdown fence, parse the string into an object, canonicalize units or casing. Downstream tasks and your own code then see the cleaned form rather than the model's raw phrasing, and the cleanup lives in testable Python instead of in prompt wording.
  • When is a plain-language guardrail preferable to a Python callable?
    Only for criteria that resist code — tone, whether the answer actually engages the question, absence of speculation. It costs an extra model call on every attempt, gives non-deterministic verdicts, and the judge can share the worker's blind spots. Anything expressible as a length check, regex, cross-field rule or lookup should be code, which is faster, free and reproducible in tests.

saying these in an interview costs you the question

  • Thinking a failed guardrail only logs a warning
  • Assuming rejection retries forever until it passes
  • Using a guardrail to enforce shape that the schema already guarantees
  • Returning a vague message like 'invalid output' as feedback
  • Raising the retry limit to fix a guardrail that fails constantly

context