skip to content

How do you run GitHub Models prompt evaluations from a GitHub Actions workflow?

level: middleimportance: should knowfreq 35%

answer

  1. prompts become repo files
  2. one permission on the job
  3. CLI extension runs and scores
  4. pull request is the useful trigger
  5. every case spends shared quota

basics

~20 s

Store prompts as .prompt.yml files in the repository, grant the job models: read so the workflow token can call inference, install the gh-models CLI extension, and run gh models eval on those files so prompt changes are reviewed and scored like code.

solid answer

~50 s

The pieces are a **prompt file**, a **permission**, and the **CLI**. A `.prompt.yml` file checked into the repo captures the model, its parameters and the message template, plus test inputs and the evaluators to score responses — so a prompt becomes a reviewable artifact rather than a string buried in code. In the workflow, declare `permissions: models: read` (alongside `contents: read` for checkout) so the built-in `GITHUB_TOKEN` can reach inference without a stored secret. Then install the `github/gh-models` CLI extension and run `gh models eval <file>` on pull requests; you can gate the job on its result so a prompt edit that regresses quality is caught in review. The same CLI gives you `gh models list` and `gh models run` for ad-hoc calls. Keep evaluation runs small and on modest models — CI shares the account's rate-limit budget with everyone else.

code

yaml · 16 lines
yaml
name: prompt-eval
on: pull_request
permissions:
  contents: read
  models: read
jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: gh extension install github/gh-models
        env:
          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
      - run: gh models eval prompts/summarize.prompt.yml
        env:
          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}

go deeper

for a junior

Know that prompts can live as .prompt.yml files in the repo and that the gh-models CLI extension can run and evaluate them from a workflow.

for a middle

Explain the three moving parts — the prompt file, the models: read permission on the job, and gh models eval — and why the workflow token removes the need for a stored secret.

for a senior

Show that you treat eval scores as a relative signal on a noisy system: enough cases per behaviour, deterministic checks where possible, advisory rather than hard gates for borderline movement, and CI load kept off the team's shared quota.

for a principal

Own where prompt evaluation sits in the quality strategy — between unit tests and production monitoring — and what it costs in shared inference capacity once every team adopts it.

## The problem being solved Prompts are production logic that traditionally escapes every engineering control: no diff review, no tests, no regression signal. GitHub Models addresses that by making a prompt a file in the repository and giving CI a way to execute and score it. ## Prompt files A `.prompt.yml` file describes one prompt as data. It names the model to run against, the model parameters, and the messages — system and user, with placeholders for inputs. Alongside that it can carry test data (the sample inputs to run) and evaluators (the checks applied to each response). Because it is YAML in the repo: - a prompt change appears in a pull request diff and gets reviewed; - the model id and parameters are versioned with the code that uses them; - the same file drives the browser editor, the CLI and CI, so there is one definition rather than three drifting copies. ## The CLI `gh models` is an extension to the GitHub CLI, installed with `gh extension install github/gh-models`. The commands you will actually name in an interview: - `gh models list` — what the catalog offers. - `gh models run <model> "<prompt>"` — a one-off call from the terminal, useful for a quick sanity check or piping output into a script. - `gh models eval <file.prompt.yml>` — execute the prompt over its test data and apply its evaluators, reporting per-case results. The CLI authenticates with your `gh` login locally and with the workflow token in CI, so the same command works in both places. ## Wiring it into Actions Three requirements: 1. **Permissions.** Add `models: read` to the workflow or job `permissions` block. Remember the block is exhaustive for the scope it appears on — if you write it on the job, also list `contents: read` or the checkout step loses access. 2. **The token in the environment.** Expose `GITHUB_TOKEN` to the step (commonly as `GH_TOKEN`, which the GitHub CLI reads). No repository secret is needed, and the credential expires with the run. 3. **The trigger.** `pull_request` is the useful one: you want the signal while the change is reviewable. Some teams add a scheduled run to catch drift when nothing changed on their side. Gate the job on the evaluation outcome if you want a hard signal, or run it informationally and read the output in the job log while the practice beds in. ## Evaluation is comparison, not proof The honest framing for an interview: automated prompt evaluation gives you a *relative* signal — this prompt versus the previous one, this model versus that one, on your own cases. It is not proof of correctness. Model outputs vary between runs, so a small test set can flip on noise. Practical mitigations: keep a decent number of cases per behaviour you care about, prefer deterministic checks where the task allows one, treat borderline score movement as a prompt to look rather than an automatic block, and keep human review for the cases that matter most. ## Cost and quota discipline Every evaluation case is inference, and CI inference draws on the same account-scoped rate limits as every developer and every other job. A large test matrix across several flagship models, running on every push, will exhaust a prototyping allowance and throttle your colleagues. Sensible defaults: run on pull requests rather than every push, keep the routine matrix on small models and reserve larger ones for a scheduled or manual deeper run, and cap the case count. If evaluation becomes central to your process, that is the moment to consider paid usage or a dedicated endpoint so CI stops competing with humans. ## What this does not replace Evaluation in CI checks prompt behaviour on curated inputs. It does not replace application tests around the code that builds the prompt and parses the response, and it does not replace production monitoring of real traffic. Say that out loud: the strongest answer positions prompt evals as one layer, sitting between unit tests and observability.

  • Why keep prompts in .prompt.yml files rather than as string literals in application code?
    Because a file makes the prompt reviewable and versioned alongside the model id and parameters it depends on, and it gives one definition that the browser editor, the CLI and CI all execute. String literals hide prompt changes inside unrelated diffs, drift between environments, and cannot be run through an evaluation harness without extracting them first.
  • What stops prompt evaluation in CI from becoming a rate-limit problem for the whole team?
    Scope and scheduling. Rate limits are account-scoped, so a broad matrix on flagship models running per push will starve colleagues. Run on pull requests rather than every push, keep the routine matrix on small models, cap the number of cases, and move deeper comparisons to a scheduled or manual job — or onto paid capacity if evaluation becomes core to the process.
  • Should a failing prompt evaluation block the merge?
    Only for checks you trust to be stable. Model outputs vary run to run, so a small test set can flip on noise and a hard gate then trains people to bypass it. Block on deterministic checks and on large, clear regressions; report the rest as advisory output in the job log and let a reviewer judge borderline movement.

saying these in an interview costs you the question

  • Thinks a workflow can call models without declaring the models permission
  • Runs a large evaluation matrix on flagship models on every push
  • Treats evaluation scores as proof of correctness rather than a relative signal
  • Stores a personal access token as a secret instead of using the workflow token
  • Keeps prompts as inline strings and evaluates a copy that drifts

context