skip to content

Code Scanning and CodeQL

Static analysis wired into pull requests: CodeQL (or any SARIF-producing scanner) annotates the diff and can gate the merge. Expect questions about triage discipline — a scanner that fires mostly false positives gets muted, and then it protects nothing.

part ofGitHuboverview, primer and where to startread it →
on this pageshow

questions

5

What is the difference between CodeQL default setup and advanced setup on a GitHub repository?

level: middleimportance: must knowfreq 58%

answer

  1. One is a setting, one is a file
  2. Who owns the build step?
  3. Custom queries and path filters need one of them
  4. Mutually exclusive per repository

basics

~20 s

Default setup is a GitHub-managed configuration enabled from repository settings with no workflow file: GitHub picks languages, queries and triggers. Advanced setup commits a workflow you own that calls the CodeQL action, so you control build, queries, packs and paths.

solid answer

~50 s

**Default setup** is switched on in the repository's code security settings. There is no file in the repository — GitHub detects the languages, runs the default query suite, and chooses when to scan (pushes and pull requests against the default branch, plus a periodic run). GitHub also keeps the CodeQL version current. It is the right choice for the long tail of repositories where nobody will maintain a workflow. **Advanced setup** commits `.github/workflows/codeql.yml`, which runs `github/codeql-action/init`, then a build (`autobuild` or your own steps), then `github/codeql-action/analyze`. You now own everything: language list, `build-mode`, query suite (`security-extended`, `security-and-quality`), custom query packs, `paths-ignore`, runner choice, and the analysis `category`. You move to advanced when the build is non-standard, when you need custom queries, packs or path filtering, or when you must run on self-hosted runners. The two are mutually exclusive per repository — enabling advanced setup turns default setup off.

code

yaml · 26 lines
yaml
name: CodeQL

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]
  schedule:
    - cron: '0 3 * * 1'

jobs:
  analyze:
    runs-on: ubuntu-latest
    permissions:
      security-events: write
      contents: read
    steps:
      - uses: actions/checkout@v4
      - uses: github/codeql-action/init@v3
        with:
          languages: java-kotlin
          config-file: .github/codeql/codeql-config.yml
      - uses: github/codeql-action/autobuild@v3
      - uses: github/codeql-action/analyze@v3
        with:
          category: java-kotlin

go deeper

for a junior

Recall that scanning can be switched on from repository settings with no file, and that the alternative is a workflow in .github/workflows that runs the CodeQL action.

for a middle

Explain what default setup decides for you — languages, query suite, triggers, updates — and name concrete reasons to move to advanced setup, such as a custom build or a custom query pack.

for a senior

Show you know the failure mode: with a compiled language, an analysis whose build does not compile the real code produces a green check over a near-empty database.

for a principal

Frame it as coverage strategy across a portfolio — default setup by policy for breadth, advanced setup only where the build or query set justifies the maintenance cost.

## Two ways to configure the same engine CodeQL analysis on GitHub always does the same thing: extract a CodeQL database from the code, run queries against it, upload SARIF. What differs between default and advanced setup is *who owns the configuration*. ## Default setup Default setup is enabled from the repository's code security settings (and can be applied in bulk from organisation-level settings, including automatically for newly created repositories). Its defining property is that **nothing is committed to the repository**. There is no `.github/workflows` file to review, no YAML to drift, and no pull request needed to turn it on. GitHub takes the decisions you would otherwise write down: - **Languages** — detected from the repository, with the option to adjust the selected set. - **Queries** — the default query suite, tuned for a low false-positive rate. - **Triggers** — analysis on pushes to the default branch, on pull requests targeting it, and on a periodic schedule. - **Build** — for compiled languages, GitHub uses a build mode that does not require you to describe a build, where the language supports it. - **Maintenance** — the CodeQL version and query set are updated by GitHub. The cost is control. You cannot add a custom query pack, exclude a generated-code directory, pin to a self-hosted runner, or run a bespoke build. ## Advanced setup Advanced setup generates a starter workflow — conventionally `.github/workflows/codeql.yml` — and from that point it is an ordinary workflow you own and review like any other file. Its skeleton is three CodeQL action steps: - `github/codeql-action/init` — declares `languages`, optionally a `build-mode`, and either inline `queries` or a `config-file` pointing at something like `.github/codeql/codeql-config.yml`. - a build — `github/codeql-action/autobuild` for a guessable build, or your own build commands when autobuild cannot work out how to compile the project. - `github/codeql-action/analyze` — runs the queries and uploads SARIF; `category` labels this analysis so several analyses of the same repository do not overwrite each other. The job needs `security-events: write` to upload results, plus `contents: read` for the checkout. ## When advanced setup is actually required Be specific in an interview; "more control" is not an answer. Advanced setup is what you need when: - **The build is non-standard.** A compiled language whose build needs a specific toolchain, code generation, a private artifact repository, or environment variables. CodeQL only sees code that the build compiles, so a wrong build silently yields an empty or partial database. - **You want a different query suite or extra packs.** `security-extended` widens coverage at the cost of precision; `security-and-quality` adds maintainability queries. Custom or third-party CodeQL packs are referenced through `packs:` in a config file. - **You need path filtering.** `paths-ignore` for vendored, generated or test code keeps the alert list honest. - **Runner requirements.** Self-hosted runners for a private toolchain, or larger runners because analysis of a big codebase is memory-hungry. - **Tuning noise.** `query-filters` in the config file lets you exclude a specific rule id that is wrong for your codebase. ## Switching between them A repository uses one or the other, not both: enabling advanced setup disables default setup, and re-enabling default setup stops the workflow's results being used as the configuration. A common migration is the reverse of what people expect — teams start with a hand-written workflow copied from another repository, discover nobody maintains it, and move the low-value repositories back to default setup, keeping advanced setup only where the build or the query set demands it. ## The judgment an interviewer is testing The good answer is portfolio-shaped: default setup everywhere, by policy, so coverage is not gated on someone writing YAML; advanced setup on the handful of repositories that genuinely need a custom build or query set. The bad answer is either extreme — hand-written workflows in three hundred repositories that nobody keeps current, or default setup on a compiled service whose real build never runs, producing a reassuring green check over an almost empty database.

  • What does the category input on github/codeql-action/analyze do?
    It labels the analysis so GitHub knows which previous results this upload replaces. Without distinct categories, two analyses of the same repository — a second language, or a different tool — overwrite each other's results, and alerts appear to vanish and reappear between runs.
  • Why would you not switch every repository to security-extended?
    The extended suite adds queries with lower precision, so it raises coverage and false positives together. On a codebase nobody is currently triaging, that reliably converts a usable alert list into background noise that gets muted. Widen the suite where there is someone to triage the extra findings.
  • Which permission does the analysis job need to publish results?
    security-events: write, which is what authorises the SARIF upload to code scanning, plus contents: read for the checkout. A workflow whose token has been narrowed elsewhere in the repository will run the analysis fine and then fail at the upload step.

saying these in an interview costs you the question

  • Claiming default setup commits a workflow you can edit
  • Thinking both setups can run on one repository at once
  • Believing advanced setup is always the better choice
  • Assuming autobuild works for any compiled project
  • Saying default setup lets you add custom query packs

context

open as a page

In GitHub, what is code scanning, and where do CodeQL results show up on a pull request?

level: juniorimportance: should knowfreq 52%

basics

~20 s

GitHub code scanning stores static-analysis findings as repository alerts. On a pull request, findings on lines the diff touches appear as inline annotations plus a code scanning results check; every alert, new or old, lists under the repository's Security tab.

open as a page

How do you get results from a non-CodeQL scanner into GitHub code scanning?

level: middleimportance: should knowfreq 41%

basics

~10 s

Have the tool emit SARIF 2.1.0 and upload it — the github/codeql-action/upload-sarif action, or the code scanning SARIF API from external CI. The job needs security-events: write, and each tool needs its own category.

open as a page

How do you make GitHub code scanning block a merge, and what breaks when you do?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Require the code scanning results check in branch protection, or add the code scanning rule to a GitHub ruleset with alert-severity thresholds. Both stall pull requests when the analysis is skipped by path filters, cannot run from a fork, or never reports.

open as a page

CodeQL is enabled across your org and produced thousands of alerts nobody triages. How do you make it trustworthy?

level: principalimportance: should knowfreq 38%

basics

~20 s

Separate the historical backlog from new findings: gate only on alerts a pull request introduces, tune the configuration so recurring false positives stop firing, require a dismissal reason plus comment, and measure the false-positive rate per rule rather than the raw alert count.

open as a page