skip to content

Why does a taint-tracking scanner report unexploitable findings and still miss real ones?

level: middleimportance: must knowfreq 60%

answer

  1. It reasons about possible, not feasible
  2. Two mirror-image causes
  3. Undeclared helpers cut both ways
  4. Run-time wiring is invisible to it
  5. Budgets and timeouts truncate long paths

basics

~20 s

It over-approximates flows it cannot prove infeasible, and it knows nothing about undeclared in-house validators, so it reports paths that are safe. It misses defects reached through undeclared sources, dynamic wiring, storage round-trips or truncated analysis budgets.

solid answer

~50 s

False positives come mostly from over-approximation and missing model entries: the engine reports possible rather than feasible paths, it is largely path-insensitive, and it has never heard of the team's own validation or encoding helper, so every flow through it is reported. Sink rules also fire in contexts that are not exploitable, such as query building in a migration script. False negatives come from the mirror image: an in-house entry point nobody declared as a source, call edges that only exist at run time through dynamic dispatch or container wiring, second-order flows where the value is written to a store in one place and read back in another, and per-function timeouts or depth limits that silently truncate long paths. Whole categories - authorisation and business-rule defects - have no source-to-sink signature at all. Treat both as model debt: fix the family, not the finding.

code

pseudocode · 8 lines
pseudocode
# service A - the write side
function saveTripNote(request):
    store.insert("trip_notes", request.body["note"])   # tainted value leaves the process

# service B - the read side, scanned separately
function renderTripNote(tripId):
    note = store.selectNote(tripId)      # not modelled as a source
    return renderMarkup("<p>" + note + "</p>")   # sink reached, path never seen

go deeper

for a junior

Be ready to say that a report is a hypothesis, not a verdict, and that the two error kinds have names: a false positive is a report that is not a defect, a false negative is a defect that was not reported.

for a middle

This is your level's core: name concrete causes on both sides - over-approximated paths, undeclared validators, undeclared entry points, run-time wiring, storage round-trips, analysis timeouts - and explain the fix as a model change.

for a senior

Demonstrate triage discipline on a real backlog: cluster findings by rule and by root cause, close families with model edits, and quantify what you are still blind to instead of reporting a clean run as an all-clear.

for a principal

Own the tradeoff explicitly. Decide which error kind your program can afford at each tier, fund the model work that raises precision permanently, and document the categories this technique is not expected to cover along with their compensating controls.

### The two error kinds, named precisely A **false positive** is a reported finding that is not a real defect: the path cannot carry attacker-controlled data, or the sink is not exploitable with it. A **false negative** is a real defect the analyzer did not report. Precision is the share of reports that are true; recall is the share of true defects that are reported. A security scanner is a classifier and, like every classifier, it trades one against the other. What makes this a *middle-level mechanics* question rather than a slogan is that each error kind has identifiable causes you can name and act on. ### Why unexploitable findings appear **Over-approximation of the flow graph.** To be tractable, the engine reasons about *possible* paths, not *feasible* ones. If a value can reach a sink along some route in the graph, that route is reported, even when a condition earlier in the program means the two branches can never both be taken. Most engines are **path-insensitive** or only partially path-sensitive precisely because tracking feasibility of every branch combination explodes. **Unmodelled sanitizers.** The largest single source of noise in a mature codebase. The team has its own validation or encoding helper; the engine has never heard of it; every flow through it is reported. Until that helper is declared as a barrier for the relevant sink family, the same finding comes back on every scan. **Sinks that are not sinks here.** A query-building call in a migration script, a command launcher in a developer tool, markup assembly in a template that is never served - the rule fires on the shape without knowing the deployment context. **Values that are attacker-*influenced* but not attacker-*controlled*.** A source model may mark a whole request object tainted, when in practice a value has already been narrowed to one of a small set upstream in a layer the engine cannot see. **Aggregation artefacts.** The same underlying defect reported once per reaching path can look like nine findings, which inflates the apparent noise even when every report is technically true. ### Why real defects are still missed **Unmodelled sources.** As above, inverted: the custom decoder, the internal transport, the configuration loader that nobody declared. The engine cannot flag what it never considered tainted. **Dynamic dispatch, reflection and configuration-driven wiring.** If the concrete implementation is selected at run time, the call graph the analyzer built may not contain the edge the defect travels along. Code generated at build time and code injected by a framework's container are the classic gaps. **Second-order flows.** The attacker's value is written to a store, a queue or a cache in one service and read back somewhere else. The write and the read are in different analysis units, so the path is cut in half and neither half reaches a sink. **Budget cuts.** Real engines impose depth limits, context-sensitivity limits and per-function timeouts. On a large module those cuts silently truncate long paths, and the report is a *partial* answer presented with the same confidence as a complete one. **Whole categories outside the technique.** Authorisation gaps, business-rule mistakes and design flaws have no source-to-sink signature at all; a taint engine's recall against them is near zero by construction, which is why a clean scan is evidence about one class of defect and about nothing else. ### Soundness, completeness, and the honest answer In analysis terms, **sound** means "reports every real defect" (no false negatives) and **complete** means "reports only real defects" (no false positives). Deciding the general question exactly is not possible, so every practical tool gives up at least one - and shipping engines quietly give up both, trading soundness for a run time and a signal-to-noise ratio that humans will tolerate. A candidate who claims their scanner is sound has usually mistaken *aggressive* for *exhaustive*. ### What a strong answer does about it Treat both error kinds as **configuration debt, not tool defects**. When you close a false positive, ask what model change would prevent the whole family: declare the team validator as a sanitizer, narrow a sink rule to the calling context, mark an internal accessor as a source. When you close a finding as a real bug, ask what other route the same source could take. Keep triage verdicts attached to the code, so that a re-scan does not re-ask a question that has been answered. A concrete example from a ride-hailing dispatcher: an 11-person team's first full scan produced 213 findings on the trip-search module. 26 were real. Of the remaining 187, 141 were the same undeclared zone-validation helper appearing on many paths - one model line closed all of them - and the rest were query-building in migration scripts. Meanwhile, the defect that later mattered lived in a value read back out of the ride-history store, written by an earlier request: a second-order flow the scan never had the chance to see. That pair of outcomes is the shape of the whole answer.

  • What does it mean for a static analysis to be sound, and are shipping security scanners sound?
    Sound means it reports every real defect in the class it targets - no false negatives - while complete means it reports only real ones. Deciding the general question exactly is not possible, so a tool must give something up. In practice shipping scanners give up both: they cut depth, context and path feasibility to keep run time and noise tolerable. A candidate who claims their scanner is sound has usually mistaken aggressive rules for exhaustive ones.
  • You close a false positive. What should you change so the same family does not come back?
    Change the model rather than the verdict. If an in-house validator produced the flow, declare it as a sanitizer for that sink family and the whole family closes at once. If the sink is not exploitable in that context, narrow the rule to the contexts where it is. If nothing generalises, record the suppression with a written justification attached to the code so a re-scan does not re-ask a question that has already been answered.
  • A scan of a service returns zero findings. What have you actually learned?
    That this technique found nothing in the code it managed to analyse under the rules it had. It is evidence about one class of defect - data reaching a modelled dangerous operation - and no evidence at all about authorisation gaps, business-rule mistakes or design flaws, which have no source-to-sink signature. Before believing it, check the analyzer's own coverage: files skipped, rules that failed to load, functions that hit the timeout.

It is like a map that shows every road but not which ones are closed today: it routes you through streets you cannot actually drive, and it omits the private lane that is nevertheless open to anyone who knows about it.

saying these in an interview costs you the question

  • Blames all noise on a bad tool rather than the model
  • Treats a clean scan as proof the service is secure
  • Claims a commercial scanner has no false negatives
  • Suppresses findings one by one with no justification
  • Thinks precision and recall can both be maximised
  • Ignores that generated and dynamically wired code is unseen

context