In security static analysis, what are a source, a sink and a sanitizer in taint tracking?
answer
- Follow the data, not the keyword
- Three declared roles per rule
- Where untrusted input enters
- Where attacker control changes behaviour
- A path with nothing neutralising on it
basics
~20 sA source is where untrusted input enters the program, a sink is a sensitive operation that must not receive it, and a sanitizer neutralises the data. A finding is a source-to-sink path with no sanitizer on it.
solid answer
~50 sTaint tracking asks a dataflow question rather than a text-matching one: can attacker-controlled data reach an operation where that control causes harm? A **source** is a modelled entry point for untrusted input - a request field, a queue message, an environment value. A **sink** is a modelled operation where attacker control changes what happens rather than just what is computed - text assembled into a store query, a value handed to a command launcher or an interpreter, an assembled file path, markup returned to a browser. A **sanitizer** is a declared step that removes taint for that sink family: an encoder for the target grammar, a strict allow-list validation, a parse that discards anything off-shape. The analyzer propagates taint through concatenation, copies and calls, and reports a **path** with its trace. All three are model declarations: an unmodelled source produces silence, an unmodelled sanitizer produces noise.
code
pseudocode · 5 linesfunction handleTripSearch(request):
zone = request.body["zone"] # SOURCE: modelled untrusted input
validateZone(request.body["city"]) # sanitizer, but on a different value
filter = buildFilter(zone) # propagator: taint copied into the object
return store.runQuery("zone = '" + filter.zone + "'") # SINK: query textgo deeper
Be ready to define all three words without hesitating and to say that a finding is a path, not a line. Knowing that the tool follows data across assignments and calls is the whole point of the question.
An interviewer expects the mechanics: how taint propagates through concatenation and function summaries, why sanitisation is tied to a specific sink family, and why an unmodelled source or barrier changes the result so drastically.
Show that you read the reported trace rather than the headline: check that the source is genuinely attacker-controlled, that the path really carries the value, and that the sink treats it as structure. Then fix the model, not the single finding.
Own the model as an asset. Decide who maintains source, sink and sanitizer declarations for in-house frameworks, and treat that catalogue as the thing that determines whether analysis results across the estate mean anything at all.
### The question a taint analysis is actually asking A security static-analysis engine that does taint tracking is not asking "does this file contain a dangerous call?" It is asking a dataflow question: **can a value that an attacker controls reach an operation where attacker control causes harm, without being made safe on the way?** Everything in the vocabulary follows from that one question. ### Source A **source** is a program location that introduces untrusted data. Typical sources are the fields of an inbound request, a message pulled off a queue, a command-line argument, an environment value, a file whose contents the attacker can influence, a row read back from a store that the attacker previously wrote. The engine does not discover sources by intuition; each one is declared in a rule model, usually as "the return value of this accessor is tainted" or "parameter 2 of this handler signature is tainted". A value that comes from a source is said to be **tainted**. The practical consequence is blunt: *an unmodelled source produces silence.* If a team builds its own request-decoding layer and no rule marks its accessors as sources, the analyzer will faithfully report nothing on code that is full of real problems. ### Sink A **sink** is an operation that must not be reached by tainted data, because attacker control of the argument changes what the operation does rather than merely what it computes. The recurring shapes are: text assembled into a query and handed to a data store; text handed to an interpreter or to an operating-system command launcher; a path assembled and handed to file access; a destination handed to an outbound request; markup assembled and returned to a browser; a serialized blob handed to a decoder that can instantiate types. Which sink a rule is about matters, because it determines what counts as safe - the neutralisation that makes a value safe for markup is not the neutralisation that makes it safe for a file path. ### Sanitizer A **sanitizer** (some engines say *cleanser* or *barrier*) is a node in the flow that removes taint. It is context-specific: an encoder for the target grammar, a strict allow-list validation, a type-narrowing parse that discards anything not matching a fixed shape, a lookup that replaces the attacker's value with an internal one. When the engine sees a flow pass through a declared sanitizer, the value stops being tainted from that point on and the path is not reported. Two mistakes cluster here. First, a check that only *rejects some* inputs is not a sanitizer for the analyzer unless it is declared as one - and often should not be, because a length check or a null-check removes no attack. Second, sanitisation is a property of the path, not of the file: a validator called on a different variable, or after the value has already reached the sink, does not clear the flow, and an engine worth its price will say so. ### Propagation, and why this is not text search Between source and sink, the analyzer must decide how taint moves. **Propagators** (or *summaries*) say that if an argument is tainted, the result is tainted: string concatenation, formatting, collection insertion and retrieval, mapping between record shapes, copying a field into a builder. Following those steps across function boundaries is **interprocedural** analysis, and it is what separates a real taint engine from a text-search rule that flags every occurrence of a sensitive call name. Doing it at scale usually means computing a reusable **summary** per function - "taint on parameter 1 flows to the return value" - so that callers can be analysed without re-analysing the callee every time. ### The finding A finding is therefore **a path**: a specific source, a specific sink, and the ordered steps between them, with no sanitizer on the path. Good engines report the whole trace rather than a single line, because the trace is what a reviewer needs in order to judge the finding. When you triage one, you are checking three things in order: is the source really attacker-controlled, does the path really carry the value (or does it drop it somewhere), and does the sink really treat it as structure rather than as data. ### A worked shape In a ride-hailing dispatcher, an inbound trip-search handler reads a `zone` field from the request body - the modelled source. It is copied into a filter object, joined into a text fragment by a helper, and the fragment is passed to a store-query call - the modelled sink. The path is four steps long across three functions. If the team's own zone allow-list check sits on that path and is declared as a sanitizer for that sink family, no finding is produced; if it is not declared, the engine reports a real path that a human then has to close by hand. That asymmetry - silence when a source is missing, noise when a sanitizer is missing - is the whole reason the model, not the tool brand, is what an interviewer wants you to discuss.
- Why does the same value need different sanitizers depending on which sink it reaches?Neutralisation is defined against a target grammar. Escaping for markup makes a value safe to place in a document but does nothing about a value used as a file path or as a destination address, where the dangerous characters and the parsing rules are entirely different. A taint rule therefore ties each sanitizer to a sink family, and an engine that treated one barrier as universal would silently clear paths it has not actually made safe.
- Your team wrote its own input validator. What does the analyzer do with it until you tell it about it?Nothing helpful. The validator is just another function call on the path, so taint propagates straight through it and every flow that uses it is reported. The fix is a model change - declare the helper as a sanitizer for the relevant sink family - which closes the whole family at once rather than dismissing findings one by one. The same asymmetry applies in reverse to an in-house request decoder: until it is declared a source, the analyzer reports nothing at all.
- What is the difference between a taint finding and a rule that flags every call to a risky function?The name-matching rule reports uses; the taint rule reports reachability. Matching on a call name flags safe uses with constant arguments and misses nothing about where the argument came from, so it is noisy and shallow. Taint tracking follows values across assignments, concatenations and function boundaries, and reports a trace a reviewer can judge. The cost is that it needs a call graph and a model, and it is much slower to run.
Think of tracing dye in plumbing: you inject it where outside water enters, watch which outlets it reaches, and a filter only counts if the dye actually passes through it - a filter on another pipe proves nothing.
saying these in an interview costs you the question
- Describes it as searching for dangerous function names
- Thinks any call to a sensitive operation is a finding
- Cannot say what makes a value tainted
- Believes any validation in the file clears the path
- Calls a length or null check a sanitizer
- Assumes the tool already knows in-house helpers