You inherit a large, long-lived codebase with thousands of queries assembled by string concatenation. How would you drive injection risk toward zero over time, and what limits the damage in the meantime?
answer
- Unrepresentable > absent: type-level API
- Frozen baseline ratchet, one-way door
- Rank by reachability × privilege × data sensitivity
- Blast radius: per-service accounts, read-only replica, no multi-statement
- WAF buys time, never counts as fixed
basics
~20 sMake the unsafe path unrepresentable rather than hunting bugs: one query-construction API that accepts a static template plus bound values, a build-time ratchet so no new concatenation is added, risk-ranked migration of the existing sinks, and least-privilege plus data-level controls to cap blast radius while the work runs.
solid answer
~50 sTreat it as an invariant-enforcement programme, not a bug backlog. First, **stop the bleeding**: introduce a single data-access API whose signature only accepts a compile-time-constant template plus bound values, then add a static-analysis or taint rule with a frozen baseline so existing violations are tolerated and new ones fail the build. Second, **burn down the baseline in risk order** — unauthenticated internet-facing paths first, then anything running on a high-privilege connection, then internal batch jobs — measuring percentage of sinks converted rather than tickets closed. Third, **cap the blast radius now**, because migration takes quarters: per-service database accounts with no DDL and no cross-schema reads, read-only replicas for reporting, multi-statement execution disabled in the driver, sensitive columns encrypted or tokenised so exfiltration yields little, and alerting on error-rate and query-shape anomalies. A WAF buys detection time; it is never the control you count on.
code
text · 8 linesrule: untrusted value reaches statement text -> ERROR
baseline.txt: 2,431 known sites (grandfathered)
CI: violations_now \ baseline != empty -> fail the build
baseline entry removed -> permanent, cannot be re-added
reported metric: sinks_converted / sinks_total (monotonic, converges)
NOT reported: tickets closed this sprint (unbounded, gameable)go deeper
Focus on the mechanical half: parameterise the queries, add a lint rule so no new concatenation is added, and start with the internet-facing code.
Add the frozen-baseline ratchet and a sensible risk ordering, and mention least-privilege database accounts as interim containment.
Cover blast-radius controls in depth — per-service accounts, read-only replicas, disabled multi-statement, data-level encryption — plus detection signals and regression testing of converted paths.
Lead with the invariant-enforcement framing and a converging metric, discuss the organisational mechanics (baseline as a one-way door, who owns the number, when the programme ends), and be explicit about which controls are compensating versus permanent.
## Frame the problem correctly Thousands of concatenated sinks is not a security finding, it is a systemic property of the codebase. Any plan that consists of "scan, ticket, fix" loses, because new sinks appear as fast as old ones are closed and the backlog never converges. The goal is to change the codebase so that the defect becomes *impossible to express*, and then to drain the residue. ## 1. Make the safe path the only path Define one internal data-access entry point. Its signature should make the unsafe call unrepresentable: the statement argument accepts only a constant template type (many languages have a compile-time-constant string or a template-literal type), and values arrive as a separate typed collection. Where identifiers must vary, the API accepts an enum, not a string. If the language cannot express that, the next-best is a single small module where all statement assembly lives, so the reviewed surface is one file rather than the whole repository. This matters more than any scanner: a control enforced by a type signature holds for every future author without their cooperation, while a guideline in a wiki decays. ## 2. Ratchet, do not sweep Add a static-analysis rule that flags untrusted data flowing into statement text (a taint rule if the toolchain supports it; a structural rule against concatenation at the sink otherwise), and record the current violations as a frozen baseline. The build fails on any violation *not* in the baseline. Every new line of code is therefore safe from day one, while the existing thousands do not block delivery. Removing an entry from the baseline is a one-way door — the file can never regress. This converts an unbounded remediation project into a monotonically shrinking number, which is the only shape that finishes. Pair it with a pre-merge check that fails if the baseline grows, and publish the count as a visible metric. ## 3. Sequence the burn-down by real risk Not all sinks are equal. Rank by three factors: reachability (is the input attacker-controlled and reachable without authentication?), privilege (what can the connection this code uses actually do?), and data sensitivity (what tables are in reach). The order that follows is usually: unauthenticated public endpoints, then authenticated but tenant-crossing endpoints, then admin tooling that runs on a powerful account, then batch jobs whose inputs are internal but which often read attacker-planted rows — the second-order case that scanners miss precisely because the injecting request and the triggering request differ. Measure conversion percentage of sinks, not tickets closed. Tickets reward slicing; percentage of the invariant achieved rewards finishing. ## 4. Cap blast radius while the work runs Migration takes quarters, so assume some sinks remain exploitable and reduce what an exploit yields. - **Least privilege per service account.** Separate credentials per service, granted only the tables and operations it uses, no DDL, no access to other schemas. An injection in the reviews service should not be able to read the payments tables. - **Read-only replicas for reporting.** Report paths are the most dynamic and the most concatenated; pointing them at a replica with a read-only account removes tampering entirely from the highest-risk code. - **Disable multi-statement execution** in the driver where the option exists, removing the stacked-statement escalation. - **Data-level controls.** Encrypt or tokenise the crown-jewel columns so a mass `SELECT` returns ciphertext; keep the keys outside the database's own trust boundary. Rate-limit and cap result-set sizes so bulk exfiltration is slow and noisy. - **Detection.** Alert on database error-rate spikes (blind injection is noisy in errors), on unusual query shapes or unusually large result sets per endpoint, and on unexpected tables appearing in a service's access pattern. A WAF or database firewall belongs here: it buys time to respond and raises the cost of automated scanning, but it is a signature-matching heuristic in front of a parser you do not control, so never treat it as remediation or as a reason to defer a sink. ## 5. Verify, do not assume Add regression tests at the boundary for the shapes that matter — a payload in a filter, in a sort key, in a stored value read back later — and run a query-level fuzz pass in CI against the endpoints you have converted, so "converted" means demonstrated rather than declared. Include the database side: procedures containing dynamic SQL live outside the repository and must be inventoried explicitly, or they will be the last thing anyone looks at. ## 6. Know when to stop The programme is done when the baseline is empty, the type-level API is the only way to reach the driver, and the residual controls (least privilege, replicas, encryption) stay in place as defence in depth rather than as compensating controls. Retain the ratchet permanently; the cost is a few seconds of build time and it is what prevents the next decade of drift.
- Leadership offers to buy a web application firewall instead of funding the migration. How do you respond?Accept it as a detection and time-buying layer, refuse it as a substitute. A WAF inspects requests with heuristics in front of a parser whose exact dialect and session settings it does not share, so bypasses are a research commodity, and it is blind to second-order payloads that arrive as innocuous stored data and are triggered later by a batch job. Fund it as monitoring and rate-limiting value, keep the burn-down metric on the roadmap, and be explicit that the WAF does not change the risk rating of any sink.
- Which single control would you put in place first if you had one week?The build-time ratchet with a frozen baseline, because it stops the problem growing and turns an unbounded backlog into a converging number — everything else is a rate question after that. If the toolchain cannot support it in a week, the fallback is per-service least-privilege database accounts, since it caps the blast radius of every existing sink at once without touching application code.
saying these in an interview costs you the question
- Proposing a scan-and-ticket backlog with no mechanism preventing new sinks.
- Counting a WAF or database firewall as remediation rather than detection.
- Assuming least privilege protects you when stored procedures run with definer's rights.
- Measuring tickets closed instead of the fraction of sinks converted.
- Forgetting database-resident code — procedures, triggers, views with dynamic SQL — because it is not in the repository.
- Treating internal batch jobs as low risk when they read attacker-planted rows (second-order injection).