skip to content

A training pipeline restores a cached preprocessed feature shard - what integrity risk does that add?

level: seniorimportance: nice to knowfreq 22%

answer

  1. restored data is a model input
  2. no producer identity on the entry
  3. tampering does not crash training
  4. could you recompute this shard?
  5. digest the shard, record it with the run

basics

~20 s

The shard is an unverified input to the model with no record of who produced it, and tampering breaks nothing: training succeeds and only the model's behaviour changes. Caching derived data moves integrity risk into data nobody reviews.

solid answer

~50 s

Restoring a shard instead of recomputing it means the model is trained on bytes whose producer is unknown. Whoever can write that cache - a low-privilege data engineer, or any compromised preprocessing job holding the cache credential - can alter labels or inject crafted rows, and nothing downstream objects: the schema still parses, the job still finishes, aggregate metrics can look normal. There is no diff to review the way there is for code, and no manifest saying what the shard should hash to. Controls that work: restrict cache writes to the pipeline's own identity; record each shard's content digest with the training run, so you can name exactly what the model consumed; keep the evaluation set behind a different credential; and periodically recompute a sampled shard and compare. Restore only what you could reproduce on demand.

go deeper

for a junior

Recognise that cached preprocessed data is an input to the model exactly as a cached library is an input to a binary, and that nothing checks it on restore.

for a middle

Be ready to explain why the tampering is silent - values change, structure does not - and what a content digest recorded with the run would let you answer afterwards.

for a senior

Show that you restrict writes to the pipeline identity, record per-shard digests as run inputs, sample-recompute to detect drift, and keep evaluation data behind a separate credential.

for a principal

Own the cost call: which pipeline stages may cache derived data at all, what evidence a training run must record about its inputs, and who is accountable when a model has to be retrained from clean data.

## Why this is a supply-chain question, not an ML question Pipelines that train models look like ordinary build pipelines with a different payload. Raw data is read, a preprocessing stage produces feature shards - fixed, expensive-to-compute chunks of transformed data - and a training stage consumes them. Because preprocessing is slow, shards are cached: the pipeline computes a key from the raw-data pointer plus the preprocessing code version, and if an entry exists it restores it rather than recomputing. That is the same structure as any build cache, so the same rule applies: **the cached input's trust level equals whoever can write the key**. What differs is the payload and, crucially, the failure signature. ## The failure signature is what makes it dangerous A poisoned code cache usually produces something you could in principle inspect - a binary that differs from a cold build. A poisoned feature shard produces: - a training job that completes normally; - a schema that still validates, because the tampering is in values and labels, not structure; - aggregate metrics that can sit inside their normal band, because a small fraction of altered rows moves them little; - a model whose behaviour differs only on inputs the attacker cares about. Nobody reviews a shard the way a pull request is reviewed. There is no diff, no reviewer, and typically no recorded expected hash. The asset at risk is the model itself - the organisation's model IP and the correctness of every decision it makes - and the compromise persists into every artifact derived from that model until someone retrains from clean data. ## Who can write it Two realistic positions. First, a legitimate insider: a data engineer with write access to the feature store or cache bucket, which is normal and usually broadly granted, since "writing processed data" is their job. Second, a compromised preprocessing job: it holds the credential that writes shards, so anything that achieves code execution inside it - an untrusted transform library, a notebook-derived step, a dependency install hook - inherits the ability to write any shard, including ones for pipelines it has nothing to do with. Neither position requires touching the training code, the model registry or the deployment path. That is the point: the weakest link is the unreviewed intermediate. ## Controls, ordered by what they actually buy 1. **Restrict who may write.** The cache should be writable by the pipeline's own workload identity and by nothing else. Human write access to a path that feeds training is the finding, not the mitigation. 2. **Content-address the shard and record it with the run.** Have the producing job compute a digest of each shard and register it as part of the training run's inputs. Then the question "what data was this model trained on?" has a byte-exact answer instead of a path and a timestamp. This also turns a later tamper into a detectable event rather than an unanswerable one. 3. **Verify on restore.** The consuming job recomputes the digest and compares against the recorded value - and, as always, that recorded value must live where the writer of the shard cannot rewrite it. 4. **Recompute and compare, on a sample.** Periodically rebuild one shard from raw data with the same code version and compare it to the cached entry. This is the only routine check that catches a tamper nobody reported, and it is cheap when sampled. 5. **Keep evaluation data out of the cached path.** If the same write access covers both the training shards and the held-out set used to judge the model, the attacker can poison the training data and adjust the yardstick. Separate credentials, separate storage, separate lifecycle. 6. **Prefer recompute for small or high-stakes stages.** Caching is a cost decision. Where a stage is cheap or the model is high-impact, recomputing removes the entire question. ## The rule that generalises **Restore only what you could reproduce, and only from a writer you would trust to produce it.** If deleting the cache would change the model, the cache is not a cache - it is an unversioned, unreviewed input to a production system, and it must be treated with the controls you would apply to source code: restricted writes, recorded identity, and a way to prove what was consumed. ## What an interviewer is listening for That you spotted the silence of the failure, that you named a realistic writer rather than a movie attacker, and that your controls distinguish *preventing* the write from *detecting* the tamper - because in this domain detection is genuinely hard, which is why the write restriction and the recorded digest do most of the work.

  • Why is a poisoned feature shard harder to detect than a poisoned code cache?
    Code has a reviewable diff, a build that can be repeated cold and compared, and often a recorded expected hash. A shard has none of that: the tampering hides in values, the schema still validates, the job still succeeds, and aggregate metrics barely move when a small fraction of rows changes. Detection has to be engineered in - digests and sampled recomputation - because nothing surfaces it on its own.
  • The team says recomputing every shard is too expensive. What do you propose instead?
    Keep the cache but change who can write it and what you record: writes limited to the pipeline identity, a digest per shard registered with each training run, and a sampled recompute-and-compare on a schedule. That gets most of the assurance for a fraction of the cost, and it lets you name exactly which runs consumed a suspect shard if one is ever found.
  • Where should the evaluation dataset live relative to this cache?
    Behind different credentials and a different storage path. If one write capability covers both the training shards and the held-out set, an attacker can shift the model and adjust the measurement that would have caught it. Separating them means a poisoning attempt has to defeat two boundaries, and the evaluation results stay meaningful as an independent check.

saying these in an interview costs you the question

  • Assuming tampered data would crash the training job
  • Treating cached datasets as outside supply-chain scope
  • Giving humans write access to paths that feed training
  • Recording only a path and timestamp as the run's inputs
  • Keeping evaluation data behind the same write credential

context