An archival tier's compressed copy of a corpus came out larger than the originals — what does that tell you?
answer
- divide growth by block count
- framing tax versus coded-block defect
- already-compressed input is the usual cause
- compressing ciphertext can never win
- bound is on the map, not the coder
basics
~20 sA corpus that grows says its blocks are already compressed or encrypted: nearly every block falls back to a raw copy and pays a header, so the total rises by about one header per block. No different coder fixes that.
solid answer
~50 sStart by turning the delta into a per-block number, because the cost shape identifies the cause. Growth of a few bytes per block is the stored-raw fallback doing its job on input that has no skew and no repeats left — already-compressed or encrypted payloads. Growth of hundreds of bytes per block means coded blocks are being emitted where a raw copy would have been smaller, which is an emit-decision defect. Growth proportional to nothing in particular usually means the block size is small enough that framing dominates. Then check where compression sits relative to encryption in the pipeline, since compressing ciphertext can never win. The counting bound is why no substitution helps: the limit is a property of the map, not of the implementation, so the real decisions are block size, pipeline order, and whether to attempt coding on this class at all.
go deeper
Recall that an archive can legitimately end up slightly larger, and that the usual reason is input which was already compressed rather than a broken tool.
Explain the per-block shape of the cost and why no alternative coder removes it, since the limit follows from reversibility rather than from any implementation.
Diagnose from the size delta: growth per block, fallback rate, pipeline order relative to encryption, and then a decision about which classes are worth attempting.
Balance the processor cost of coding a corpus that cannot repay it against a saving counting rules out, and set a per-class policy with a stated growth bound.
## Read the delta before proposing a fix The single most useful number is not the total growth but the growth **per block**, because each cause has its own signature and they do not overlap much. | Growth per block | Most likely cause | What it means | |---|---|---| | A few bytes | Stored-raw fallback on every block | The input has nothing left to code; the format is behaving correctly | | Tens to hundreds of bytes | Coded blocks emitted where a raw copy was smaller | The emit decision is not comparing sizes, or a model description is shipped per block | | A few bytes, but a huge total | Block size far too small | Framing dominates because there are far too many blocks | Getting this number costs one division and settles most of the argument before anybody opens a tuning guide. ## The three causes, ranked 1. **The corpus is already compressed or encrypted.** A well-coded payload has had its symbol skew and its repeated sequences consumed; a cipher's output is required to look uniform to anyone without the key. Neither leaves the coder anything to trade on, so every block falls back and pays its header. This is the common case in an archival tier, whose inputs are usually other systems' outputs. 2. **Compression sits after encryption in the pipeline.** Same effect, different reason, and the fix is structural rather than about settings: move the compression step ahead of the encryption step wherever the design allows it, so the coder sees structured plaintext. 3. **The writer has no working fallback.** Instead of storing a block raw, it emits a coded block that is longer than what it read, sometimes with a model or code-table description attached. This is the only one of the three that is a genuine defect, and the per-block arithmetic is what distinguishes it. ## What the counting bound rules out as a fix The bound is about the map, not the code: a lossless map is injective, there are fewer short strings than inputs, and any scheme that shrinks something must grow something else. So an entire family of proposals is dead on arrival. | Proposal | Verdict | |---|---| | Substitute a stronger coder | No. The obstruction is the absence of structure in the input, which no coder creates | | Raise the effort level | No. More search over a space with nothing in it costs processor time and returns the same bytes | | Run a second compression pass | No. Pass two sees the near-uniform output of pass one and adds another layer of framing | | Enlarge the blocks | Helps. Fewer headers, so the tax falls toward zero, but never below it | | Move compression ahead of encryption | Helps, where the pipeline permits it, because the coder then sees plaintext | | Skip compression for this class of input | The honest answer. Exact parity, and the processor time comes back | ## What to measure before deciding - The growth per block, as above, and the block count implied by the configured block size. - What fraction of blocks fell back to stored form. Near 100% is the signature of a corpus that should not be coded at all. - The processor time spent coding blocks that then fell back, which is the real cost: the bytes are a rounding error, the compute may not be. - Whether the corpus is homogeneous. A mixed corpus usually wants a per-class policy, since the structured part still pays for itself. ## What to tell the team Say the mechanism and the decision, not the tool. 'These payloads are already compressed, so there is no skew and no repetition left for a coder to trade on. Counting guarantees some inputs must fail to shrink, and this is the class that pays; the payment is one block header each. Swapping the coder cannot change that because the limit is a property of any lossless map. Our options are larger blocks, compressing earlier in the pipeline, or not compressing this class — and the last one also gives us the processor time back.' The reason this belongs in a senior conversation rather than a lecture is that the wrong reflex is expensive: a team that reads growth as a defect will spend a sprint on coder selection, effort levels and second passes, and end with the same bytes plus more compute. The counting argument is the cheapest way to redirect that effort on day one, and the per-block arithmetic is what turns it from an assertion into a measurement anyone on the team can repeat.
- How would you tell the framing tax apart from a writer with no working raw fallback?Divide the total growth by the number of blocks. A few bytes per block is the header on stored blocks, which is the format working correctly. Tens or hundreds of bytes per block means coded blocks are being emitted where a raw copy would have been smaller, which is a defect in the emit decision.
- The team asks for a guarantee that the archive never grows at all. What do you offer instead?Offer the bound, not the guarantee. Any lossless format that ever shrinks an input must expand another, so 'never grows' is only achievable by never coding anything. What is real: growth capped at one header per block, tunable by block size, or exact parity for the classes you choose not to compress.
- Does the same reasoning apply if the corpus is a mix of compressed and plain payloads?Yes, per class. The structured part still repays coding, so a blanket 'stop compressing' is as wrong as a blanket 'compress everything'. Measure the fallback rate per class and set the policy where the rate actually is, rather than on the aggregate number.
saying these in an interview costs you the question
- Concludes the compressor is defective and files a bug
- Proposes a stronger coder or a higher effort level as the fix
- Assumes growth means blocks were corrupted or duplicated
- Expects a second compression pass to recover the loss
- Treats the growth as fixed overhead rather than one header per block