skip to content

In an automated coding assessment, why can code that passes the sample cases still score low?

level: middleimportance: must knowfreq 71%

answer

  1. Examples show format, not coverage
  2. The grading suite is never shown
  3. Constraints predict the hidden tests
  4. Empty, single, duplicate, maximum
  5. Score tracks the fraction of tests passed

basics

~20 s

Sample cases are illustrations, not the grading suite. The hidden tests add boundary inputs and inputs at the constraint ceiling, and where scoring is per test, the problem score is the fraction of that hidden suite your submission passes.

solid answer

~40 s

The examples printed in the problem statement exist to show input and output format; the grading is done by a hidden suite you never see. That suite typically includes the boundaries the examples skip — empty input, a single element, all-equal values, the extremes of the stated range — plus timing tests built at the top of the constraint ceiling. Because many graders award credit per test case, a submission that handles the shown examples and nothing else can land a small fraction of the problem. So before I submit I stop reading the examples and start reading the constraints, hand-build the boundary cases they imply, and run those. That habit converts a partial score into a full one more reliably than another rewrite does.

go deeper

for a junior

Remember that the examples in the statement are not the grading tests, and that you are expected to invent your own boundary inputs before submitting. Empty and single-element input are the two that catch the most first-time candidates.

for a middle

Explain the mechanics: what categories a hidden suite usually contains, how per-test credit turns a problem score into a fraction, and why a correct slow solution still collects most of the available points.

for a senior

Demonstrate the judgment of buying the right tests with the minutes you have left — recognising when the remaining failures are cheap boundary fixes and when they demand a different approach you no longer have time for.

for a principal

Own what per-test partial credit does to a screen: it rewards steady coverage over ambition and can rank a careful partial solver above a candidate who nearly finished something harder. That is a defensible bias, but it is a choice.

## Two different sets of tests Every automated assessment problem carries two test sets. The **sample cases** are printed in the statement, usually one to three of them, and they exist to disambiguate the wording: what the input looks like, what the output looks like, whether there is a trailing line, whether the answer is a count or a list. The **hidden suite** is what the grader actually runs, and it is written by whoever authored the problem with the explicit intent of separating candidates. The two sets do not overlap in purpose. A submission that satisfies the samples has proven that it reads the format correctly. It has proven almost nothing about correctness. ## What the hidden suite usually contains Across platforms the pattern is consistent enough to plan around. On the invented assessment used through this leaf — a mid-size product consultancy screening junior frontend candidates with three problems in an 82-minute window — the third problem carried 17 hidden tests behind 2 shown examples. A plausible breakdown, and a good default guess for any problem: - **Format and happy path** (the samples again, plus a few variations). - **Boundary inputs**: empty, one element, two elements, all elements identical, already-ordered input, input where the answer is the first or last item. - **Value extremes**: the smallest and largest values the stated range allows, including negatives and zero when the range permits them. - **Scale tests**: inputs built at the constraint ceiling, run under a time limit. These are the tests that fail with a timeout rather than a wrong answer. - **Output pedantry**: rounding, ordering of a returned collection, whitespace, the exact wording of a sentinel value. ## Partial credit changes the strategy Where scoring is per hidden test, the problem score is proportional to the fraction passed. On a 180-point problem with 17 hidden tests, passing 16 of them earns roughly 169 points; passing 5 earns roughly 53. Three practical consequences follow. First, **an unfinished problem is worth attempting**. A solution that handles only the simple case often passes the format and happy-path tests and several boundary tests, which is a meaningfully non-zero score against a blank editor's zero. Second, **a slow but correct solution is not a failure**. It typically clears everything except the scale tests, which is most of the suite. That is why submitting a correct brute force before optimising is the safe order of operations. Third, **the last few tests are the expensive ones**. Moving from 5 of 17 to 14 of 17 is usually a matter of boundary handling and costs minutes; moving from 16 to 17 can require a different approach entirely. Know which one you are buying with the time you have left. ## Reading the constraints as a test plan The constraints paragraph is the most under-read text on the page, and it is effectively a summary of the hidden suite. If a length may be zero, there is a test with zero. If values may be negative, there is a test with negatives. If the ceiling is large enough that a naive scan of every pair would be slow, there is a timing test built exactly there. Treat each clause as a test case and run it yourself before submitting. ## Reading the score report afterwards When the run is over, the platform's score report shows the per-problem hidden-test summary: how many passed, how many failed, and often the category or name of each failing test. That summary is the highest-value artefact of the whole exercise for your own practice, because it tells you whether you lose points to correctness, to boundaries, or to scale — three different problems with three different fixes. It also exposes the compounding version of the classic failure: attempting the problems in the order given and timing out on the hardest leaves no minutes for self-testing anywhere, so the earlier problems come back with boundary failures too.

  • How would you guess what the hidden tests cover when the statement shows only two examples?
    Read the constraints as a test plan. Every clause implies a case: a lower bound of zero implies an empty input, a range crossing zero implies negatives, a large ceiling implies a timing test built at that ceiling. Add the output pedantry the statement hints at — ordering, rounding, the sentinel for no answer — and you have most of a hidden suite without seeing it.
  • One hidden test fails and you cannot see its input — what do you do with the remaining minutes?
    Do not rewrite wholesale, because the passing submission is already banked. Re-read the constraints for the case you have not handled, hand-run the boundaries in order of cheapness — empty, single element, duplicates, the extreme of the stated range — and check the output format for ordering and rounding. If nothing surfaces in a few minutes, leave the banked score and spend the time elsewhere.
  • Does partial credit mean a problem you cannot fully solve is still worth starting?
    Yes, and this is the strongest practical argument against leaving a problem blank. A correct handling of the simple case usually clears the format and happy-path tests and several boundary tests. Against a blank editor scoring nothing, that is a real gain for a few minutes of work.

The worked examples printed in a problem statement are like the two practice words on a spelling-test handout: they show you how the exam is laid out, not the words you will be graded on.

saying these in an interview costs you the question

  • Submitting the moment the shown examples pass
  • Skipping the constraints section when choosing an approach
  • Assuming one failing hidden test zeroes the whole problem
  • Never running empty or single-element input before submitting
  • Reading a timeout as a platform fault rather than a scale test
  • Leaving a problem blank because a partial score seems worthless

context