How do you validate an S2C2F maturity self-score when the audit practice was never exercised?
answer
- a score is a claim, not a fact
- who has ever exercised this?
- documented versus demonstrated
- measure the latency of the answer
- a control that never blocked anything
basics
~20 sRun the drill. A maturity level claims a capability exists, and the only evidence is exercising it: ask which shipped artifacts contain one specific package version, and time the answer against real records rather than a spreadsheet.
solid answer
~50 sA self-assessed level measures whoever was in the workshop, so I would convert each green into a demonstration. Pick a package and a version at random and ask three questions against a clock: which artifacts we shipped in the last ninety days contain it, where is the copy of the exact file we consumed, and who approved its ingestion. The **latency** of that answer is the score - a two-week reconstruction is not a capability. Then grade the evidence rather than the intent: a statement is worth nothing, a document very little, and only a query you can re-run now or a control that has demonstrably blocked something is worth a green. Inventory backed by a hand-kept spreadsheet fails on freshness and on transitive depth. An Enforce control that has never blocked anything is indistinguishable from one that was never wired up, so trigger a harmless violation and watch.
go deeper
Understand that a maturity score is something an organisation gives itself, and that a practice can be marked done because a document exists rather than because anyone has ever used it.
Be ready to describe what evidence would actually back a practice: a query you can re-run, a log line from a control that fired, a retained artifact - rather than a policy page.
Expect to design the test. Explain the drill, why you measure the time to a complete answer, why a transitive package makes a better probe than a famous one, and how you tell an obeyed control from a broken one.
Own the incentive problem. Decide who is allowed to lower a score, how self-assessment results are used in customer conversations, and how you keep an assessment programme from becoming a document-production exercise.
## Why a self-score drifts The S2C2F is self-assessed by design, and its levels describe capability rather than compliance. That makes it cheap to adopt and easy to inflate. A scoring workshop measures one thing reliably: what the people in the room believe their organisation intends to do. Intent and capability diverge quietly, and nothing in the scoring process notices. The classic shape is a workshop where **Inventory** scores green because a spreadsheet exists, and **Audit** scores green because someone says "we could pull that if we were asked" - and nobody ever has been. The asset at stake in that room is audit truth: the organisation's ability to make a true statement about itself later, to a customer, a regulator, or its own responders. An inflated score does not damage anything today. It damages the day the claim is tested. ## Documents are not capabilities Grade evidence in tiers, and be explicit about which tier each green rests on: 1. **A statement** - someone asserts the practice happens. Worth nothing. 2. **A document** - a policy, a runbook, a diagram. Evidence of intent only. 3. **An artifact** - a spreadsheet, an exported report. Evidence that the practice happened once, at a date printed on it. 4. **A query you can re-run now** - derived from live data, answering in seconds. 5. **A control that has actually stopped something** - with a log entry to show for it. Only the last two justify a green. The distinction that matters is between *documented* and *demonstrated*: an untested capability decays silently, because between the spreadsheet's last edit and today the estate changed and nobody told the spreadsheet. ## Design the drill The fastest way to test a set of scores is one unannounced exercise with a stopwatch: - Pick a package **and a specific version**, ideally a transitive dependency rather than a headline framework, because the direct list is the part people maintain by hand. - Ask: which artifacts shipped in the last ninety days contain it? Where is the copy of the exact file we consumed? Who or what approved its entry, and when? - Record the wall-clock time to a **complete** answer, and record how the answer was produced. The production method is the finding. "A query over build records, thirty seconds" is a green Inventory. "Three engineers grepped repositories for two days" is the same answer with a different score, because it was reconstructed rather than retrieved, and it will not survive the person who did it leaving. Notice also that the second question - where is the file - is the Ingest practice being tested at the same time, which is why this single drill grades several practices at once. ## The special case of Enforce An enforcement control that has never blocked anything produces exactly the same evidence as an enforcement control that was silently misconfigured six months ago: zero blocks. You cannot tell them apart from the outside, and a green score on that basis is a guess. The test is to introduce a deliberate, harmless violation - a throwaway build that tries to consume something policy forbids - and check that it is stopped **and** logged. A control that blocks without logging is only half a control, because the Audit practice has nothing to read. ## Exercising the drill is the practice The neat part is that running the drill is not preparation for the Audit practice; it **is** the Audit practice being performed. Putting it on a schedule and keeping the results is what converts an aspirational green into a defensible one, and the schedule is also what catches the decay you cannot otherwise see. ## Reporting honestly When the drill deflates a score, re-score down and say so. This is uncomfortable because self-assessments are usually produced with a customer conversation somewhere in the background, and the incentive runs one way. But the score's real value is as a planning input - it tells you where to spend next - and an inflated one hides exactly the work that most needs funding. Worse, a level you cannot demonstrate becomes a misstatement the moment anyone asks for evidence behind it, which converts a security gap into a credibility problem. The repair, in order: attach a named evidence owner to every practice; make inventory **derived** from builds rather than maintained by hand, so freshness stops being a human responsibility; wire enforcement to log; and schedule the drill so that at least one practice is exercised every cycle rather than every never.
- What single question exposes an inflated inventory score fastest?Which artifacts we shipped in the last ninety days contain version X of this transitive package - and where is the copy of the file we consumed? A derived inventory answers in seconds. A maintained spreadsheet returns a name and a version, cannot say whether it is still true, and usually does not reach transitive depth at all.
- How do you distinguish an enforcement control that is obeyed from one that is not wired up?You cannot from the outside: both show zero blocks. Trigger a deliberate, harmless violation in a throwaway build and check that it is stopped and that a log entry exists. Until that test has been run, a green enforcement score is an assumption dressed as a measurement.
- Should you report a lower maturity level than the workshop agreed on?Yes. The score's value is as a planning input, and an inflated one hides the work that needs funding. A level you cannot demonstrate also becomes a misstatement the moment someone asks for the evidence behind it, which turns a fixable gap into a credibility problem.
saying these in an interview costs you the question
- Treats an existing document as evidence of a capability
- Scores a control green because it has never fired
- Assumes a hand-maintained inventory is current
- Never asks when the practice was last exercised
- Inflates the score ahead of a customer conversation