skip to content

A shared fleet credential pulled back data no collector should need - why can the access record not name the host?

level: middleimportance: should knowfreq 51%

answer

  1. the record names what was presented
  2. one value, one identity column
  3. the identity field partitions nothing
  4. exclusion is what you actually lose
  5. a stolen credential authenticates successfully

basics

~20 s

An access record can only name what the caller presented. When several hundred collectors present one value, every entry carries that same identity, so the record proves a credential acted and can never say which of its holders acted.

solid answer

~50 s

Attribution is decided by how many holders share one identity, not by how good the logging is. The destination writes down a time, the identity it authenticated, the operation, the object and usually a source address - and with one shared value the identity field holds the same string in every entry, so it partitions nothing. You are left with weaker evidence: a source address, which names a network position often shared by a whole site behind one egress point; a hostname the caller asserts, which any holder of the value can forge; and correlation against each host's own local logs, which are exactly the logs you trust least on a compromised machine. Per-holder credentials fix this at the root, because the identity field becomes the answer. The incident cost is scope: until something else rules a host out, all several hundred are suspects.

code

pseudocode · 24 lines
pseudocode
# two entries from the destination's access record, minutes apart
entry_1:
  at         = '04:12:07'
  actorId    = 'fleet-upload-credential'   # the value presented
  operation  = 'read'
  objectPath = '/uploads/2026-06/'
  sourceAddr = '203.0.113.9'               # the site's shared egress address

entry_2:
  at         = '04:12:09'
  actorId    = 'fleet-upload-credential'   # identical - 400 holders, one value
  operation  = 'write'
  objectPath = '/uploads/2026-06/site-14.dat'
  sourceAddr = '203.0.113.9'

# the question the incident asks of this record
ask: which collector performed the read at 04:12:07 ?

resolve_by(actorId)    -> 400 candidates   # every holder presents this value
resolve_by(sourceAddr) -> the site         # ~20 hosts behind one egress address
resolve_by(at)         -> nothing on its own

# with one credential per collector, actorId alone answers it
resolve_by(actorId)    -> 1 candidate

go deeper

for a junior

Recall that a record can only name the credential that was presented. If four hundred machines present the same one, the record cannot tell them apart.

for a middle

Explain why the identity field has no discriminating power under sharing, and why retention, search and alerting all operate on a distinction the entries never carried.

for a senior

Demonstrate the incident consequence: which questions stay answerable, which fallbacks you would reach for, what each is worth, and why you cannot exclude a single host.

for a principal

Frame attribution as something bought at design time by choosing how many holders share one identity, and decide what resolution the estate needs against the identities it must then operate.

## What an access record can resolve Any destination that accepts a credential can write down what happened: a **timestamp**, the **identity it authenticated**, the **operation**, the **object touched**, and usually a **source address**. That record answers questions of the form *what was done, when, and to what*. It answers *by whom* only to the resolution of the identity it authenticated - and that identity is the credential presented, not the machine that presented it. With one value shared by several hundred collectors, the identity field is the same string in every entry. Nothing about it is wrong; it simply has no discriminating power. Group a million entries by it and you get one group. ## The questions you actually ask in the first hour | question | one shared credential | one credential per collector | |---|---|---| | What was reached? | answered | answered | | When did it start? | answered | answered | | Which host did it? | not answerable | the identity field | | Which hosts can be ruled out? | none can | all the others | | Was it one actor or several? | not answerable | visible as several identities | The two rows in the middle are the expensive ones. **Exclusion matters as much as accusation**: an investigation proceeds by shrinking the suspect set, and a shared credential makes the suspect set un-shrinkable from the destination's side. ## The fallbacks, and what each is worth - **Source address.** It names a network position, which is often a whole site behind one shared egress point, and it can change between requests. It is evidence about *where*, offered as if it were evidence about *who*. - **A hostname the caller puts in the request.** This is a claim made by whoever holds the value. An adversary holding the shared credential can assert any host's name, including a neighbour's, which makes it worse than nothing during an incident even though it is useful in ordinary operations. - **Host-local logs correlated by timestamp.** Only as good as your collection, and on the one machine that matters they are the logs an intruder was best placed to edit or delete. - **The shape of the operation.** A read from a fleet that only ever writes is a very strong signal about *what happened*, and no signal at all about *who did it*. ## Why better logging is the wrong fix Retaining more entries, shipping them to a search system, or alerting on them all operate on records that never carried the distinction in the first place. The identity axis is where attribution is decided, and it is decided at the moment you choose how many holders share one value. Everything downstream inherits that resolution. A companion mistake is expecting authentication-failure alerting to have caught this. **A stolen credential succeeds.** It authenticates correctly every time, because it is the real value; failure alerting sees nothing. The signals that do catch it are about the shape of successful use - a read from a fleet that only writes, activity at an hour the fleet is idle, a volume nobody needs - and none of them tell you which holder. ## What losing attribution costs 1. **Scope.** No holder can be excluded, so containment is fleet-wide by default and forensics multiplies by the fleet size. 2. **Decision quality.** The judgement about whether to withdraw a value every host holds gets made without knowing whether one host or twenty were involved. 3. **Cleanup.** You cannot say which machines to rebuild, so you either rebuild everything or accept that you did not. 4. **The aftermath.** When an external reviewer later asks which host performed the read, the honest answer is *one of four hundred*, and that answer is a finding in itself. ## The fix, and its price Give each holder an identity of its own, and the record's identity field becomes the answer to every question above. The price is the one this whole leaf is about: several hundred identities to enrol, to withdraw when a host is retired, and to keep grants on. Teams that cannot pay it in full usually buy a fraction - an identity per site or per role - which does not give per-host attribution but does shrink the suspect set from the estate to a group, and that is a real improvement rather than a failed one.

  • The collectors send their hostname in every request and it is recorded. Does that give you attribution?
    No. The hostname is asserted by the caller, and any holder of the shared value can assert any name - including the name of a neighbouring collector. It is useful for ordinary operational triage, where nobody is lying, and worthless against an adversary who holds the credential, which is exactly the case where you need it.
  • What does losing attribution cost you in the first hour of an incident?
    Scope. With no per-holder identity you cannot exclude any holder, so every host stays a suspect: containment becomes fleet-wide, forensic effort multiplies by the fleet size, and the decision about withdrawing a value everyone holds is made without knowing whether one host or twenty were involved.
  • Can the access record still tell you anything useful while the fleet shares one credential?
    Yes - it distinguishes values even though it cannot distinguish holders. It will tell you what was reached and when, whether the operation was one the fleet should ever perform, and, during a replacement, whether anybody is still presenting the old value. That last one is how you confirm disuse before withdrawing it.

saying these in an interview costs you the question

  • Says better logging or longer retention would have identified the host
  • Assumes a source address always identifies one machine
  • Expects failed-authentication alerts to catch a stolen valid credential
  • Believes a hostname supplied by the caller is evidence during an incident
  • Treats attribution as an audit nicety rather than an input to containment