skip to content

A Kubernetes pod failed overnight, and by morning kubectl get events shows nothing about it. How long do Kubernetes Events survive, why, and how do you keep that evidence?

level: seniorimportance: should knowfreq 44%

answer

  1. just another API object
  2. etcd lease, not a controller
  3. kube-apiserver --event-ttl, one hour
  4. updates renew the lease
  5. export, or it is gone

basics

~20 s

Events are ordinary API objects stored in etcd with a lease set by kube-apiserver's --event-ttl, default one hour, renewed whenever the Event is updated. After that they vanish, so keeping them means exporting them continuously to a durable store.

solid answer

~40 s

An Event is a normal namespaced API object, but kube-apiserver writes it to etcd with a lease whose length is `--event-ttl`, one hour by default, and each update — such as a repeat bumping the `count` — renews it. An Event that stopped recurring at 01:12 is gone by 02:12, and `kubectl describe` shows nothing either, because it reads the same objects. The TTL exists because Events are the highest-churn object in a cluster. To keep the evidence, run a small controller that watches Events and writes them into the log pipeline, capture `describe` output during incidents, and lean on status that lives as long as the pod — the `PodScheduled` condition, a container's `lastState`. Raising `--event-ttl` is possible but grows etcd; a separate etcd for events via `--etcd-servers-overrides` limits that.

code

bash · 5 lines
bash
kubectl events -n media --for pod/transcoder-7f9c4d6b8-x2lqp --types=Warning
kubectl get events -n media \
  --field-selector involvedObject.name=transcoder-7f9c4d6b8-x2lqp,type=Warning
kubectl get pod transcoder-7f9c4d6b8-x2lqp -n media \
  -o jsonpath='{.status.conditions[?(@.type=="PodScheduled")].message}'

go deeper

for a junior

Recall that Kubernetes Events disappear about an hour after they last changed, so an empty Events list does not prove nothing happened.

for a middle

Explain that the expiry is an etcd lease set by kube-apiserver's --event-ttl, renewed on update, and how count and series record repeats.

for a senior

Show the incident habit of capturing Events early, reading status that outlives them, and exporting Events so overnight failures leave evidence.

for a principal

Weigh longer Event retention against etcd size and write load, and decide whether exporting Events or a separate events etcd fits the platform.

## What an Event is A Kubernetes **Event** is a namespaced API object that a component writes to report something about another object: the scheduler could not place a pod, the kubelet pulled an image, a probe failed. It is served in two API groups that share the same storage: | Meaning | `v1` (core) field | `events.k8s.io/v1` field | |---|---|---| | object the Event is about | `involvedObject` | `regarding` | | short machine-readable cause | `reason` | `reason` | | human-readable text | `message` | `note` | | `Normal` or `Warning` | `type` | `type` | | who reported it | `source` | `reportingController`, `reportingInstance` | | repetition | `count`, `firstTimestamp`, `lastTimestamp` | `series` (count and last observed time) | | when it happened | `firstTimestamp`, `lastTimestamp` | `eventTime` | Components that use the older client-go recorder, such as the kubelet, repeat a similar Event by updating `count` and `lastTimestamp`. Components on the newer events API, such as kube-scheduler, set `eventTime` and keep a `series` instead. The older recorder also collapses bursts: when many similar events arrive within a short window it emits one message prefixed `(combined from similar events)`. ## Why Events expire kube-apiserver stores Events in etcd **with a lease**. The lease length is the API server flag `--event-ttl`, default **1h**, and it is applied on create and again on every update. So: - an Event that fired once lives about one hour; - an Event that keeps repeating — a `BackOff` updated every few minutes — keeps getting renewed and lives until about an hour after its **last** update; - once the lease expires, etcd deletes the key and the Event disappears from every read path: `kubectl get events`, `kubectl events` and the Events section of `kubectl describe`. The TTL exists because Events are the **highest-churn object** in most clusters: every probe failure, pull and scheduling attempt writes or updates one. Without expiry they would dominate etcd's size and write load. ## A worked example On a single-node development cluster on a laptop, a video-transcoding worker is scaled up at 01:12. The node's CPU requests are already at **91%** of allocatable, so the new pod stays `Pending` and kube-scheduler writes a `FailedScheduling` Warning with `Insufficient cpu`. The job is cancelled at 01:30 and the pod deleted. At 09:00 the developer finds nothing: the pod object is gone, and the Event lease ran out about an hour after its last update. Had the pod still existed, the evidence would partly survive in its **status**: the `PodScheduled` condition with `status: "False"`, `reason: Unschedulable` and the scheduler's message. Status lives as long as the object, while Events do not. ## Reading Events before they go - `kubectl events -n media --for pod/<name> --types=Warning` filters to one object and sorts by the best available time — series last-observed time, then `lastTimestamp`, then `eventTime`. - `kubectl get events -n media --field-selector involvedObject.name=<name>,type=Warning` does the same filtering on the server. - `kubectl get events --sort-by=.lastTimestamp` is the common idiom, but Events written through `events.k8s.io/v1` can leave `lastTimestamp` empty, so scheduler events can sort to the wrong place. `kubectl events` avoids that. - `kubectl get events --watch` during a reproduction catches Events as they happen. ## Keeping the evidence Ordered from cheapest to most thorough: 1. **Capture during the incident.** Save `kubectl describe` and `kubectl events` output into the incident notes before anything is deleted. 2. **Rely on status that outlives Events.** Pod conditions, a container's `lastState.terminated` reason and exit code, and a controller's `status.conditions` persist with their objects. 3. **Export Events continuously.** Run a small controller that watches Events cluster-wide and writes each one as a structured log line into the cluster's log pipeline. That is the only option that covers a pod deleted overnight. 4. **Raise `--event-ttl` carefully.** A longer TTL multiplies how many Events etcd holds at once. If you do raise it, `--etcd-servers-overrides` can put the `events` resource on a separate etcd cluster so they cannot crowd out the rest. ## Common misreadings - **"No Events means nothing happened."** It may mean it happened more than an hour ago. - **"`count: 1` means it happened once."** For newer-API Events the repetition lives in `series`, and bursts can be collapsed. - **"Events are an audit trail."** They are best-effort: recorders rate-limit and drop events under load. Use audit logging for an authoritative record of API requests.

  • Why not just set `--event-ttl` to a week?
    Events are the highest-churn object in the cluster, and the TTL decides how many exist at once. A week instead of an hour keeps roughly 168 times as many objects in etcd, which grows the database, slows full lists of Events and raises write load on the store every other object depends on. If longer retention is required, export Events instead, or at least move them to their own etcd with `--etcd-servers-overrides`.
  • Why can `kubectl get events --sort-by=.lastTimestamp` put a scheduler's Events in the wrong order?
    kube-scheduler records through the `events.k8s.io/v1` API, which sets `eventTime` and a `series` rather than the legacy `lastTimestamp`. Read through the core API those Events can have an empty `lastTimestamp`, so they sort as if they were the oldest. `kubectl events` sorts by the series' last observed time, then `lastTimestamp`, then `eventTime`, which gives a sensible order.
  • Can Events serve as an audit record of what happened to a pod?
    No. Event recording is best-effort: recorders rate-limit, collapse similar events and give up after repeated write failures, and the API server expires what was written. For who changed what, use the API server's audit log. For what a workload did, use its logs. Events are hints for live triage.

saying these in an interview costs you the question

  • Events are kept as long as the object they describe exists.
  • A garbage-collection controller in kube-controller-manager deletes old Events.
  • The kubelet's eventRecordQPS setting controls how long Events are retained.
  • If kubectl describe shows no Events, nothing went wrong with the pod.
  • Events are a complete, reliable audit trail of cluster activity.
  • Raising --event-ttl has no cost because Events are small.