skip to content

What is Trogdor and when would you use it instead of the kafka-*-perf-test.sh scripts?

level: seniorimportance: nice to knowfreq 18%

answer

  1. Coordinator + Agents, REST API, JSON task specs
  2. workloads: ProduceBench, ConsumeBench, RoundTripWorkload
  3. faults: NetworkPartition, ProcessStop, latency
  4. distributed multi-node vs single-box perf scripts
  5. load + chaos together; soak/resilience tests

basics

~20 s

Trogdor is Kafka's distributed test framework: a coordinator plus agents that run coordinated workloads and fault injections across many nodes. Use it for large-scale, multi-node, repeatable load and chaos tests; use the perf-test scripts for quick single-box benchmarks.

solid answer

~50 s

Trogdor is the test harness that ships in Apache Kafka's source tree for running scalable, distributed workloads and fault injection. It has a central Coordinator that schedules tasks and multiple Agents (one per host) that execute them. You submit JSON task specs over its REST API describing 'tasks' that are either workloads (e.g. ProduceBench, ConsumeBench, RoundTripWorkload) or faults (NetworkPartitionFault, ProcessStopFault, latency injection). Unlike the simple perf-test CLIs — which spin up a single producer/consumer on one box for a quick number — Trogdor coordinates many producer/consumer agents simultaneously to generate cluster-scale load, runs reproducible scripted scenarios, and injects failures to test resilience under load. You reach for Trogdor when you need multi-node aggregate throughput, soak/endurance tests, or combined load-plus-chaos experiments; you reach for kafka-producer-perf-test.sh / kafka-consumer-perf-test.sh when you just need a fast, ad-hoc throughput/latency reading from one machine.

go deeper

for a junior

Know Trogdor exists as Kafka's distributed test framework, separate from the simple perf CLIs.

for a middle

Describe coordinator/agent architecture and that it runs workloads and faults.

for a senior

Choose between Trogdor and perf scripts by scale/chaos/reproducibility needs and name workload/fault task types.

for a principal

Design cluster-scale soak and resilience test suites combining Trogdor workloads with fault injection in CI.

## What Trogdor is Trogdor is Apache Kafka's **distributed test framework**, living under the project's tooling. Where `kafka-producer-perf-test.sh` and `kafka-consumer-perf-test.sh` are single-process CLIs for a quick benchmark, Trogdor is built to orchestrate complex workloads and faults across an entire cluster of test machines, repeatably. ## Architecture - **Coordinator**: a single central server. You talk to it over a REST API, submitting and querying tasks. It schedules tasks with start times and durations and aggregates status. - **Agents**: a process running on each worker host. Agents receive task assignments from the coordinator and actually execute the workload or fault on that node. This coordinator/agent split is what lets Trogdor drive many machines in lockstep — something the single-JVM perf scripts cannot do. ## Task types Tasks are submitted as JSON specs and fall into two families: 1. **Workloads** — generate load: - `ProduceBench`: distributed producers pushing to topics with configurable rate, size, partitions. - `ConsumeBench`: distributed consumers reading topics/groups. - `RoundTripWorkload`: produces then verifies the same records are consumed, checking correctness and end-to-end latency. 2. **Faults** — inject failures: - `NetworkPartitionFault`: split nodes apart. - `ProcessStopFault`: pause/stop a broker process. - latency/IO faults. Combining workloads and faults in one scenario lets you measure how throughput and latency behave *while the cluster is degraded* (e.g. during a broker failure or network partition) — true resilience testing. ## When to choose which | Need | Tool | |------|------| | Quick single-box throughput/latency number | `kafka-producer-perf-test.sh` / `kafka-consumer-perf-test.sh` | | Cluster-scale aggregate load from many hosts | Trogdor (ProduceBench/ConsumeBench) | | Reproducible scripted scenarios in CI/soak tests | Trogdor | | Load + fault injection (chaos) together | Trogdor | | End-to-end produce→consume correctness + latency | Trogdor RoundTripWorkload | ## Trade-offs Trogdor is far more powerful but heavier: you must deploy a coordinator and agents, author JSON specs, and operate the REST API. For a developer wanting a five-minute capacity sanity check on one node, the perf-test scripts are the right, lighter tool. For Kafka-developer-grade system tests, scale benchmarks, and resilience experiments, Trogdor is the framework.

  • What are Trogdor's two main components and how do they relate?
    A single Coordinator that schedules and aggregates tasks via a REST API, and per-host Agents that execute the assigned workloads or faults on each node. The coordinator/agent split enables synchronized multi-node testing.
  • What does the RoundTripWorkload add over a plain ProduceBench?
    It both produces records and verifies they are consumed back, so it checks end-to-end correctness and measures end-to-end latency, not just one-way produce throughput.

saying these in an interview costs you the question

  • Calling Trogdor just another name for kafka-producer-perf-test.sh — it's a distributed coordinator/agent framework.
  • Saying Trogdor only does load and cannot inject faults.
  • Claiming Trogdor runs on a single JVM like the perf scripts.
  • Reaching for Trogdor for a trivial one-box sanity check.

context