Why can adding a combiner to a Hadoop MapReduce job make an average come out wrong?
answer
- it may run, or it may not
- free bandwidth, but only sometimes
- think about the algebra of the function
- mean of means loses the group sizes
- carry the count alongside the sum
basics
~20 sA combiner is a reduce-style function the framework may run zero, one or many times over a mapper's own output. An arithmetic mean is not decomposable that way, so averaging partial groups and then averaging those averages gives a wrong result.
solid answer
~50 sA combiner is an optional **map-side mini-reduce**: it runs in the map task on that mapper's output to shrink what has to cross the network. It is set with `job.setCombinerClass(...)` and must be a `Reducer` whose input *and* output types both equal the map output types, because its output is fed back into the same sorted stream. The framework decides when to invoke it — on each spill, again during the merge of spills, or not at all — so the job must produce identical results whether it runs zero times or five. That only holds for functions that are commutative and associative: sum, count, min, max, and bitwise/set unions are fine; mean, median, and "top-N by a percentile" are not, because the mean of means weights small groups equally with large ones. The standard fix is to emit `(sum, count)` pairs, let the combiner add pairs, and divide once in the reducer.
code
java · 12 lines// WRONG: reusing the averaging reducer as a combiner
public class AvgReducer extends Reducer<Text, DoubleWritable, Text, DoubleWritable> {
public void reduce(Text key, Iterable<DoubleWritable> values, Context ctx)
throws IOException, InterruptedException {
double sum = 0;
long n = 0;
for (DoubleWritable v : values) { sum += v.get(); n++; }
ctx.write(key, new DoubleWritable(sum / n));
}
}
job.setCombinerClass(AvgReducer.class); // means of means: not associativego deeper
Recall that a combiner is a map-side pre-aggregation whose job is to shrink data before the network hop, and that word count is the classic case where it helps a lot.
Explain the contract: it runs on spills and merges, zero or more times, with input and output types equal to the map output types. Be able to show concretely why an average breaks and how sum-and-count pairs fix it.
Demonstrate that you check whether the combiner is earning its keep — duplication per mapper, the combine counters, map output bytes — and that you can spot the class of silent correctness bugs that only appear once data volume makes the framework invoke it.
Frame it as a general rule for distributed aggregation: only algebraic, mergeable summaries are safe to compute in stages, and every pipeline you own should carry the smallest mergeable state and finalize exactly once. That principle outlives MapReduce.
## What a combiner actually is In Hadoop MapReduce, a combiner is an optional function that runs **inside the map task**, over that single mapper's intermediate output, before anything crosses the network. It is configured with `job.setCombinerClass(SomeReducer.class)` and is written as a `Reducer` subclass — there is no separate `Combiner` base class. Because a combiner's output is spliced back into the same sorted, partitioned stream that the map task is producing, its input and output types must both be the map output types `(K2, V2)`. If your mapper emits `(Text, IntWritable)` and your reducer emits `(Text, DoubleWritable)`, that reducer cannot be reused as the combiner — the types do not close. ## Where and how often it runs Map output accumulates in an in-memory buffer (`mapreduce.task.io.sort.mb`, 100 MB by default) and is spilled to local disk when it passes `mapreduce.map.sort.spill.percent` (0.80). The combiner may run as each spill is written, and again when several spill files are merged into the map task's final output. It may also **not run at all** — for example when a map task produces so little output that there is only one small spill. This is the crux: the framework treats the combiner as a pure optimization it is free to skip or repeat, so the job's result must be invariant to the number of invocations. ## Why the average breaks The operations that survive arbitrary re-application are the commutative and associative ones. `sum(sum(a, b), c) == sum(a, sum(b, c))`, so summing partial sums is exact. An arithmetic mean has no such property: `mean(mean(2, 4), 9) = mean(3, 9) = 6`, but the true mean of `2, 4, 9` is 5. The combiner has thrown away the group sizes, so the second averaging pass weights a two-element group exactly like a one-element group. What makes this insidious is that the job still succeeds, the numbers still look plausible, and the error changes shape when the data volume changes — you can pass a small test and be quietly wrong at scale, because with small inputs the combiner may never fire. ## The sum-and-count fix Make the intermediate value carry enough information to be merged. Instead of emitting the value alone, emit a `(sum, count)` pair as a custom `Writable`. The combiner adds pairs componentwise — associative and commutative — and only the reducer performs the single division at the end. The same trick generalizes: variance is combinable if you carry `(count, sum, sum-of-squares)`; an approximate distinct count is combinable if you carry a sketch and merge sketches. The rule is to find the smallest *algebraic* summary that can be merged, then finalize once. ## What a combiner actually buys The payoff is proportional to how much duplication exists **within one mapper's output**. Word count over natural-language text is the textbook win: one split contains the word "the" thousands of times, and the combiner collapses that to a single record, cutting local spill volume, the bytes fetched in the copy phase, and the reduce-side merge work. If keys are nearly unique per mapper — session IDs, order IDs, UUIDs — the combiner reads and rewrites everything and saves nothing, so it costs CPU for no benefit. Watch the `COMBINE_INPUT_RECORDS` and `COMBINE_OUTPUT_RECORDS` counters: if they are nearly equal, remove the combiner. ## Combiner versus in-mapper aggregation A related technique is to aggregate inside the mapper itself — keep a small `HashMap` in the `Mapper` instance, update it in `map()`, and emit its contents in `cleanup()`. That guarantees the aggregation happens (no framework discretion) and skips the serialize-sort-deserialize round trip a combiner pays, at the cost of unbounded map-task memory if the key cardinality per split is large. Senior candidates are expected to know both and to say when the bounded, framework-managed combiner is the safer choice. ## The engine-comparison trap It is fair to observe that Spark's `reduceByKey` performs an equivalent map-side combine while `groupByKey` does not — that is the same idea under a different name. But do not describe Spark's behaviour as if the framework rules were identical: Spark decides map-side combining from the operator you chose, whereas MapReduce leaves it as a class you register and a decision the framework makes at runtime. ## What the interviewer is testing The question is rarely about combiners as trivia. It is a proxy for whether you understand that distributed aggregation only works for functions with the right algebra, and whether you can spot a correctness bug that no test on a small dataset will surface.
- When is it legitimate to pass your reducer class straight to setCombinerClass?Only when the reduce function is commutative and associative *and* its input and output types are identical to the map output types. Summing counts satisfies both, which is why word count reuses its reducer. As soon as the reducer's output type differs — a double average from integer counts — the reuse will not even type-check as a combiner.
- How would you tell from a finished job whether the combiner was worth having?Compare the `COMBINE_INPUT_RECORDS` and `COMBINE_OUTPUT_RECORDS` counters and look at map output bytes. A big drop between input and output records means real duplication was collapsed and the copy phase moved far fewer bytes. If the two counters are close, the combiner is burning CPU serializing and re-sorting records for no reduction.
- What is in-mapper combining and when would you prefer it?Keep an aggregation map in the `Mapper` instance, update it per record, and emit the aggregates in `cleanup()`. It always runs and avoids the combiner's serialize-sort-deserialize round trip, so it can be substantially faster. The tradeoff is unbounded memory: with high key cardinality per split the map task can OOM, while a combiner's memory is bounded by the sort buffer.
Averaging averages is like judging a school's mean exam score by averaging each classroom's mean: a class of five and a class of forty get equal weight, and the number you get belongs to no one.
saying these in an interview costs you the question
- Says a combiner is guaranteed to run exactly once per map task
- Claims a combiner reduces the number of reduce tasks needed
- Uses the averaging reducer as a combiner and calls it an optimization
- Thinks a combiner runs on the reduce side after fetching
- Believes a combiner can change the map output value type