How do you decide which legacy Hadoop MapReduce pipelines to rewrite on Spark and which to leave alone?
answer
- not every job pays the same tax
- count the passes over the data
- some of it is a config change
- dual-run before you retire anything
- leave the cheap tail alone
basics
~20 sRank jobs by how much the MapReduce model actually costs them. Multi-job chains and iterative algorithms pay an HDFS round trip between every stage and repay a rewrite; stable single-pass map-only jobs are already IO-bound and rarely justify the risk.
solid answer
~50 sScore each pipeline on **shape**, **pain** and **risk**. Shape: a chain of jobs or an iterative algorithm materializes its intermediate result to HDFS — replicated, by default three ways — between every pass, restarts a JVM for every task and submits a fresh YARN application each time, so collapsing it into one DAG in a single application is where the large wins are. A single-pass map-only job that reads once and writes once is already near the IO floor and will gain little. Pain: runtime against its SLA, cost, and how often it breaks. Risk: business criticality, test coverage, and whether you can dual-run old and new and diff outputs for a period. Then check for cheaper moves than a rewrite — Hive and Pig workloads switch execution engines by configuration rather than by rewriting logic, and shipped tools such as `distcp` should be left alone. Both engines run on the same YARN cluster, so migrate incrementally rather than as a cutover.
code
text · 9 linesChained MapReduce, three passes over one dataset
job 1: read HDFS -> map/shuffle/reduce -> write HDFS (replicated)
job 2: read HDFS -> map/shuffle/reduce -> write HDFS (replicated)
job 3: read HDFS -> map/shuffle/reduce -> write HDFS (final)
= 3 YARN applications, 3 HDFS reads, 3 replicated HDFS writes, fresh JVM per task
Same logic as one Spark application
read HDFS -> narrow ops pipelined -> shuffle (local disk) -> ... -> write HDFS
= 1 application, 1 HDFS read, 1 HDFS write, dataset reusable across iterationsgo deeper
Recall the core reason a chain of MapReduce jobs is slow: each job writes its intermediate result to HDFS and the next job reads it back, instead of keeping data moving inside one application.
Explain the mechanics behind the cost — replicated intermediate writes, a fresh JVM per task, one YARN application per job, no reuse of a cached dataset — and note that a single map-shuffle-reduce pass narrows the gap considerably.
Show you would profile the estate before proposing anything: job chain lengths, cluster hours, shuffle volume, SLA headroom. Then propose a dual-run verification plan and migrate the expensive pipelines first rather than starting wherever the code is easiest.
Own the whole tradeoff: which parts of the estate you will deliberately never migrate, the cost of maintaining two operating models, whether the real requirement is lower latency rather than faster batch, and how you keep a partially migrated platform from becoming permanent debt.
## Start from what the model actually costs The question is not "is Spark faster" but "how much does this specific pipeline pay for MapReduce's shape". Four costs are structural: **Materialization between jobs.** MapReduce has no way to hand a result to the next computation except through the filesystem. A three-job chain writes its intermediate results to HDFS twice, and every one of those writes is replicated (`dfs.replication` defaults to 3), then read back in full. A DAG engine keeps the whole chain inside one application, pipelines narrow operations within a stage, and only spills at shuffle boundaries — to local disk, unreplicated. **Process startup.** Each MapReduce task runs in its own container and its own fresh JVM. For jobs made of thousands of short tasks, or for iterative algorithms that run the same tiny job repeatedly, JVM start and container allocation can dominate real work. Spark's executors are long-lived within an application and run many tasks per JVM. **No reuse of a hot dataset.** An iterative algorithm — anything that sweeps the same dataset repeatedly, such as a training loop or a graph traversal — re-reads it from HDFS every pass. Being able to `persist` it once and reuse it is precisely the case Spark was built for, and the case where the difference is largest. **Expressiveness.** Multi-way joins, windowed aggregations and anything with an optimizer behind it are painful hand-written MapReduce and ordinary SQL or DataFrame code elsewhere. That is a maintenance cost that compounds, even where runtime is acceptable. Be honest about the limits of the argument. Both engines shuffle, and both write shuffle data to local disk; for a single map-shuffle-reduce pass over a large dataset the gap is far narrower than the marketing suggests. The dramatic differences come from chains and iteration. ## Build the inventory before you build the plan Pull the job history and characterize each pipeline: number of jobs in the chain, wall-clock time, cluster hours consumed, shuffle volume, failure rate, SLA headroom, and how many people understand the code. That inventory usually shows the familiar shape — a small number of pipelines consuming most of the cluster, and a long tail of small jobs that cost almost nothing and are therefore not worth touching. ## The decision axes - **Rewrite first:** long multi-job chains, iterative algorithms, pipelines missing their SLA, pipelines whose logic changes frequently, and anything already forcing hand-rolled joins. - **Rewrite later or never:** single-pass map-only jobs (filters, conversions, copies) that are IO-bound and stable; jobs that run rarely and cheaply; jobs slated for decommission. - **Do not rewrite — reconfigure:** Hive and Pig workloads where the logic is declarative and only the execution engine underneath needs to change; shipped Hadoop tools such as `distcp` that happen to be MapReduce internally. - **Reconsider the requirement entirely:** if a nightly batch chain exists to reduce latency, the right target may be a streaming engine rather than a faster batch engine, which is a product decision before a technology one. ## Managing the risk A rewrite is a correctness event. The discipline that makes it survivable is dual-running: keep the MapReduce pipeline in production, run the new implementation on the same inputs, and diff outputs — full equality where feasible, otherwise row counts, per-key aggregates and distribution checks — for enough cycles to cover month-end and other edge conditions. Only then retire the old job. Freeze the logic during the migration; combining a rewrite with a feature change makes every difference ambiguous. Migrate the boundaries of a chain last, so downstream consumers see a stable contract. ## The organizational half Two engines mean two sets of tuning knowledge, two failure modes and two on-call runbooks, so a migration that stalls halfway is more expensive than either end state. Decide up front whether the goal is *all* pipelines or only the expensive tail, and if it is the tail, say so explicitly so no one treats the remainder as unfinished work. Weigh team skills honestly: a team fluent in MapReduce internals but new to executor memory, partitioning and shuffle tuning will trade one class of production incident for another until they build that experience. Since both engines run on the same YARN cluster and read the same HDFS data, incremental migration is genuinely available — use it rather than planning a cutover. ## What the interviewer is listening for Not "Spark is faster". They want to hear that you would measure, that you can name *why* certain job shapes pay more than others, that you know cheaper alternatives to rewriting exist, that you have a correctness-verification plan, and that you would deliberately choose to leave part of the estate alone.
- Why is an iterative algorithm the strongest case for moving off MapReduce?Each iteration is a separate job that reads its input from HDFS and writes the result back replicated, so a twenty-iteration algorithm makes twenty full round trips through the filesystem and pays fresh JVM startup for every task. A DAG engine reads the dataset once, keeps it cached across iterations, and runs the whole loop inside one application, which removes the dominant cost rather than shaving it.
- Which MapReduce workloads gain the least from a rewrite?Single-pass map-only jobs — filters, format conversions, projections, bulk copies. They have no shuffle and no intermediate materialization, so their runtime is essentially the cost of reading and writing the data, which any engine must also pay. Rewriting them buys a marginal gain and spends the same review, testing and dual-run effort as a valuable migration would.
- How do you verify a rewritten pipeline produces the same results?Dual-run: keep the original in production, run the new implementation on identical inputs to a shadow location, and compare. Prefer full row-level diffs where volume permits; otherwise compare row counts, per-key aggregates, null rates and value distributions. Run through at least one full business cycle including month-end and a backfill, and freeze the logic so any difference is a defect rather than an intended change.
- When is switching execution engines cheaper than rewriting?When the logic is already declarative. Hive and Pig workloads express intent rather than map and reduce functions, so the same queries can run over a different execution engine with a configuration change and no rewrite of business logic. You still need dual-run verification, but you avoid re-deriving semantics from imperative Java, which is where migration defects usually originate.
saying these in an interview costs you the question
- Says rewrite everything because MapReduce is obsolete
- Claims Spark avoids disk entirely, so it is always faster
- Ignores dual-running and correctness verification
- Treats a rewrite as free because the logic already exists
- Overlooks that Hive and Pig can change engine without rewriting