Your platform still runs nightly MapReduce jobs on Hadoop 3 — how do you decide which to migrate and to what?
answer
- the cluster does not have to go
- count the round trips to storage
- two axes: cost and change rate
- chained jobs first, stable ones never
- same YARN, new framework, job by job
basics
~20 sRank jobs by cost and change rate, not by age. Multi-stage pipelines that materialize to HDFS between jobs gain most from Spark on the same YARN cluster; stable, cheap, rarely-touched jobs can stay on MapReduce indefinitely.
solid answer
~50 sI would not treat this as a platform replacement. Spark runs as a YARN application on the same Hadoop 3 cluster and reads the same HDFS data, so migration is job-by-job with no cutover event. I rank candidates by two axes: **runtime cost** and **change rate**. The biggest wins are pipelines chained as several MapReduce jobs, because each job materializes its output to HDFS and the next re-reads it from replicated storage — collapsing that chain into one Spark application removes those round trips. The second group is jobs under active development, where the DataFrame or SQL API pays back in maintenance. Jobs that are stable, cheap and untouched for years I leave alone; rewriting them buys risk, not throughput. For SQL-shaped work the cheaper move is often switching the execution engine under Hive rather than hand-writing Spark. I sequence with a correctness gate: run both versions over the same input and diff outputs before retiring the old job.
code
bash · 6 lines# legacy: two chained MapReduce jobs, each materializing to HDFS
hadoop jar pipeline.jar StageOne /data/in /data/tmp/stage1
hadoop jar pipeline.jar StageTwo /data/tmp/stage1 /data/out
# same pipeline as one Spark application on the same YARN cluster
spark-submit --master yarn --deploy-mode cluster pipeline.py /data/in /data/outgo deeper
Know that MapReduce and Spark are two processing frameworks that can run on the same Hadoop cluster over the same HDFS data, and that Spark is what new batch work is normally written in today.
Be ready to explain the mechanical reason a chain of MapReduce jobs is slow: each job writes its output to HDFS and the next reads it back, whereas one Spark application runs multiple stages without that round trip.
Demonstrate that you would rank the job portfolio by runtime cost and change rate, migrate incrementally on the existing YARN cluster, and gate each cutover on a parallel-run output diff rather than a scheduled switchover.
Own the argument that a partial migration is the correct end state. Be able to price the residual legacy footprint — maintainer skills, cluster capacity, risk — against the rewrite cost, and defend deliberately leaving the tail of the portfolio where it is.
## What this question is really testing This is a modernization judgment question, not a MapReduce mechanics question. The interviewer wants to see whether you treat a legacy Hadoop platform as something to be replaced wholesale (expensive, risky, usually stalls halfway) or as something to be drained job by job while the cluster keeps running. The strongest answers are boring: measure, rank, migrate the top of the list, leave the tail alone. ## The key structural fact: it is not a cluster migration Apache Spark can run as a YARN application. On a Hadoop 3 cluster you submit it with `spark-submit --master yarn`, YARN's ResourceManager allocates containers for it exactly as it does for a MapReduce ApplicationMaster, and it reads and writes the same HDFS paths. That means MapReduce jobs and Spark jobs coexist on one cluster indefinitely, sharing the same scheduler queues and the same data. Nothing forces a cutover date, and nothing forces you to abandon HDFS. Candidates who answer "first we move off HDFS to object storage, then we move off MapReduce" have merged two independent decisions and made the project ten times larger than it needs to be. ## Why chained jobs are the top of the list A MapReduce job ends by writing its output to HDFS, which means replicated, durable, disk-backed storage. A pipeline expressed as three MapReduce jobs therefore materializes two intermediate datasets to HDFS and reads them back. Spark expresses the same pipeline as one application whose DAG contains several stages; data crosses a stage boundary through a shuffle to local disk, but the pipeline does not round-trip through replicated storage between steps, and results a later step reuses can be held with `persist`/`cache`. Be careful with the popular shorthand here. "Spark is faster because it is in memory" is only half true and interviewers push on it: Spark's sort-based shuffle also writes to local disk, and a shuffle-heavy Spark job on a memory-starved executor spills just like anything else. The reliable, defensible saving is the removal of inter-job HDFS materialization and of per-job JVM and scheduling overhead, plus a planner (Catalyst) that can push filters and prune columns that hand-written MapReduce code cannot. ## Ranking the portfolio A workable ordering: 1. **Multi-job pipelines with long wall-clock time.** Largest, most measurable win. 2. **Jobs under active change.** Every feature request is currently paid in hand-written mapper and reducer code; a DataFrame or SQL rewrite pays that back repeatedly. 3. **Jobs whose logic is already SQL-shaped.** Often these are Hive queries, and the cheapest move is changing the execution engine under Hive rather than writing a Spark application at all. 4. **Everything else — leave it.** A job that runs in four minutes once a night, has not been edited in three years, and nobody is paged about is not a migration candidate. Rewriting it consumes engineer time and introduces a correctness risk in exchange for nothing. ## Sequencing and the correctness gate Run the new implementation in parallel against the same input for a period and diff the outputs, rather than cutting over and hoping. Byte-identical output is rare and not the bar: floating-point aggregation order, null and empty-string handling, and default date parsing all differ between hand-written Java and a DataFrame API, so agree in advance on what counts as equivalent. Only after the diff is clean do you retire the old job — and keep the old code in version control, because the fastest rollback is resubmitting it. ## What to say about staying It is a strong answer, not a weak one, to say that some MapReduce stays. Hadoop 3 still ships MapReduce and still runs it; there is no forced end-of-life pushing you. The costs of the legacy footprint are real but bounded: engineers who can maintain the code, and cluster capacity. State them, then argue that they are cheaper than a rewrite for the tail of the portfolio. ## Version note This assumes Hadoop 3.x with Spark 3.5 or later submitted to YARN. On a Kubernetes-based platform the same ranking logic applies, but the coexistence argument does not — there you really are moving the workload to a different substrate, and the migration is correspondingly larger.
- Why is "Spark is faster because it keeps data in memory" an incomplete justification for the migration?Spark's sort-based shuffle writes to local disk, and a memory-starved executor spills, so Spark is not disk-free. The defensible savings are removing inter-job materialization to HDFS, removing per-job JVM and scheduling overhead, and gaining a planner that prunes columns and pushes filters. Framing it as "in memory" invites an interviewer to ask why your Spark job is still slow.
- How would you validate that a rewritten job produces the same results as the MapReduce version?Run both over identical input for several cycles and diff the outputs before retiring the old job. Byte equality is rarely the right bar: aggregation order for floating point, null versus empty-string handling, and date parsing defaults legitimately differ. Agree the equivalence criteria up front, keep the old code in version control, and treat resubmitting it as the rollback plan.
- What argues for leaving some jobs on MapReduce permanently?Hadoop 3 still ships and runs MapReduce, so nothing forces an end-of-life. For a job that is cheap, stable, and untouched for years, a rewrite spends engineering time and adds correctness risk for no measurable gain. The real costs of the residue are maintainer skills and cluster capacity; name them explicitly and argue they are cheaper than the rewrite.
It is like replacing a building's plumbing while people still live there: you re-pipe the bathrooms that leak and get used every day, and you leave the tap in the garage alone, because shutting off the whole building to do it in one weekend is what turns a repair into a disaster.
saying these in an interview costs you the question
- Claims MapReduce was removed from Hadoop 3 and must be migrated
- Says HDFS must be replaced before Spark can be adopted
- Proposes a big-bang rewrite of every job in one cutover
- Justifies the move purely as 'Spark runs in memory, so no disk'
- Migrates stable, cheap, rarely-run jobs first for no measured gain