skip to content

Processing Engines

The engines that actually crunch large datasets — Spark and Flink for distributed compute, and the Hadoop stack that supplies the storage and scheduling underneath them. Data-engineering interviews spend most of their technical time here, because this is where batch and streaming pipelines either scale or fall over.

on this pageshow

explore

→ has its own guide

questions

444 · 7 sections

In spark-submit, what do the --master, --class and application-jar arguments specify?

level: juniorimportance: must knowfreq 55%
basics
~20 s

In spark-submit, --master names the cluster manager and its address (local[*], yarn, k8s://https://host:6443, spark://host:7077), --class is the fully-qualified class holding main() inside the jar, and the trailing application jar is the code Spark ships to the cluster.

open as a page

In Spark, what is the difference between a job, a stage and a task?

level: juniorimportance: must knowfreq 82%
basics
~10 s

In Spark, an action submits a job; the driver cuts that job into stages at shuffle boundaries; each stage runs one task per partition. Tasks are the smallest unit executors actually execute.

open as a page

In Spark, how do the MEMORY_ONLY and MEMORY_AND_DISK persist levels differ when a partition will not fit?

level: juniorimportance: must knowfreq 70%
basics
~20 s

With MEMORY_ONLY, a partition that does not fit in storage memory is simply not cached and is recomputed from lineage the next time it is needed. With MEMORY_AND_DISK, that partition is written to the executor's local disk and read back instead of recomputed.

open as a page

In Spark MLlib, what is the difference between a Transformer and an Estimator?

level: juniorimportance: must knowfreq 55%
basics
~20 s

A Transformer implements transform() and converts one DataFrame into another, usually by appending columns. An Estimator implements fit(), which learns from a DataFrame and returns a Model — and that Model is itself a Transformer.

open as a page

In Spark, what is a partition and what decides how many partitions a DataFrame starts with?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A partition is Spark's unit of parallelism: one partition is processed by one task on one core. The starting count comes from the input — file splits packed to spark.sql.files.maxPartitionBytes (128 MB by default) — not from a fixed number.

open as a page

In a Hadoop 3 cluster, which daemons run on the master nodes and which run on every worker node?

level: juniorimportance: must knowfreq 52%
basics
~20 s

Master nodes run the coordinators: the HDFS NameNode (plus JournalNodes and ZKFC when HA is on) and the YARN ResourceManager. Every worker runs a DataNode for storage and a NodeManager for compute, co-located so containers read local blocks.

open as a page

In Hive, what happens to the data when you DROP a managed table versus an external table?

level: juniorimportance: must knowfreq 50%
basics
~20 s

Dropping a managed table removes its metastore entry and deletes its data files. Dropping an external table removes only the metastore entry and leaves the files untouched. Declare EXTERNAL whenever another system owns the data.

open as a page

In HDFS, what does the NameNode store and what do the DataNodes store?

level: juniorimportance: must knowfreq 60%
basics
~10 s

The NameNode keeps the filesystem namespace in memory: directories, files, permissions, and which blocks make up each file. DataNodes store the actual block bytes on local disks. File data never flows through the NameNode.

open as a page

What phases does a Hadoop MapReduce job move through from input split to final output?

level: juniorimportance: must knowfreq 55%
basics
~20 s

A Hadoop MapReduce job runs map, then shuffle and sort, then reduce. One map task handles each InputSplit; its output is partitioned and sorted on local disk, fetched by reducers over the network, merged by key, and written out by the OutputFormat.

open as a page

In YARN, what do the ResourceManager, NodeManager and ApplicationMaster each do?

level: juniorimportance: must knowfreq 58%
basics
~20 s

YARN splits cluster management three ways: the ResourceManager schedules containers across the whole cluster, a NodeManager runs and monitors containers on each machine, and every application gets its own ApplicationMaster that requests containers and drives that job.

open as a page

How does HDFS differ from cloud object storage like S3 for a Spark job's output?

level: middleimportance: should knowfreq 42%
basics
~20 s

HDFS is a hierarchical filesystem where directory rename is an atomic metadata operation, so committing output is nearly free. S3 is a flat key store with no rename: a commit copies every object, so Spark needs a dedicated committer.

open as a page

When would you still build a new platform on HDFS instead of object storage?

level: principalimportance: nice to knowfreq 30%
basics
~20 s

Rarely, and only on-premises: an air-gapped or regulated cluster, or a latency-sensitive workload like HBase where compute sits on the same disks as the data. New cloud platforms default to object storage plus an open table format.

open as a page

When is YARN still the right cluster manager for Spark or Flink instead of Kubernetes?

level: seniorimportance: should knowfreq 40%
basics
~20 s

YARN still wins on an existing Hadoop estate: data already in HDFS on the same machines, Kerberos and queue policies already configured, and no container images to build or maintain. Greenfield platforms reading object storage normally choose Kubernetes.

open as a page

Your platform still runs nightly MapReduce jobs on Hadoop 3 — how do you decide which to migrate and to what?

level: seniorimportance: should knowfreq 35%
basics
~20 s

Rank jobs by cost and change rate, not by age. Multi-stage pipelines that materialize to HDFS between jobs gain most from Spark on the same YARN cluster; stable, cheap, rarely-touched jobs can stay on MapReduce indefinitely.

open as a page

The first 100,000 rows of an input file are used as a development cut. What does that cut hide about the whole input?

level: juniorimportance: must knowfreq 56%
basics
~20 s

The first rows of a file are whatever was written first - one day, one source, one writer's share - so they carry neither the whole input's key distribution nor the rare record shapes that actually break the job.

open as a page

Which part of a distributed job can a plain unit test call without a cluster, and what stops it?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The per-record rule can: a plain function taking ordinary values and returning ordinary values. It stops being callable once its signature mentions a runtime type, or it reads configuration, storage or the clock itself instead of taking them as arguments.

open as a page

Why is sleeping in a test a poor way to prove a job groups records by when they happened?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Sleeping proves only that the machine's wall clock advanced. Instead hand the job records stamped with chosen moments and advance from the test the job's own claim that nothing older will arrive, so groups close on command.

open as a page

Why does comparing a distributed job's output line by line against a stored expected result fail even when every value is right?

level: juniorimportance: must knowfreq 70%
basics
~10 s

A distributed run promises the right records, not an order, a file layout or bit-identical arithmetic. A line-by-line comparison asserts all three at once, so it fails on arrangement while every value is correct.

open as a page

What does a cluster engine know about a named operator that it does not know about a function you hand it to run per record?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A named operator carries its meaning - the engine knows it keeps rows, or groups by a key - so it may reorder, narrow, fuse or skip it. A handed-over body means only 'call this per record', so it runs exactly where it stands.

open as a page