Processing Engines
The engines that actually crunch large datasets — Spark and Flink for distributed compute, and the Hadoop stack that supplies the storage and scheduling underneath them. Data-engineering interviews spend most of their technical time here, because this is where batch and streaming pipelines either scale or fall over.
on this pageshowhide
explore
- Apache Spark61 questions
- RDDs and DataFrames6 questions
- Spark SQL and Catalyst6 questions
- Partitioning and Shuffles19 questions
- Structured Streaming6 questions
- Cluster Execution and Tuning18 questions
- MLlib Pipelines6 questions
- Hadoop30 questions
- HDFS Storage Model6 questions
- YARN Resource Management6 questions
- MapReduce Programming6 questions
- Cluster Topology & Ops6 questions
- Hadoop Ecosystem Tools6 questions
- HDFS2 questions
- YARN1 questions
- MapReduce1 questions
- Apache Flink (has its own guide)49 questions
- DataStream API6 questions
- Event Time & Watermarks6 questions
- Windowing6 questions
- State Management7 questions
- Checkpointing & Exactly-Once6 questions
- Deployment & Scaling6 questions
- Flink SQL & Table API6 questions
- Flink SQL Joins & Windows6 questions
- Processing Engine Concepts300 questions
- Choosing a Processing Engine24 questions
- Job as a Dataflow25 questions
- Splitting the Work28 questions
- Moving Data Between Workers27 questions
- Skew and Stragglers20 questions
- Memory, Spill and Caching20 questions
- Time in a Stream20 questions
- Bounding an Endless Stream21 questions
- Long-Lived State21 questions
- Surviving Failure22 questions
- Machines for the Job23 questions
- Operating Jobs in Production27 questions
- Testing and Changing Jobs22 questions
→ has its own guide
questions
444 · 7 sectionsIn spark-submit, what do the --master, --class and application-jar arguments specify?
basics
~20 sIn spark-submit, --master names the cluster manager and its address (local[*], yarn, k8s://https://host:6443, spark://host:7077), --class is the fully-qualified class holding main() inside the jar, and the trailing application jar is the code Spark ships to the cluster.
In Spark, what is the difference between a job, a stage and a task?
basics
~10 sIn Spark, an action submits a job; the driver cuts that job into stages at shuffle boundaries; each stage runs one task per partition. Tasks are the smallest unit executors actually execute.
In Spark, how do the MEMORY_ONLY and MEMORY_AND_DISK persist levels differ when a partition will not fit?
basics
~20 sWith MEMORY_ONLY, a partition that does not fit in storage memory is simply not cached and is recomputed from lineage the next time it is needed. With MEMORY_AND_DISK, that partition is written to the executor's local disk and read back instead of recomputed.
In Spark MLlib, what is the difference between a Transformer and an Estimator?
basics
~20 sA Transformer implements transform() and converts one DataFrame into another, usually by appending columns. An Estimator implements fit(), which learns from a DataFrame and returns a Model — and that Model is itself a Transformer.
In Spark, what is a partition and what decides how many partitions a DataFrame starts with?
basics
~20 sA partition is Spark's unit of parallelism: one partition is processed by one task on one core. The starting count comes from the input — file splits packed to spark.sql.files.maxPartitionBytes (128 MB by default) — not from a fixed number.
In a Hadoop 3 cluster, which daemons run on the master nodes and which run on every worker node?
basics
~20 sMaster nodes run the coordinators: the HDFS NameNode (plus JournalNodes and ZKFC when HA is on) and the YARN ResourceManager. Every worker runs a DataNode for storage and a NodeManager for compute, co-located so containers read local blocks.
In Hive, what happens to the data when you DROP a managed table versus an external table?
basics
~20 sDropping a managed table removes its metastore entry and deletes its data files. Dropping an external table removes only the metastore entry and leaves the files untouched. Declare EXTERNAL whenever another system owns the data.
In HDFS, what does the NameNode store and what do the DataNodes store?
basics
~10 sThe NameNode keeps the filesystem namespace in memory: directories, files, permissions, and which blocks make up each file. DataNodes store the actual block bytes on local disks. File data never flows through the NameNode.
What phases does a Hadoop MapReduce job move through from input split to final output?
basics
~20 sA Hadoop MapReduce job runs map, then shuffle and sort, then reduce. One map task handles each InputSplit; its output is partitioned and sorted on local disk, fetched by reducers over the network, merged by key, and written out by the OutputFormat.
In YARN, what do the ResourceManager, NodeManager and ApplicationMaster each do?
basics
~20 sYARN splits cluster management three ways: the ResourceManager schedules containers across the whole cluster, a NodeManager runs and monitors containers on each machine, and every application gets its own ApplicationMaster that requests containers and drives that job.
How does HDFS differ from cloud object storage like S3 for a Spark job's output?
basics
~20 sHDFS is a hierarchical filesystem where directory rename is an atomic metadata operation, so committing output is nearly free. S3 is a flat key store with no rename: a commit copies every object, so Spark needs a dedicated committer.
When would you still build a new platform on HDFS instead of object storage?
basics
~20 sRarely, and only on-premises: an air-gapped or regulated cluster, or a latency-sensitive workload like HBase where compute sits on the same disks as the data. New cloud platforms default to object storage plus an open table format.
When is YARN still the right cluster manager for Spark or Flink instead of Kubernetes?
basics
~20 sYARN still wins on an existing Hadoop estate: data already in HDFS on the same machines, Kerberos and queue policies already configured, and no container images to build or maintain. Greenfield platforms reading object storage normally choose Kubernetes.
Your platform still runs nightly MapReduce jobs on Hadoop 3 — how do you decide which to migrate and to what?
basics
~20 sRank jobs by cost and change rate, not by age. Multi-stage pipelines that materialize to HDFS between jobs gain most from Spark on the same YARN cluster; stable, cheap, rarely-touched jobs can stay on MapReduce indefinitely.
In Flink, what does enabling checkpointing do for a long-running streaming job?
basics
~20 sCheckpointing makes Flink periodically snapshot every operator's state and every source's read position to durable storage. On failure the job restarts from the last completed snapshot and replays from there, so accumulated state survives without being double-counted.
In Flink's DataStream API, what does keyBy() produce and why is it required before keyed state?
basics
~10 skeyBy() turns a DataStream into a KeyedStream by hash-partitioning records so every record sharing a key reaches the same parallel subtask. Keyed state, keyed timers and keyed windows exist only on a KeyedStream.
In a Flink cluster, what do the JobManager and the TaskManager each do?
basics
~20 sFlink's JobManager is the coordinator: it accepts submissions, turns a job into a schedule, allocates task slots and triggers checkpoints. TaskManagers are the worker processes that run the operator subtasks and exchange records directly with each other.
In Flink, what is the difference between event time and processing time?
basics
~20 sEvent time is the timestamp carried inside the record, set when the event happened. Processing time is the clock on the machine running the operator. Event time gives reproducible, order-independent results; processing time gives lower latency and no correctness guarantees.
In Flink SQL, how do you count ad impressions per campaign per minute with the TUMBLE window TVF, and what does the query emit?
basics
~10 sCall TUMBLE(TABLE impressions, DESCRIPTOR(impression_time), INTERVAL '1' MINUTE) in FROM and GROUP BY window_start, window_end, campaign_id. Each window emits one final insert-only row per campaign when it closes, then its state is purged.
The first 100,000 rows of an input file are used as a development cut. What does that cut hide about the whole input?
basics
~20 sThe first rows of a file are whatever was written first - one day, one source, one writer's share - so they carry neither the whole input's key distribution nor the rare record shapes that actually break the job.
Which part of a distributed job can a plain unit test call without a cluster, and what stops it?
basics
~20 sThe per-record rule can: a plain function taking ordinary values and returning ordinary values. It stops being callable once its signature mentions a runtime type, or it reads configuration, storage or the clock itself instead of taking them as arguments.
Why is sleeping in a test a poor way to prove a job groups records by when they happened?
basics
~20 sSleeping proves only that the machine's wall clock advanced. Instead hand the job records stamped with chosen moments and advance from the test the job's own claim that nothing older will arrive, so groups close on command.
Why does comparing a distributed job's output line by line against a stored expected result fail even when every value is right?
basics
~10 sA distributed run promises the right records, not an order, a file layout or bit-identical arithmetic. A line-by-line comparison asserts all three at once, so it fails on arrangement while every value is correct.
What does a cluster engine know about a named operator that it does not know about a function you hand it to run per record?
basics
~20 sA named operator carries its meaning - the engine knows it keeps rows, or groups by a key - so it may reorder, narrow, fuse or skip it. A handed-over body means only 'call this per record', so it runs exactly where it stands.