In Hadoop, what do core-site.xml, hdfs-site.xml and yarn-site.xml each configure, and what overrides them?
answer
- four files, four scopes
- the jar ships the baseline
- who wins: shipped, deployed, or submitted?
- one flag stops a job overriding you
- nothing copies itself to the workers
basics
~20 score-site.xml holds cluster-wide settings like fs.defaultFS, hdfs-site.xml configures NameNode and DataNode behaviour, yarn-site.xml the ResourceManager and NodeManagers. Each overrides the matching *-default.xml shipped in the jars, and a job can override again unless the property is marked final.
solid answer
~40 sHadoop layers configuration. Every property has a shipped default in `core-default.xml`, `hdfs-default.xml`, `yarn-default.xml` or `mapred-default.xml` inside the jars; the matching `*-site.xml` on `$HADOOP_CONF_DIR` overrides it, and a client or job can override again through its own `Configuration` or `-D`. By scope: **core-site.xml** is cluster-wide (`fs.defaultFS`, `hadoop.tmp.dir`, `ha.zookeeper.quorum`, `hadoop.security.authentication`); **hdfs-site.xml** is HDFS (`dfs.replication`, `dfs.blocksize`, `dfs.namenode.name.dir`, `dfs.datanode.data.dir`, all the HA keys); **yarn-site.xml** is YARN (`yarn.resourcemanager.hostname`, `yarn.nodemanager.resource.memory-mb`, `yarn.nodemanager.aux-services`); **mapred-site.xml** covers MapReduce (`mapreduce.framework.name=yarn`). Wrapping a property in `<final>true</final>` stops jobs from overriding it. Nothing is distributed automatically — you push the files to every node and restart, and verify with `hdfs getconf -confKey <key>` or the daemon's `/conf` endpoint.
code
xml · 16 lines<!-- core-site.xml -->
<property>
<name>fs.defaultFS</name>
<value>hdfs://prod</value>
</property>
<property>
<name>ha.zookeeper.quorum</name>
<value>zk1:2181,zk2:2181,zk3:2181</value>
</property>
<!-- yarn-site.xml: locked so no job can widen it -->
<property>
<name>yarn.nodemanager.local-dirs</name>
<value>/data/1/nm,/data/2/nm</value>
<final>true</final>
</property>go deeper
Know the four site files by name and roughly what each covers, and that you edit *-site.xml, never the *-default.xml bundled in the jars.
Explain the override chain from defaults to site files to per-job settings, what final does, and the difference between client-side properties like dfs.blocksize and server-side ones like dfs.namenode.name.dir.
Talk about operating it: rolling restarts after a config change, which commands refresh without a restart, diagnosing drift when one node has stale config, and proving effective values from the /conf endpoint.
Own configuration as managed state — templated per node shape, version-controlled, reviewed — and decide which properties are locked final for tenants versus left to teams to tune.
## Four files, four scopes Hadoop configuration lives in XML files under `$HADOOP_CONF_DIR` (usually `etc/hadoop`), each a flat list of `<property><name>…</name><value>…</value></property>` entries. The split is by subsystem, not by node role — the same files are pushed to every machine: - **`core-site.xml`** — anything shared: `fs.defaultFS` (the default filesystem URI, e.g. `hdfs://prod`), `hadoop.tmp.dir`, `io.compression.codecs`, `ha.zookeeper.quorum`, `hadoop.security.authentication` and the proxyuser rules. - **`hdfs-site.xml`** — NameNode and DataNode: `dfs.replication`, `dfs.blocksize`, `dfs.namenode.name.dir`, `dfs.datanode.data.dir`, `dfs.namenode.handler.count`, and the whole HA block (`dfs.nameservices`, `dfs.ha.namenodes.*`, `dfs.namenode.shared.edits.dir`, fencing, the failover proxy provider). - **`yarn-site.xml`** — ResourceManager and NodeManagers: `yarn.resourcemanager.hostname` or the HA `rm-ids` set, `yarn.nodemanager.resource.memory-mb` and `.cpu-vcores`, `yarn.nodemanager.local-dirs` and `log-dirs`, `yarn.nodemanager.aux-services` (set to `mapreduce_shuffle` for MapReduce), and the scheduler class. - **`mapred-site.xml`** — the MapReduce framework: `mapreduce.framework.name=yarn`, `mapreduce.map.memory.mb`, `mapreduce.job.reduces`, JobHistoryServer addresses. Alongside them sit non-XML files: `hadoop-env.sh` (JVM-level environment such as `JAVA_HOME`, `HADOOP_HEAPSIZE_MAX`, `HDFS_NAMENODE_OPTS`), `workers` (host list used by the start scripts), `capacity-scheduler.xml`, and `log4j.properties`. ## The override chain Resolution order, weakest to strongest: 1. **`*-default.xml`**, bundled inside the Hadoop jars. You never edit these — they are the documented reference for what a property means and what its default is. 2. **`*-site.xml`** on the classpath, i.e. whatever is in `$HADOOP_CONF_DIR` on *that* machine. 3. **Per-application settings** — a `Configuration.set(...)` in code, `-D key=value` on the command line, or `-conf` pointing at another file. That last layer is why `<final>true</final>` exists. Marking a property final in a site file makes the daemon refuse a job's attempt to override it, and log a warning. Operators use it for things that must not be negotiable — security settings, `yarn.nodemanager.local-dirs`, memory caps. ## Client-side versus server-side properties This catches people out. Some properties are read by the **client** at the moment it acts: `dfs.replication` and `dfs.blocksize` are sent by the client when it creates a file, which is why two users on different edge nodes can write files with different replication into the same cluster, and why editing `dfs.replication` on the NameNode changes nothing about files that already exist (use `hdfs dfs -setrep` for those). Others are purely **server-side**: `dfs.namenode.name.dir` or `yarn.nodemanager.resource.memory-mb` are meaningless in a client's config and only take effect where the daemon reads them. ## Getting configuration onto the cluster Hadoop does not distribute its own configuration. Changing a property means shipping the file to every relevant host with your configuration-management tool and restarting the affected daemons — which is why a rolling restart plan matters on a live cluster. A handful of things can be refreshed without a restart: - `hdfs dfsadmin -refreshNodes` — re-reads the include/exclude host lists (used to decommission a DataNode). - `yarn rmadmin -refreshQueues` — re-reads `capacity-scheduler.xml`. - `hdfs dfsadmin -reconfig namenode <host:port> start` — applies the subset of properties marked reconfigurable at runtime. Everything else needs the daemon restarted, and a mismatched config across nodes (one worker still pointing at the old ResourceManager) produces some of the most confusing failures in Hadoop operations. ## Verifying what is actually in effect Never assume the file you edited is the file the daemon loaded. `hdfs getconf -confKey dfs.blocksize` prints the resolved value from the local client configuration. Each daemon's web UI exposes a `/conf` endpoint that dumps the effective configuration *and the file each value came from* — the fastest way to prove that your change reached the NameNode. `hadoop conftest` validates that the XML files parse. And remember that a job's own configuration is visible in the ApplicationMaster / JobHistory UI, which settles arguments about whether a value came from the cluster or from the submitting user. ## Common mistakes Editing `*-default.xml` instead of the site file; editing the site file on one node only; expecting `dfs.replication` to retro-apply; setting a per-node value like `yarn.nodemanager.resource.memory-mb` uniformly across heterogeneous hardware; and forgetting that `hadoop-env.sh` — not any XML file — is where daemon heap is set.
- An operator lowers dfs.replication from 3 to 2 in hdfs-site.xml and restarts. What happens to existing files?Nothing. Replication factor is per-file metadata, chosen by the client when the file is created, so existing files keep the factor they were written with. Only newly created files pick up 2. To change files already in HDFS you run `hdfs dfs -setrep -R 2 /path`, which asks the NameNode to schedule the extra replicas for deletion.
- What does wrapping a property in <final>true</final> achieve?It makes the value non-negotiable for jobs: an application that tries to override it through its own `Configuration` or `-D` is ignored and a warning is logged. Operators use it for settings a tenant must not weaken — security properties, local directory paths, container memory ceilings. It has no effect on the site-over-default layering itself.
- How do you confirm which value a running NameNode actually loaded?Open the daemon's `/conf` endpoint on its web UI: it dumps the effective configuration together with the source file for each property, so you can see whether your edit reached that host or whether a default is still winning. On the client side, `hdfs getconf -confKey <key>` resolves the value from the local configuration directory.
saying these in an interview costs you the question
- Edits hdfs-default.xml instead of hdfs-site.xml
- Expects a config change to propagate to workers by itself
- Thinks changing dfs.replication rewrites existing files
- Says daemon heap is set in an XML property rather than hadoop-env.sh
- Believes a job can never override a cluster property