skip to content

Crawlers and the Data Catalog

The metadata half of Glue: crawlers infer schema and partitions from files in S3, and the resulting catalog is what Athena, EMR and Redshift Spectrum query against. Asked because a mis-crawled partition scheme is the usual reason a lake query returns nothing.

on this pageshow

questions

6

What does an AWS Glue crawler create in the Data Catalog after it scans an S3 prefix?

level: juniorimportance: must knowfreq 68%

answer

  1. metadata only, no bytes move
  2. something has to tell Athena what the files are
  3. a database, a table, and one row per folder
  4. columns, location, SerDe, partition keys
  5. classifier decides the format first

basics

~20 s

A Glue crawler produces metadata, never data. It creates or updates a table inside a Data Catalog database — inferred columns and types, the S3 location, the format's SerDe — plus one partition entry for each detected sub-prefix.

solid answer

~50 s

An AWS Glue crawler points at one or more targets (usually an S3 prefix), assumes an IAM role that can read those objects, and samples files to work out what they are. A **classifier** identifies the format — CSV, JSON, Parquet, ORC, Avro and others have built-in classifiers, and you can add custom ones — and from that the crawler infers a column list and types. It then writes a **table definition** into the Data Catalog database you nominated: columns, partition keys, the S3 `Location`, the input/output formats and SerDe, and properties such as `classification`. If the prefix has sub-folders that look like partitions, it also writes a partition entry per folder. Nothing is copied or converted: the files stay exactly where they are, and Athena, EMR and Redshift Spectrum read that one shared table definition.

code

text · 11 lines
text
s3://my-lake/events/dt=2026-08-19/part-0000.parquet
s3://my-lake/events/dt=2026-08-20/part-0000.parquet

-- after crawling s3://my-lake/events/
database: analytics
table:    events
  columns:        user_id bigint, action string, ts timestamp
  partition keys: dt string
  location:       s3://my-lake/events/
  serde:          parquet
  partitions:     dt=2026-08-19, dt=2026-08-20

go deeper

for a junior

Be able to say in one breath what a crawler produces: a database, a table with columns and a location, and partition entries — metadata only. Know that Athena then queries that table.

for a middle

Explain the pipeline inside a run: classifier picks the format, schema is inferred and reconciled across files, folders are grouped into tables, and the table records SerDe, location and partition keys.

for a senior

Show judgment about when a crawler is the wrong tool — cost on huge prefixes, the lag between a write and the partition appearing, and inferred types that shift under downstream consumers.

for a principal

Own the question of who is authoritative for schema across the lake: crawler inference for unknown third-party drops versus declared schema-as-code for tables your own pipelines produce.

## What a crawler actually is An AWS Glue crawler is a managed metadata scanner. You give it a **target** (most often an S3 prefix such as `s3://my-lake/events/`, but JDBC databases, DynamoDB tables and others are supported), an **IAM role** that can read the target, and a **destination database** in the AWS Glue Data Catalog. When it runs, it reads a sample of the objects it finds and writes table metadata. It never moves, copies, rewrites or converts the underlying data. This is the single most important mental model for the service: **Glue has two halves**, and the crawler belongs entirely to the metadata half. Glue *jobs* move and transform data; crawlers only describe what already exists. ## The run, step by step 1. **List the target.** The crawler enumerates objects under the prefix, applying any exclude patterns you configured (useful for `_temporary/`, `_SUCCESS` and other writer debris). 2. **Classify.** Each file is offered to a chain of **classifiers**. Built-in classifiers recognise Parquet, ORC, Avro, JSON, CSV, XML and several log formats; custom classifiers (grok, CSV with a declared delimiter, JSON path, XML) run first if you attach them. The winning classifier decides the format and yields a schema. 3. **Infer and reconcile a schema.** For self-describing formats such as Parquet or Avro the schema is read from the footer or header. For JSON and CSV the crawler infers types from sampled rows. Schemas from many files under the same folder are reconciled into one. 4. **Group folders into tables.** The crawler decides whether sub-folders are partitions of one table or separate tables, based on whether their schemas and structure are compatible. 5. **Write to the catalog.** It creates or updates the table and adds partition entries. ## What a catalog table actually holds A Data Catalog table is a Hive-metastore-shaped record: - `Name` and the owning `DatabaseName` - `StorageDescriptor.Columns` — the non-partition columns and their types - `PartitionKeys` — the columns derived from folder structure - `StorageDescriptor.Location` — the S3 prefix - `InputFormat` / `OutputFormat` and `SerdeInfo.SerializationLibrary` — how a reader deserialises the bytes - `Parameters` — free-form properties such as `classification=parquet`, `compressionType`, row counts and `CrawlerSchemaSerializerVersion` Each partition is its own record with its own `Location` and storage descriptor, which is why a partition can technically point somewhere outside the table's prefix and can carry a schema that differs from the table's. ## Who reads it The Data Catalog is a per-account, per-Region metastore that speaks the Hive metastore model, so several engines consume it without any further schema definition: - **Amazon Athena** uses it as its metastore outright — a crawled table appears in the Athena query editor. - **Amazon EMR** can be configured to use Glue as its Hive metastore instead of a local one. - **Amazon Redshift Spectrum** builds an external schema from a Glue database. - **AWS Glue ETL jobs** read tables from it via `GlueContext`. That sharing is the point: one crawl, and every engine agrees on what the files mean. ## What it does not do - It does not create data, compact files, or change formats. - It does not run continuously — it runs on demand, on a schedule, or triggered by S3 events if you configure event-based crawling. - It is not mandatory. You can create the same table by hand with `CREATE EXTERNAL TABLE` in Athena or with the `CreateTable` API, and for tables your own pipelines write that is often the better choice, because a crawler infers types from a sample and can therefore change its mind. ## Practical notes The crawler's IAM role needs read access to the S3 objects **and** Glue permissions to create and update tables and partitions; a KMS-encrypted bucket also requires `kms:Decrypt`. Crawlers are billed for the time they run, so pointing one at a prefix containing millions of small files and scheduling it every five minutes is a common way to be surprised by the bill. Start with a manual run, inspect the resulting table with `aws glue get-table`, and only then schedule it.

  • Do you have to use a crawler at all to get a table into the Glue Data Catalog?
    No. The catalog is just a metastore, so you can create the table with `CREATE EXTERNAL TABLE` in Athena, with the Glue `CreateTable` API, or from infrastructure-as-code. For tables your own pipelines produce, declaring the schema explicitly is usually safer than letting a crawler infer it, because inference from a file sample can silently change a column's type between runs.
  • What does a classifier do that a SerDe does not?
    A classifier runs at crawl time and answers "what format is this file, and what schema does it imply?" — its output is the table definition. A SerDe runs at query time inside the engine and answers "how do I turn these bytes into rows?" The crawler picks the SerDe and records it in the table; after that the classifier is out of the picture.
  • Why would a Glue crawler need kms:Decrypt in its IAM role?
    Because it reads the objects themselves to infer a schema. If the S3 objects are encrypted with SSE-KMS using a customer managed key, the crawler role must be allowed `kms:Decrypt` on that key — and the key policy must allow the role — or the crawl fails or produces an empty table.

A crawler is a librarian walking the stacks: it does not move a single book, it fills in the card catalogue so that anyone who comes to the desk can find and read them.

saying these in an interview costs you the question

  • Saying the crawler copies or converts the data into Glue
  • Thinking rows are stored inside the Data Catalog table
  • Assuming a crawler is required before Athena can query S3
  • Believing the crawler runs continuously and keeps the table live
  • Confusing a crawler with a Glue ETL job that transforms data

context

open as a page

Athena returns zero rows for a partition whose files are in S3 — what went wrong in the Glue Data Catalog?

level: middleimportance: must knowfreq 74%

basics

~20 s

Almost always the partition is not registered. Athena reads only partition entries in the Glue Data Catalog, so a new prefix nobody added is invisible. Re-run the crawler, run MSCK REPAIR TABLE, add the partition explicitly, or use partition projection.

open as a page

How does an AWS Glue crawler decide which parts of an S3 key prefix become partition columns?

level: middleimportance: should knowfreq 52%

basics

~20 s

It walks the folder levels below its target prefix. Hive-style folders named key=value become a partition column called key; any other consistent folder level becomes an unnamed column partition_0, partition_1 and so on, in depth order.

open as a page

An AWS Glue crawler retyped a table column on its next run and broke Athena queries — which crawler settings govern that?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The crawler's schema change policy. UpdateBehavior decides whether a re-crawl rewrites the existing table (UPDATE_IN_DATABASE) or only logs the difference (LOG); DeleteBehavior decides what happens to tables and partitions whose data vanished. Every write also creates a new table version.

open as a page

How would you decide whether scheduled AWS Glue crawlers or the writing pipelines maintain a lake's Data Catalog?

level: principalimportance: should knowfreq 29%

basics

~20 s

Split by ownership. Crawlers are discovery tools for data you did not produce; for tables your own pipelines write, register partitions at write time and declare the schema as code, so metadata is never stale, inferred or surprising.

open as a page

In AWS Lake Formation, why does an IAM role with full Glue and S3 access still get access denied in Athena?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Lake Formation adds a second authorization layer over Data Catalog databases, tables and columns. Once a location is registered with it, IAM alone is not enough — the principal also needs an explicit Lake Formation grant such as SELECT on that table.

open as a page