What does an AWS Glue crawler create in the Data Catalog after it scans an S3 prefix?
answer
- metadata only, no bytes move
- something has to tell Athena what the files are
- a database, a table, and one row per folder
- columns, location, SerDe, partition keys
- classifier decides the format first
basics
~20 sA Glue crawler produces metadata, never data. It creates or updates a table inside a Data Catalog database — inferred columns and types, the S3 location, the format's SerDe — plus one partition entry for each detected sub-prefix.
solid answer
~50 sAn AWS Glue crawler points at one or more targets (usually an S3 prefix), assumes an IAM role that can read those objects, and samples files to work out what they are. A **classifier** identifies the format — CSV, JSON, Parquet, ORC, Avro and others have built-in classifiers, and you can add custom ones — and from that the crawler infers a column list and types. It then writes a **table definition** into the Data Catalog database you nominated: columns, partition keys, the S3 `Location`, the input/output formats and SerDe, and properties such as `classification`. If the prefix has sub-folders that look like partitions, it also writes a partition entry per folder. Nothing is copied or converted: the files stay exactly where they are, and Athena, EMR and Redshift Spectrum read that one shared table definition.
code
text · 11 liness3://my-lake/events/dt=2026-08-19/part-0000.parquet
s3://my-lake/events/dt=2026-08-20/part-0000.parquet
-- after crawling s3://my-lake/events/
database: analytics
table: events
columns: user_id bigint, action string, ts timestamp
partition keys: dt string
location: s3://my-lake/events/
serde: parquet
partitions: dt=2026-08-19, dt=2026-08-20go deeper
Be able to say in one breath what a crawler produces: a database, a table with columns and a location, and partition entries — metadata only. Know that Athena then queries that table.
Explain the pipeline inside a run: classifier picks the format, schema is inferred and reconciled across files, folders are grouped into tables, and the table records SerDe, location and partition keys.
Show judgment about when a crawler is the wrong tool — cost on huge prefixes, the lag between a write and the partition appearing, and inferred types that shift under downstream consumers.
Own the question of who is authoritative for schema across the lake: crawler inference for unknown third-party drops versus declared schema-as-code for tables your own pipelines produce.
## What a crawler actually is An AWS Glue crawler is a managed metadata scanner. You give it a **target** (most often an S3 prefix such as `s3://my-lake/events/`, but JDBC databases, DynamoDB tables and others are supported), an **IAM role** that can read the target, and a **destination database** in the AWS Glue Data Catalog. When it runs, it reads a sample of the objects it finds and writes table metadata. It never moves, copies, rewrites or converts the underlying data. This is the single most important mental model for the service: **Glue has two halves**, and the crawler belongs entirely to the metadata half. Glue *jobs* move and transform data; crawlers only describe what already exists. ## The run, step by step 1. **List the target.** The crawler enumerates objects under the prefix, applying any exclude patterns you configured (useful for `_temporary/`, `_SUCCESS` and other writer debris). 2. **Classify.** Each file is offered to a chain of **classifiers**. Built-in classifiers recognise Parquet, ORC, Avro, JSON, CSV, XML and several log formats; custom classifiers (grok, CSV with a declared delimiter, JSON path, XML) run first if you attach them. The winning classifier decides the format and yields a schema. 3. **Infer and reconcile a schema.** For self-describing formats such as Parquet or Avro the schema is read from the footer or header. For JSON and CSV the crawler infers types from sampled rows. Schemas from many files under the same folder are reconciled into one. 4. **Group folders into tables.** The crawler decides whether sub-folders are partitions of one table or separate tables, based on whether their schemas and structure are compatible. 5. **Write to the catalog.** It creates or updates the table and adds partition entries. ## What a catalog table actually holds A Data Catalog table is a Hive-metastore-shaped record: - `Name` and the owning `DatabaseName` - `StorageDescriptor.Columns` — the non-partition columns and their types - `PartitionKeys` — the columns derived from folder structure - `StorageDescriptor.Location` — the S3 prefix - `InputFormat` / `OutputFormat` and `SerdeInfo.SerializationLibrary` — how a reader deserialises the bytes - `Parameters` — free-form properties such as `classification=parquet`, `compressionType`, row counts and `CrawlerSchemaSerializerVersion` Each partition is its own record with its own `Location` and storage descriptor, which is why a partition can technically point somewhere outside the table's prefix and can carry a schema that differs from the table's. ## Who reads it The Data Catalog is a per-account, per-Region metastore that speaks the Hive metastore model, so several engines consume it without any further schema definition: - **Amazon Athena** uses it as its metastore outright — a crawled table appears in the Athena query editor. - **Amazon EMR** can be configured to use Glue as its Hive metastore instead of a local one. - **Amazon Redshift Spectrum** builds an external schema from a Glue database. - **AWS Glue ETL jobs** read tables from it via `GlueContext`. That sharing is the point: one crawl, and every engine agrees on what the files mean. ## What it does not do - It does not create data, compact files, or change formats. - It does not run continuously — it runs on demand, on a schedule, or triggered by S3 events if you configure event-based crawling. - It is not mandatory. You can create the same table by hand with `CREATE EXTERNAL TABLE` in Athena or with the `CreateTable` API, and for tables your own pipelines write that is often the better choice, because a crawler infers types from a sample and can therefore change its mind. ## Practical notes The crawler's IAM role needs read access to the S3 objects **and** Glue permissions to create and update tables and partitions; a KMS-encrypted bucket also requires `kms:Decrypt`. Crawlers are billed for the time they run, so pointing one at a prefix containing millions of small files and scheduling it every five minutes is a common way to be surprised by the bill. Start with a manual run, inspect the resulting table with `aws glue get-table`, and only then schedule it.
- Do you have to use a crawler at all to get a table into the Glue Data Catalog?No. The catalog is just a metastore, so you can create the table with `CREATE EXTERNAL TABLE` in Athena, with the Glue `CreateTable` API, or from infrastructure-as-code. For tables your own pipelines produce, declaring the schema explicitly is usually safer than letting a crawler infer it, because inference from a file sample can silently change a column's type between runs.
- What does a classifier do that a SerDe does not?A classifier runs at crawl time and answers "what format is this file, and what schema does it imply?" — its output is the table definition. A SerDe runs at query time inside the engine and answers "how do I turn these bytes into rows?" The crawler picks the SerDe and records it in the table; after that the classifier is out of the picture.
- Why would a Glue crawler need kms:Decrypt in its IAM role?Because it reads the objects themselves to infer a schema. If the S3 objects are encrypted with SSE-KMS using a customer managed key, the crawler role must be allowed `kms:Decrypt` on that key — and the key policy must allow the role — or the crawl fails or produces an empty table.
A crawler is a librarian walking the stacks: it does not move a single book, it fills in the card catalogue so that anyone who comes to the desk can find and read them.
saying these in an interview costs you the question
- Saying the crawler copies or converts the data into Glue
- Thinking rows are stored inside the Data Catalog table
- Assuming a crawler is required before Athena can query S3
- Believing the crawler runs continuously and keeps the table live
- Confusing a crawler with a Glue ETL job that transforms data