How would you decide whether scheduled AWS Glue crawlers or the writing pipelines maintain a lake's Data Catalog?
answer
- ask who knows the schema first
- two zones, two different answers
- the lag is the whole problem
- metadata belongs to the job that wrote the data
- sometimes the best partition record is none
basics
~20 sSplit by ownership. Crawlers are discovery tools for data you did not produce; for tables your own pipelines write, register partitions at write time and declare the schema as code, so metadata is never stale, inferred or surprising.
solid answer
~50 sThe deciding question is whether you already know the schema. For **third-party or exploratory drops** you do not, so a crawler earns its keep: it discovers structure, formats and partitions you did not design. For **tables your pipelines produce**, you know the schema before a byte lands, and a crawler then adds three liabilities — a scheduling lag between the write and the partition appearing, per-run cost that scales with object count, and inference that can retype a column under downstream consumers. There, declare the table in code and have the writing job register the partition (`BatchCreatePartition`, an `ADD PARTITION` statement, or a Glue job configured to update the catalog), or skip partition metadata altogether with Athena partition projection where the layout is predictable. Keep a crawler on those tables only in `LOG` mode, as a drift detector. Then layer the operational concerns on top: partition indexes for very large tables, bounded recrawl scope for cost, and a single owner for governance grants.
code
python · 13 lines# the job that wrote the partition registers it, so there is no crawl lag
import boto3
glue = boto3.client("glue")
table = glue.get_table(DatabaseName="analytics", Name="events")["Table"]
sd = dict(table["StorageDescriptor"])
sd["Location"] = "s3://lake/events/dt=2026-08-20/"
glue.batch_create_partition(
DatabaseName="analytics",
TableName="events",
PartitionInputList=[{"Values": ["2026-08-20"], "StorageDescriptor": sd}],
)go deeper
Know that a crawler is not the only way to get a table into the catalog, and that a pipeline can register its own partitions.
Compare the mechanisms concretely — crawl lag and cost versus write-time registration versus projection — and explain what each does to query planning.
Design the write path so metadata lands with the data, keep crawlers as drift detectors on owned tables, and handle large partition counts deliberately.
Set the lake-wide policy by zone, tie metadata freshness to the consumer contract, and make schema, partition registration and access grants all have named owners rather than emerging from whichever crawler someone created.
## Frame it as ownership, not tooling The wrong version of this decision is "crawlers versus DDL" as a taste question. The right version is: **who is authoritative for the schema of this table?** Every table in the lake has an answer, and the answer determines the mechanism. **You are not authoritative** — a partner drops files, a SaaS export lands, an analyst dumps something into a sandbox. You cannot declare a schema you do not know, so inference is the only option and a crawler is exactly the right tool. Accept `UPDATE_IN_DATABASE`, because a stale definition is worse than a changing one when there is no contract to break. **You are authoritative** — your job writes the files, and marts, dashboards and other teams read the result. Now schema is a published contract. Letting a scheduled inference process rewrite it is a governance hole, and the fix is to move metadata into the same change control as the code: the table definition lives in version control, is applied by the deploy pipeline, and changes through review. ## What a crawler costs on tables you own - **Freshness lag.** Data is invisible until the next crawl. Every "the dashboard is missing this morning's data" incident traces back to this. - **Cost that scales with objects, not with news.** A crawl of a prefix with millions of small files is expensive and mostly re-derives what it already knew. Bounding the scope with "crawl new sub-folders only" or event-based crawling helps, and at that point you have half-built partition registration anyway. - **Inference risk.** One malformed CSV row retypes a column and breaks queries. - **No relationship to your deploy.** Nothing ties the catalog state to the code version that produced the data, so you cannot reason about them together or roll them back together. ## The alternatives, and when each fits **Register from the writer.** The job that wrote `dt=2026-08-20` calls `BatchCreatePartition`, or issues `ALTER TABLE … ADD PARTITION … LOCATION …`, or runs as a Glue ETL job configured to update the catalog as part of the write. Metadata and data become atomic-ish and there is no lag. This is the default for owned tables. **Partition projection.** For predictable, date-shaped layouts, describe the partition space in table properties and let Athena compute prefixes. No partition records exist, so none can go missing, and query planning stops depending on catalog lookups entirely. The constraints are real: the projection must match the layout, values outside the declared range are unreachable, and the properties are honoured by Athena rather than by every engine reading the same catalog table — check your consumer set first. **Schema as code.** `CREATE EXTERNAL TABLE` or `CreateTable` from your IaC, reviewed like any other change, with a migration when the shape changes. ## The operational layer Whatever mechanism you choose, three things bite at scale. **Partition count.** Query planning has to resolve partitions, and a table with hundreds of thousands of them makes that slow. **Partition indexes** on the leading partition keys let filtered lookups use an index rather than a full enumeration, which is the standard remedy when you cannot re-partition. The better remedy is usually to coarsen the grain: partition by day, not by minute, and let file size carry parallelism. **Cost control.** Crawl scope, schedule frequency and exclude patterns are the levers. A crawler that runs every five minutes over the whole lake is a line item somebody will eventually ask about. **Governance ownership.** If Lake Formation governs the catalog, someone must own grants, and tag-based grants are what keep that tractable past a few dozen tables. Decide this at the same time as the registration mechanism, because a table created by a pipeline still has to be granted to its consumers, and "the crawler created it and nobody can read it" is a recurring incident. ## The answer to give A policy, per zone, with reasons: - **Raw / landing (foreign data):** crawlers, scheduled, `DEPRECATE_IN_DATABASE`, scope-bounded, treated as discovery. - **Curated / published (owned data):** schema declared in code, partitions registered by the writing job or replaced by projection, crawler present only in `LOG` mode as a drift alarm, partition indexes where partition counts are large, grants owned by the data-platform team. - **Sandbox:** crawlers, cheap schedule, nobody downstream, lifecycle expiry on the bucket. And one closing principle worth saying out loud: **metadata freshness is a contract with consumers, not a background chore.** If a table promises hourly data, the mechanism that makes an hour visible has to be part of the pipeline that produces it — not a separate schedule that happens to run nearby.
- Where does partition projection stop being the right answer?When the layout is not derivable — irregular keys, backfilled gaps, partition values you cannot express as a range or enum — or when partitions must live outside the templated location. It is also Athena-side: other engines reading the same catalog table do not honour the projection properties, so a mixed consumer set pushes you back to real partition records.
- What does a partition index actually improve, and when is it not the fix?It speeds up filtered partition lookups on tables with very large partition counts, so query planning stops enumerating everything. It does not reduce bytes scanned or help a query that filters on a non-indexed leading key. If planning is slow because the table is partitioned far too finely, coarsening the grain fixes the cause rather than the symptom.
- You still want a crawler on an owned table. What is it for?As an alarm, not an actuator. Run it on a modest schedule with the update behaviour set to LOG, so it reports that inferred schema has diverged from the declared one without touching the table, and alert on that finding. It becomes an early warning that an upstream writer changed shape, while the contract stays under change control.
- How do you keep crawler spend from growing quietly?Bound scope and frequency: crawl new sub-folders only or use event-driven crawling instead of full re-scans, exclude staging and temp prefixes, and drop the schedule to what freshness actually requires. Then audit which crawlers still exist — most lakes accumulate crawlers for tables that have long since moved to pipeline registration.
A crawler is a surveyor: indispensable for land you have just been handed, absurd for a building whose blueprints you drew yesterday.
saying these in an interview costs you the question
- Treating crawlers as the default mechanism for every table
- Ignoring the freshness lag a scheduled crawl introduces
- Letting inference define a schema other teams depend on
- Partitioning finely and then reaching for indexes instead of fixing the grain
- Choosing the mechanism without deciding who owns access grants