How do push-based and pull-based metadata ingestion into a data catalog differ, and how does each one go stale?
answer
- who initiates the transfer
- scheduled crawl versus emitted event
- stale between crawls
- missed events drift silently
- many teams combine both
basics
~20 sPull ingestion has the catalog crawl sources on a schedule, so it is stale between crawls and misses short-lived objects. Push ingestion has systems emit metadata when things change, so it is fresh but depends on every producer emitting correctly.
solid answer
~50 sIn **pull** ingestion the catalog runs connectors that **crawl** each source on a schedule — read the system tables, list files, fetch dashboard definitions. It is easy to add a source and needs nothing from producers, but the catalog is **only as fresh as the last crawl**, frequent crawls load the source, and objects created and dropped between crawls are never seen. In **push** ingestion, pipelines and systems **emit metadata events** when something changes — a schema change, a finished run, a new table — so the catalog updates within moments and captures operational facts like runs. The cost is that every producer must integrate and keep emitting; a missed or failed event leaves the catalog **silently wrong** with no crawl to correct it. Many platforms push from pipelines and still pull periodically as a reconciliation pass.
go deeper
Know that pull means the catalog crawls sources and push means sources send metadata when things change.
Explain the freshness and failure trade-offs of each and why deletions are hard for pull.
Design a combined approach with reconciliation crawls and monitoring of ingestion freshness per source.
Decide which producers must integrate push as a platform standard and how to fund and enforce it across teams.
## Two directions of flow A catalog is only useful if its metadata matches reality. Getting metadata **into** it happens in one of two directions. - **Pull**: the catalog initiates. Connectors connect to each source on a schedule and read what is there. - **Push**: the source initiates. Systems send metadata to the catalog (through an API or an event stream) when something happens. ## Comparison | Aspect | Pull (crawl) | Push (emit) | |---|---|---| | Who does the work | the catalog's connectors | each producing system or pipeline | | Freshness | as of the last crawl | near real time | | Onboarding a new source | configure a connector | integrate the producer | | Load on sources | periodic scans, heavier when frequent | small events per change | | Short-lived objects | missed if created and dropped between crawls | captured when emitted | | Operational facts (runs, row counts) | hard to reconstruct later | natural — emitted as they happen | | Failure mode | stale until the next crawl | silently missing if an emitter fails | ## How each one goes stale **Pull** goes stale **predictably**: between crawls the catalog lags reality, and a crawl that fails or times out extends the lag. Deletions are only noticed when a crawl finds an object gone, so dropped tables can linger. **Push** goes stale **unpredictably**: if a pipeline stops emitting (a library upgrade, a new pipeline built without the integration, an event lost in transit), nothing tells the catalog that it has missed something. The catalog looks current while describing a system that has moved on. ## Combining them Mature platforms commonly: 1. **Push** from pipelines and orchestrators, so runs, lineage and schema changes arrive immediately. 2. **Pull** on a slower schedule as a **reconciliation**: the crawl detects objects the push path missed and deletions nobody announced. 3. **Monitor ingestion itself**: alert when a source's last metadata update is older than expected, the same way data freshness is monitored. ## Choosing for a source - A database or warehouse with good system tables → pull is simple and reliable. - A pipeline or streaming system where runs matter → push captures what a crawl cannot. - A dashboarding tool with an export API → usually pull. ## Why interviewers ask it Catalog projects fail when the catalog drifts from reality and users stop trusting it. The strong answer explains the **freshness trade-off**, names the **distinct failure mode** of each direction, and proposes monitoring and reconciliation rather than assuming either is complete.
- How would you detect that a push integration has silently stopped?Track the time of the last metadata event per source and alert when it exceeds the source's normal interval, and compare periodic crawl results with pushed state to spot objects the push path never reported.
- Why can dropped tables linger in a pull-based catalog?A crawl only sees what exists now. Unless the connector compares with the previous crawl and marks missing objects as removed, a dropped table stays listed, and users may keep trying to use it.
Pull is buying the newspaper each morning; push is a news alert sent the moment something happens. Alerts are fresher, but you never hear about the one that failed to send.
saying these in an interview costs you the question
- Assuming a pushed catalog is always current because it is event-driven
- Crawling production sources so often that the scans hurt their performance
- Never removing objects that no longer exist
- Having no alert when metadata for a source stops arriving