How would you configure an Amazon Data Firehose delivery stream so JSON events land in S3 as partitioned Parquet that Athena can query efficiently, and what constrains that setup?
answer
- columnar format needs a declared schema
- Glue table supplies it
- prefix built from record content
- jq inline or Lambda partitionKeys
- one buffer per active partition
basics
~20 sEnable record format conversion, which reads the schema from an AWS Glue Data Catalog table and writes Parquet, and enable dynamic partitioning to build the S3 prefix from fields in each record. Both must be planned at stream creation, and partition keys must be low cardinality.
solid answer
~60 sTwo Firehose features combine here. **Record format conversion** turns incoming JSON into Parquet or ORC; it needs a table in the AWS Glue Data Catalog to supply the schema, and the input must be JSON — anything else has to be converted by a Lambda transformation first. **Dynamic partitioning** derives the S3 prefix from the record's own content, either by inline parsing with a jq expression or from `partitionKeys` returned by the transformation Lambda, and you reference the result in the prefix as `!{partitionKeyFromQuery:name}` or `!{partitionKeyFromLambda:name}`, usually alongside `!{timestamp:yyyy/MM/dd}`. The constraints matter more than the syntax: dynamic partitioning can only be enabled when the delivery stream is created, not added later; an `ErrorOutputPrefix` is required; each active partition buffers separately, so there is a higher minimum buffer size and a quota on active partitions; and both features carry extra per-GB charges. Choose low-cardinality keys — tenant, region, event type — because a key like user id explodes the partition count and produces exactly the small files Parquet was meant to avoid.
code
json · 17 lines{
"Prefix": "events/country=!{partitionKeyFromQuery:country}/!{timestamp:yyyy/MM/dd}/",
"ErrorOutputPrefix": "errors/!{firehose:error-output-type}/!{timestamp:yyyy/MM/dd}/",
"DynamicPartitioningConfiguration": { "Enabled": true },
"ProcessingConfiguration": {
"Enabled": true,
"Processors": [
{
"Type": "MetadataExtraction",
"Parameters": [
{ "ParameterName": "MetadataExtractionQuery", "ParameterValue": "{country:.country}" },
{ "ParameterName": "JsonParsingEngine", "ParameterValue": "JQ-1.6" }
]
}
]
}
}go deeper
Know that Firehose can write Parquet instead of raw JSON and can build the S3 path from fields in the record, and that partitioned columnar data is cheaper for Athena to query.
Explain the mechanics: the Glue Data Catalog table that supplies the schema, jq inline extraction versus Lambda-supplied partition keys, and the !{partitionKeyFromQuery:...} prefix syntax.
Demonstrate production judgment on cardinality and file size — why each active partition buffers separately, what the partition quota implies, and that dynamic partitioning is a creation-time decision you cannot retrofit.
Own the lake layout as a platform contract: which partition keys every stream standardises on, when the per-GB conversion and partitioning charges beat running a downstream ETL job, and how schema evolution is governed in the Glue catalog.
## Why this is a real interview question "Get the events into S3" is easy. "Get them into S3 in a shape the analytics team can query without setting money on fire" is the actual job, and Firehose can do most of it as configuration. The answer has three parts: the file format, the partition layout, and the constraints that stop you doing it naively. ## Part one: record format conversion JSON is a terrible storage format for analytics. It is row-oriented, untyped and uncompressed by column, so a query that reads one field still pays to scan every byte. **Parquet** is columnar: Athena reads only the columns the query names, and per-column encoding compresses far better. Firehose's **record format conversion** does the conversion in the pipeline. Two requirements come with it: 1. **A schema from the AWS Glue Data Catalog.** Parquet is typed, and Firehose will not guess. You create a Glue table describing the record shape, and point the delivery stream at it. That table is usually the same one Athena queries, which is convenient — but it also means schema evolution is a deliberate act: adding a field to your events without updating the Glue table means the field is not written. 2. **JSON input.** Conversion reads JSON. If producers emit something else, a Lambda transformation must normalise it to JSON first — which is one of the main reasons to attach a transformation at all. Conversion applies to the S3 destination and is billed per GB converted on top of ingest. ## Part two: dynamic partitioning A partitioned S3 layout lets Athena skip data it does not need. Without partitioning, `WHERE country = 'DE'` scans everything. With `s3://bucket/events/country=DE/2026/08/21/`, it scans one branch. By default Firehose partitions only by ingest time. **Dynamic partitioning** lets the prefix depend on the record's *content*. There are two ways to extract the key: - **Inline parsing** with a jq expression, configured as a `MetadataExtractionQuery` on a `MetadataExtraction` processor — for example `{country:.country}` — with the parsing engine set to `JQ-1.6`. - **From the transformation Lambda**, which can return a `partitionKeys` object alongside each record when the key needs logic jq cannot express. You then reference the extracted keys in the S3 prefix: ```text events/country=!{partitionKeyFromQuery:country}/!{timestamp:yyyy/MM/dd}/ ``` Use `!{partitionKeyFromLambda:name}` for the Lambda-supplied variant. The prefix must end with `/`, and the `ErrorOutputPrefix` must be set and must include `!{firehose:error-output-type}` so failures land somewhere distinguishable. Writing the key in Hive style (`country=DE`) is what lets Glue crawlers and Athena partition projection recognise the layout without extra configuration. ## Part three: the constraints that bite **Dynamic partitioning is creation-time only.** You cannot enable it on an existing delivery stream. Discovering this after six months of unpartitioned data means creating a new stream and cutting producers over — which is exactly why it is worth deciding up front. **Every active partition buffers separately.** This is the single most important operational fact. With one partition, one buffer flush produces one object. With 200 active partitions, one flush window produces up to 200 objects, each a fraction of the size. Firehose therefore enforces a **larger minimum buffer size** when dynamic partitioning is enabled, and imposes a **quota on active partitions per delivery stream** (a few hundred by default, raisable by request). Exceeding it means records that cannot be routed go to the error output. **Cardinality is a design decision, not a detail.** Good keys are low-cardinality and appear in query predicates: tenant, region, environment, event type. Bad keys are unbounded: user id, session id, request id. A high-cardinality key does not just cost money — it inverts the benefit, producing thousands of tiny Parquet files whose per-file overhead is worse than the scan you were trying to avoid. **Both features cost extra per GB**, and dynamic partitioning also charges for the objects delivered. For a low-value, rarely-queried feed, plain compressed JSON in a time-based prefix may be the better economics. ## Putting it together A sane production shape looks like: producers emit JSON; a Lambda transformation normalises and appends a newline; inline jq extraction pulls one or two low-cardinality keys; the prefix combines those keys Hive-style with a `!{timestamp:yyyy/MM/dd}` element; record format conversion writes Parquet against a Glue table; buffering is tuned toward larger objects; and an alarm watches the error output prefix. Athena then reads a partitioned, columnar table with no ETL job in between — which is the whole point of choosing Firehose over building a consumer. ## Where it goes wrong Partitioning on a high-cardinality field; forgetting the Glue table and wondering why conversion fails; assuming a Glue crawler is needed for Hive-style prefixes when partition projection would be cheaper; and treating the small-file problem as a query-engine issue rather than as the direct consequence of buffer and partition choices made here.
- Why is user id a poor dynamic-partitioning key?Because each active partition buffers independently. A high-cardinality key produces one tiny object per user per flush, multiplying small files, driving up request and delivery cost, and pushing the stream against its active-partition quota. Partition on fields that appear in query predicates and have bounded cardinality — tenant, region, event type.
- What does record format conversion need that plain JSON delivery does not?A table in the AWS Glue Data Catalog to supply the typed schema, and JSON input. Parquet is typed and columnar, so Firehose will not infer the shape. It also means schema changes are deliberate: a new event field is not written until the Glue table describes it.
- You inherited an unpartitioned delivery stream. Can you add dynamic partitioning to it?No — dynamic partitioning can only be enabled when the delivery stream is created. You create a new delivery stream with the partitioning and prefix you want, cut producers over, and if the history matters, reprocess the existing objects with a separate batch job. Plan the layout before the first stream exists.
- Where does the small-file problem actually come from in this setup?From the interaction of buffering and partition count: each active partition holds its own buffer, so objects per flush scale with partitions while bytes per object shrink. The levers are fewer, lower-cardinality partitions and larger buffer sizes; compaction downstream is the remedy once the damage is done, not a substitute for getting this right.
saying these in an interview costs you the question
- Partitions on a unique id like user or request id
- Expects Firehose to infer the Parquet schema itself
- Thinks dynamic partitioning can be turned on later
- Ignores that each partition buffers and flushes separately
- Assumes Parquet conversion is free beyond ingest cost