In an AWS Glue job, what does resolveChoice do to a column with a ChoiceType?
answer
- a field that is two types needs a decision
- five ways to settle the argument
- one of them adds columns, one adds a struct
- cast, project, make_cols, make_struct, match_catalog
- settle it before the schema is forced
basics
~20 sresolveChoice turns a DynamicFrame field that holds more than one type into a single well-defined shape. You choose the rule: cast to one type, project to one type, split into a column per type (make_cols), wrap the types in a struct (make_struct), or match the Data Catalog.
solid answer
~40 sA Glue `DynamicFrame` keeps a field that is a long in some records and a string in others as a **ChoiceType**. `resolveChoice` is how you decide what that field should actually be before writing it or converting to a DataFrame. The resolutions are: `cast:<type>` (convert everything, values that will not convert become null), `project:<type>` (keep only the records where the field already has that type, null elsewhere), `make_cols` (split into one column per type, e.g. `amount_long` and `amount_string`), `make_struct` (keep both under a struct with a member per type), and `match_catalog` (cast to whatever the Glue Data Catalog table says, which needs `database` and `table_name`). You apply it either per field with `specs=[("amount", "cast:double")]` or to every choice at once with `choice="make_struct"`. Resolve before `toDF()` or the conversion decides for you.
code
python · 11 linesdyf.printSchema()
# root
# |-- order_id: string
# |-- amount: choice
# | |-- long
# | |-- string
dyf = dyf.resolveChoice(specs=[("amount", "cast:double")])
# alternative: keep both types, one column each
# dyf = dyf.resolveChoice(specs=[("amount", "make_cols")])go deeper
Recall that a Glue field holding two different types is a ChoiceType and that resolveChoice is the transform that settles it — being able to name cast is enough at this level.
Name the resolutions and what each does to the data and the schema, including that make_cols adds type-suffixed columns and match_catalog needs the catalog database and table.
Argue the choice from what the ambiguity means upstream, and show how you detect a cast that is quietly nulling real values instead of finding out downstream.
Decide where type ambiguity is allowed to be settled at all: in the ingest job, in a published catalog contract, or by pushing the fix back to the producing team — and make that rule uniform across pipelines.
## Why a ChoiceType exists at all A Glue `DynamicFrame` is self-describing per record. When Glue reads a JSON or CSV dataset in which the same field is sometimes `1200` and sometimes `"1200"` — or sometimes a scalar and sometimes an array — it does not fail and it does not guess. It records the field as a **ChoiceType**: a field that has genuinely more than one type in this dataset. Printing the schema shows something like: ```text root |-- order_id: string |-- amount: choice | |-- long | |-- string ``` That is an honest description of the input, but it is not a shape you can write to Parquet or hand to Spark SQL. `resolveChoice` is the transform that turns the honest description into a decision. ## The five resolutions - **`cast:<type>`** — convert every value to the named type. Values that cannot be converted become null. This is the common choice when the ambiguity is cosmetic (numbers quoted by one producer) and you know the semantic type. - **`project:<type>`** — keep the field only where it already *is* that type; records with any other type get null. Use it when the other types are garbage you want to discard rather than convert. - **`make_cols`** — split the field into one column per type, named by suffixing the type: `amount_long`, `amount_string`. Nothing is lost and nothing is coerced; downstream SQL decides. Note that this changes the output schema, so a table already registered for this dataset needs updating. - **`make_struct`** — keep the field as a struct with one member per type, so `amount` becomes a struct with a `long` member and a `string` member, one of them populated per record. Also lossless, and keeps the column count stable. - **`match_catalog`** — cast to whatever the Glue Data Catalog says the column is, which is why this resolution needs the `database` and `table_name` arguments. It is the right answer when the catalog is your contract and the job's role is to conform to it. ## Applying it Two forms, and mixing them is an error: ```python # per-field: different rules for different columns dyf = dyf.resolveChoice(specs=[ ("amount", "cast:double"), ("tags", "make_cols"), ]) # blanket: one rule for every ChoiceType in the frame dyf = dyf.resolveChoice(choice="make_struct") # conform to the catalog's declared types dyf = dyf.resolveChoice( choice="match_catalog", database="raw", table_name="orders") ``` The `specs` list uses a dotted path for nested fields, so a choice inside a struct is addressed as `"payload.amount"`. ## Why the ordering matters `toDF()` collapses the DynamicFrame's per-record schemas into one Spark schema, and writing to a strongly typed format like Parquet does the same thing at the sink. Either way, an unresolved ChoiceType gets decided by machinery rather than by a rule you wrote down. That is how a column of amounts quietly loses every string-formatted row, and it will not show up as a job failure — the run goes green and the numbers are wrong. So: resolve first, then convert or write. In review, an unresolved choice reaching `toDF()` is worth flagging the same way you would flag a swallowed exception. ## Choosing the rule The operational question behind `resolveChoice` is *what does the ambiguity mean?* - If it is a formatting inconsistency from one upstream producer and the semantic type is obvious, `cast` is right, but you should count the nulls it creates — a spike in nulls means the cast is eating real values. - If the two types are genuinely different facts sharing a name (a legacy field reused), `make_cols` keeps them separable and forces the downstream model to be explicit. - If you are loading into a table whose schema is a published contract, `match_catalog` puts the contract in charge. - `project` is a filter dressed as a type resolution; be sure you actually want the discarded records gone. ## What interviewers listen for That you know the ChoiceType is a feature rather than an error state; that you can name at least three resolutions and what each does to the data; and that you resolve deliberately before the schema is forced, rather than discovering the coercion later in a mart nobody reconciled.
- What does the relationalize transform do, and how is it different from resolveChoice?`relationalize` restructures nesting rather than types: it flattens nested structs into dotted top-level columns and pulls each array out into a separate table joined back by a generated key, returning a `DynamicFrameCollection` and writing intermediate data to the staging path you supply. `resolveChoice` never changes the shape of the record graph — it only decides what a multi-typed field becomes.
- How would you notice that cast:double is silently discarding real values?Count nulls in the resolved column before and after, and alarm on the ratio rather than eyeballing it. A cast turns anything unconvertible into null without failing, so a producer that starts sending values in a new format shows up only as a rise in nulls — which is why the null rate on cast columns belongs in the job's emitted metrics or a downstream test.
saying these in an interview costs you the question
- Treating a ChoiceType as a read error to be suppressed
- Assuming resolveChoice always casts to string
- Calling toDF() first and resolving the choice afterwards
- Using match_catalog without passing database and table_name
- Not realising make_cols changes the output column set