skip to content

In a Solr schema, how does a dynamicField pattern decide which incoming field names it matches, and what breaks a tie?

level: middleimportance: should knowfreq 52%

answer

  1. the wildcard may only sit at one end
  2. an explicit declaration outranks any pattern
  3. two patterns match — length decides
  4. longest matching pattern wins, order is irrelevant

basics

~20 s

A dynamicField name is a glob with a wildcard only at the start or the end, such as i or attr. An explicitly declared field always wins over any pattern, and among competing patterns the longest one wins.

solid answer

~50 s

A `dynamicField` declares a type for field names you have not enumerated. Its `name` must be a simple glob with a single `*` at the beginning or the end — `*_i`, `*_txt`, `attr_*` — never in the middle. When a document arrives with a field name Solr has never seen, resolution runs in a fixed order: an explicit `<field>` declaration with that exact name always wins; otherwise Solr picks the matching `dynamicField` with the **longest** pattern, so `*_str_s` beats `*_s` for `color_str_s`; if nothing matches, the update is rejected as an unknown field. Many configsets end with a catch-all `<dynamicField name="*" type="ignored"/>` that silently swallows the rest, which is convenient for prototypes and dangerous in production because typos then disappear instead of erroring. The suffix convention (`_i` int, `_s` string, `_txt` text, `_ss` multi-valued string) is just naming discipline layered on this mechanism.

code

xml · 12 lines
xml
<field name="id" type="string" indexed="true" stored="true" required="true"/>
<field name="text" type="text_general" indexed="true" stored="false" multiValued="true"/>

<!-- explicit declaration wins over the *_s pattern below -->
<field name="sku_s" type="string" indexed="true" stored="true" docValues="true"/>

<dynamicField name="*_s"     type="string"       indexed="true" stored="true"/>
<dynamicField name="*_str_s" type="string"       indexed="true" stored="true" docValues="true"/>
<dynamicField name="*_txt"   type="text_general" indexed="true" stored="true"/>
<dynamicField name="*"       type="ignored"      multiValued="true"/>

<copyField source="*_txt" dest="text"/>

go deeper

for a junior

Recall that a dynamicField maps a name pattern such as *_i to a type, so producers can add fields by naming convention without a schema change. Knowing the common suffix conventions is enough here.

for a middle

Explain the two resolution rules — explicit fields outrank patterns, and the longest matching pattern wins — plus the restriction that the wildcard sits only at one end. Be able to describe what copyField does and when it runs.

for a senior

Show that you have operated one: a catch-all pattern mapped to ignored hides producer bugs, unbounded field names bloat segment metadata and facet heap, and any copyField or type change means a reindex. Have a position on dynamic fields versus schemaless guessing.

for a principal

Decide whether the field namespace is a contract you enforce or a free-for-all you absorb. Suffix conventions, a bounded set of patterns and a rejection path for unknown names are cheaper than discovering a hundred thousand fields in production.

## The problem dynamic fields solve Some document shapes are not knowable up front: user-defined attributes, per-tenant metadata, sensor names, the long tail of a product catalogue's specs. Declaring every field is impossible. A `dynamicField` lets the schema say "any field whose name ends in `_i` is an integer", so the application controls the type by choosing a field name. ```xml <dynamicField name="*_i" type="pint" indexed="true" stored="true"/> <dynamicField name="*_s" type="string" indexed="true" stored="true" docValues="true"/> <dynamicField name="*_ss" type="string" indexed="true" stored="true" multiValued="true"/> <dynamicField name="*_txt" type="text_general" indexed="true" stored="true"/> ``` Indexing `{"id":"1","price_i":42,"desc_txt":"blue widget"}` needs no schema change at all. ## Pattern syntax The `name` of a dynamicField is not a regular expression. It supports exactly one `*`, and only as the first or last character. `*_i` and `attr_*` are legal; `pre*post` is not. Everything else about the declaration is identical to a normal `<field>`: `type`, `indexed`, `stored`, `docValues`, `multiValued`, `required`, `default`. ## Resolution order When Solr resolves the field name in an incoming document: 1. An **explicit** `<field name="price_i">` declaration always wins, even if a dynamic pattern also matches. This is how you carve out one special case from a broad pattern. 2. Otherwise, every matching `dynamicField` is considered and the one with the **longest pattern** is chosen. For the field `color_str_s`, the pattern `*_str_s` beats `*_s`. Declaration order in the file does not matter — only pattern length. 3. If nothing matches, the document update fails with an unknown-field error, unless a catch-all pattern exists. That third case is where the ubiquitous `<dynamicField name="*" type="ignored" multiValued="true"/>` comes in: the `ignored` fieldType is declared with `indexed="false" stored="false"`, so unmatched fields are accepted and thrown away. It stops indexing from failing on unexpected input, at the cost of hiding mistakes — a misspelled `titel_txt` is silently discarded rather than rejected. ## copyField and dynamic fields `copyField` sits next to dynamic fields in the same schema and is the other half of the wiring. A copyField directive copies the **raw source value** into another field during indexing, before either field's analyzer runs, so the destination applies its own fieldType and analysis independently. It accepts wildcards: ```xml <copyField source="*_txt" dest="text"/> <copyField source="title" dest="text" maxChars="30000"/> ``` Three properties matter. First, copying happens at index time only — adding or changing a copyField has no effect on documents already in the index, so it implies a reindex. Second, if several sources copy into one destination, or a source is multi-valued, the destination must be `multiValued="true"`. Third, `maxChars` caps how much of a long value is copied, which protects a catch-all field from a single enormous document. The classic use is a catch-all `text` field that everything copies into, searched as a single default field. The modern alternative is to skip the copy and let the edismax `qf` parameter search several fields with per-field boosts — that keeps per-field scoring signals which a catch-all field destroys, at the cost of a wider query. ## Costs and limits Dynamic fields are cheap in the schema but not free in the index. Every distinct field name that actually appears creates its own term dictionary entries, and every field with `docValues` creates its own column. A multi-tenant collection where each tenant invents its own attribute names can accumulate tens of thousands of real fields, inflating segment metadata, slowing merges and bloating heap for faceting. If field names are effectively unbounded, consider a single multi-valued field holding `name=value` tokens instead, or per-tenant collections. Schemaless field guessing is a separate mechanism that adds **explicit** field declarations to the managed schema on the fly; dynamic fields never mutate the schema at all. Teams often use dynamic fields precisely so they can leave schemaless guessing switched off. ## What interviewers probe The two facts that separate a user from someone who has debugged a schema are "explicit beats dynamic" and "longest pattern wins". The third is that a catch-all `*` mapped to `ignored` turns a loud failure into a silent one.

  • What exactly does a copyField directive copy, and when?
    It copies the raw source value at index time, before either field's analyzer runs, so the destination applies its own fieldType and analysis. It never runs at query time and never touches documents already indexed, so adding or changing a copyField requires a reindex. If several sources feed one destination, or the source is multi-valued, the destination must be declared `multiValued="true"`; `maxChars` caps how much of a long value is copied.
  • Why do some teams drop the catch-all copyField into a single text field and use edismax qf instead?
    A catch-all field flattens everything into one bag of terms, so you lose per-field length normalization and cannot boost a title match above a body match. Searching the fields directly with `qf=title^5 body^1` keeps those signals and lets you tune each field's weight. The trade-off is a wider, more expensive query and more careful field-list maintenance.
  • What goes wrong when a multi-tenant collection lets every tenant invent its own dynamic field names?
    Each distinct name that actually appears becomes a real field in the index, with its own term dictionary entries and, if docValues is on, its own column. Tens of thousands of fields inflate segment metadata, slow merges and raise heap use for faceting. The mitigations are a single multi-valued field holding name=value tokens, or splitting tenants into separate collections.

Dynamic fields are like a filing cabinet whose drawers are labelled by suffix rather than by document: anything ending in _i goes in the integer drawer. A named folder for one specific document always overrides the drawer rule.

saying these in an interview costs you the question

  • Thinks dynamicField patterns support full regular expressions
  • Says the first matching pattern in the file wins
  • Believes a dynamic field wins over an explicit declaration
  • Assumes adding a copyField reindexes existing documents
  • Confuses dynamic fields with schemaless field guessing

context