In a Solr schema, what is the difference between the solr.StrField and solr.TextField field types?
answer
- one is analyzed, one is not
- think exact match versus word match
- which one is safe to facet on
- StrField indexes the entire value as one term
basics
~20 ssolr.StrField stores the value verbatim with no analysis, so it only matches, sorts and facets on the whole string. solr.TextField runs an analysis chain that tokenizes and normalizes text, enabling full-text matching on individual words.
solid answer
~50 sA `solr.StrField` is not analyzed: the whole value becomes one index term, so a document whose title is `The Matrix` matches only the exact string `The Matrix`. That makes StrField the right type for identifiers, enum-like codes, and any field you facet, sort or group on. A `solr.TextField` runs an analyzer — a tokenizer plus a chain of filters — at index time and again at query time, so `The Matrix` is broken into `the` and `matrix` and a search for `matrix` hits it. TextField is for human-readable text you want to search, and a poor choice for faceting (your buckets become tokens) or sorting (which needs at most one term per document). The usual pattern is to index both: a TextField for search plus a StrField with `docValues="true"` for facets and sorting, wired together with a `copyField`. Solr also offers `solr.SortableTextField` when one field must do both.
code
xml · 16 lines<fieldType name="string" class="solr.StrField" sortMissingLast="true" docValues="true"/>
<fieldType name="text_general" class="solr.TextField" positionIncrementGap="100">
<analyzer type="index">
<tokenizer class="solr.StandardTokenizerFactory"/>
<filter class="solr.LowerCaseFilterFactory"/>
</analyzer>
<analyzer type="query">
<tokenizer class="solr.StandardTokenizerFactory"/>
<filter class="solr.LowerCaseFilterFactory"/>
</analyzer>
</fieldType>
<field name="category" type="string" indexed="true" stored="true" docValues="true"/>
<field name="category_txt" type="text_general" indexed="true" stored="false"/>
<copyField source="category" dest="category_txt"/>go deeper
Be able to say which type is analyzed and give one example field for each: an id or category code as string, a product description as text. Knowing that exact-match failures usually trace to this choice is enough at this level.
Explain the analyzer chain a TextField runs at index and query time, why faceting and sorting want a single term per document, and how docValues fits in. Be ready to sketch the copyField pattern that gives you both behaviours.
Own the consequence: a fieldType change is a reindex, so describe how you migrate — build a new collection, reindex, swap the alias — without downtime. Also know when SortableTextField is the cleaner answer than a copyField pair.
Frame it as schema governance: who may add fields, whether types are guessed or declared, and how a schema change is versioned alongside the reindex it forces. The cost of a wrong type is paid in reindex hours, so the review gate matters more than the individual field.
## What a fieldType is in Solr A Solr schema (the `managed-schema.xml` file, or its equivalent managed through the Schema API) declares `fieldType` definitions and then `field` declarations that reference them. The fieldType decides how a value is turned into index terms, whether those terms are analyzed, and which capabilities (sorting, faceting, range queries) the field supports. `solr.StrField` and `solr.TextField` are the two you meet first, and choosing the wrong one is the most common cause of "my search returns nothing" in a new Solr project. ## solr.StrField — the raw string `solr.StrField` performs no analysis whatsoever. The exact byte sequence you send becomes exactly one term. No lowercasing, no tokenizing, no stemming, no whitespace splitting. Consequences: - A query for `title:matrix` will not match a StrField holding `The Matrix`; only `title:"The Matrix"` (or the equivalent escaped form) matches. - Because there is exactly one term per value, StrField is safe to sort on, facet on, and group on. - It supports `docValues="true"`, the column-oriented structure Solr uses for sorting, faceting and function queries; without docValues those operations must uninvert the index at runtime, which costs heap. - It is the correct type for the `uniqueKey` field, SKUs, ISO country codes, status enums, and URLs used as keys. ## solr.TextField — the analysis chain `solr.TextField` is the only common type that carries analyzers. A fieldType declares up to two of them — `<analyzer type="index">` and `<analyzer type="query">` — each made of an optional set of char filters, exactly one tokenizer, and any number of token filters. A typical `text_general` chain uses `solr.StandardTokenizerFactory`, `solr.StopFilterFactory` and `solr.LowerCaseFilterFactory`. Indexing `The Matrix` therefore produces the term `matrix` (and possibly `the`, depending on the stopword list), and a user typing `MATRIX` is lowercased at query time and matches. This is what makes full-text search work: the analysis chain is a normalizing function applied on both sides so that surface differences in casing, punctuation, word forms and accents stop mattering. It is also why TextField behaves badly for the operations StrField is good at — faceting on an analyzed field gives you buckets of tokens (`the`, `matrix`) rather than the human-readable value, and sorting requires a field that yields at most one term per document. ## Storage and indexing flags are orthogonal Both types accept `indexed`, `stored`, `docValues` and `multiValued`, and beginners often confuse them with the type itself: - `indexed="true"` — searchable. - `stored="true"` — the original value can be returned in results. - `docValues="true"` — a per-document forward column, used for sorting, faceting and function queries. `solr.TextField` does not support docValues (use `solr.SortableTextField` if you need them alongside analysis; it builds docValues from the original, truncated value). - `useDocValuesAsStored` lets docValues stand in for stored values on retrieval. A field can be indexed but not stored (searchable, not returnable) or stored but not indexed (returnable, not searchable). ## Indexing both: the copyField pattern The usual production shape is to keep the authoritative value in a StrField and a searchable projection in a TextField: ```xml <field name="category" type="string" indexed="true" stored="true" docValues="true"/> <field name="category_txt" type="text_general" indexed="true" stored="false"/> <copyField source="category" dest="category_txt"/> ``` Now `facet.field=category` returns clean buckets, and `qf=category_txt` lets edismax match individual words. The same trick builds a catch-all `text` field that many fields copy into. ## The reindex rule Changing a field's type — StrField to TextField or the reverse — changes how terms are produced, and Solr never rewrites documents that are already in the index. Existing documents keep the terms they were indexed with, so behaviour becomes inconsistent between old and new documents until you reindex everything. Treat any fieldType or analyzer change as "requires a full reindex", ideally into a new collection you then alias to. ## Common mistakes Declaring a product description as `string` (StrField) and wondering why keyword search misses; faceting on a `text_general` field and getting tokens; sorting on an analyzed field and getting an error or nonsense ordering; assuming `stored="true"` is what makes a field searchable. Each is a fieldType choice, not a query bug.
- If you need both keyword search and sorting on a product title, how do you declare it?Two options. Either declare the title as a TextField for searching and copy it with `copyField` into a `string` field carrying `docValues="true"` that you sort and facet on, or use `solr.SortableTextField`, which analyzes the value like a TextField while also building docValues from the original value, truncated beyond a configured character limit. The copyField approach is more explicit and lets the two fields use different types.
- Why is a solr.StrField a bad choice for the field a user types into?Because nothing normalizes either side. The user must reproduce casing, punctuation and word order exactly, and partial-word or multi-word matching is impossible without wildcards. StrField is for values your application generates and compares exactly — keys, codes, enums — not for language.
- Does changing a field from string to text_general re-analyze the documents already indexed?No. Solr stores the terms produced at index time; it does not keep a re-analyzable copy of the source. Documents indexed under the old type keep their old terms, so the index becomes a mix of two behaviours. Any fieldType or analyzer change requires reindexing the affected documents, usually into a fresh collection swapped in via an alias.
saying these in an interview costs you the question
- Says StrField is just a shorter TextField
- Facets on an analyzed TextField and expects whole values
- Thinks stored=true is what makes a field searchable
- Believes changing a field's type reindexes existing documents
- Uses TextField for the uniqueKey id field