Compare DelimitedLineTokenizer and FixedLengthTokenizer. How do you configure each, and what do field names give you?
answer
- Delimited = split on char + quote support
- FixedLength = Range(start,end), 1-based inclusive
- setNames → FieldSet by name → drives BeanWrapper
- strict=true throws IncorrectTokenCountException
- delimiter is literal, not regex
basics
~10 sDelimitedLineTokenizer splits a line on a delimiter like a comma. FixedLengthTokenizer splits by fixed column positions (Ranges). Both produce a FieldSet; setting names lets you read fields by name instead of index.
solid answer
~40 sBoth implement LineTokenizer and return a FieldSet, but they split differently. DelimitedLineTokenizer separates tokens by a delimiter (default comma; configurable via setDelimiter) and supports a quote character so delimiters inside quotes are ignored. FixedLengthTokenizer ignores delimiters and instead uses column positions defined by an array of Range objects (setColumns), e.g. Range(1,10), Range(11,20) — 1-based, inclusive. On both you call setNames(...) so the resulting FieldSet is addressable by name (fieldSet.readString("lastName")) rather than only index; those same names are what BeanWrapperFieldSetMapper matches to bean properties. By default tokenizers are strict — the wrong number of columns throws; setStrict(false) tolerates trailing missing columns. DelimitedLineTokenizer also supports includedFields to keep only selected columns.
code
java · 14 lines// Delimited
DelimitedLineTokenizer delim = new DelimitedLineTokenizer();
delim.setDelimiter("|");
delim.setQuoteCharacter('"');
delim.setNames("firstName", "lastName", "age");
// Fixed length: 1-based, inclusive ranges
FixedLengthTokenizer fixed = new FixedLengthTokenizer();
fixed.setColumns(new Range(1, 10), new Range(11, 20), new Range(21, 23));
fixed.setNames("firstName", "lastName", "age");
fixed.setStrict(true);
FieldSet fs = fixed.tokenize("John Smith 042");
int age = fs.readInt("age"); // 42go deeper
Knows delimited splits on a comma and fixed-length splits by position.
Configures both, knows Range semantics, quote char, and that names drive the FieldSet.
Discusses strict mode, includedFields, and the delimiter-is-not-regex and 1-based-Range gotchas.
Weighs fixed-width legacy formats, quoted-CSV limitations (embedded newlines) and when a custom tokenizer is warranted.
## Two tokenizers, one FieldSet contract A `LineTokenizer` turns one raw line into a `FieldSet`. Spring Batch ships two production implementations. ### DelimitedLineTokenizer Splits on a **delimiter character**. - `setDelimiter(String)` — default is a comma (`DelimitedLineTokenizer.DELIMITER_COMMA`). Can be a tab, pipe, etc. Note it is a single delimiter string, not a regex. - `setQuoteCharacter(char)` — default `"`. Tokens wrapped in the quote char may contain the delimiter without being split, and the quotes are stripped. This is what makes real CSV (`"Smith, John",42`) parse correctly. - `setNames(String...)` — assigns names to columns so the `FieldSet` supports name-based access. - `setIncludedFields(int...)` — keep only certain column indexes (drop columns you don't care about). - `setStrict(boolean)` — default `true`; a line whose token count differs from the number of names throws `IncorrectTokenCountException`. `false` pads/tolerates. ### FixedLengthTokenizer Splits by **column position**, not by any delimiter — used for legacy fixed-width/mainframe files where each field occupies a fixed number of characters. - `setColumns(Range...)` — the heart of it. Each `Range` is **1-based and inclusive**: `new Range(1, 10)` is characters 1–10. A `Range` with only a start (`new Range(46)`) runs to end of line. - `setNames(String...)` — one name per column range. - `setStrict(boolean)` — whether a line shorter than the last column is an error. ### The FieldSet result Both return a `FieldSet`. Without names you read by index (`readString(0)`); with names you read by name (`readString("age")` / `readInt("age")`). **Names are load-bearing**: `BeanWrapperFieldSetMapper` matches FieldSet names to JavaBean property setters, so `names("firstName")` must equal the property `firstName`. ### Gotchas - Delimiter is a literal string, **not** a regex — you can't say `\s+`. - Fixed-length `Range` is 1-based inclusive; off-by-one column definitions are the classic bug. - Strict mode: a trailing empty field in delimited files can change the token count and trip `IncorrectTokenCountException`. - Embedded newlines inside quoted CSV fields are **not** handled by the line-based reader out of the box (the reader splits on physical lines first). ### When to use which - Modern exports (CSV/TSV, pipe-delimited): `DelimitedLineTokenizer`. - COBOL/mainframe/legacy fixed-width extracts: `FixedLengthTokenizer`.
- Are FixedLengthTokenizer Range positions 0-based or 1-based?1-based and inclusive: new Range(1, 10) covers characters 1 through 10 of the line.
- How does DelimitedLineTokenizer handle a comma that is part of a value, like "Smith, John"?If the value is wrapped in the configured quote character (default double-quote), the delimiter inside the quotes is ignored and the quotes are stripped from the token.
saying these in an interview costs you the question
- Saying the delimiter accepts a regex
- Treating FixedLengthTokenizer Ranges as 0-based
- Claiming you must always read fields by index — names enable name-based access and bean mapping
- Thinking quoted embedded commas are split into separate fields