skip to content

In a bash script you must split each line of a colon-delimited file into fields. Explain how `IFS=: read -ra fields <<< "$line"` works, what scope the `IFS=:` has, and what `read -ra` produces for the line `x::y`.

level: middleimportance: should knowfreq 52%

answer

  1. one command, one separator
  2. assignment before the command name
  3. -a fills an array
  4. a colon is not whitespace
  5. repeats leave holes

basics

~20 s

The IFS=: prefix applies to that single read, leaving the script's IFS untouched. read -ra splits the input on colons and assigns the fields to array elements. Because a colon is not whitespace, x::y yields three fields: x, an empty string, and y.

solid answer

~40 s

`IFS=:` written before the command name is a per-command assignment: `read` runs with a colon as its only field separator, and once it returns the script's `IFS` is exactly what it was. `read` splits the input it consumed at those separators, and `-a fields` assigns the fields to successive elements of the array `fields` instead of to named variables; `-r` keeps backslashes literal, as always. The key subtlety is that IFS splitting treats whitespace and non-whitespace separators differently: runs of whitespace collapse and leading or trailing runs are ignored, but each non-whitespace separator delimits its own field. A colon is non-whitespace, so `x::y` produces three fields — `x`, an empty string, and `y` — and `${#fields[@]}` is 3. That difference is exactly what makes colon- and comma-delimited data usable: empty fields survive.

code

bash · 8 lines
bash
line="x::y"
IFS=: read -ra fields <<< "$line"
echo "${#fields[@]}"            # 3
printf '[%s]\n' "${fields[@]}"  # [x] [] [y]
printf '%q\n' "$IFS"            # $' \t\n' - unchanged

read -ra words <<< "   x   y   "
echo "${#words[@]}"             # 2 - whitespace runs collapse

go deeper

for a junior

Know the shape and use it: IFS=: read -ra fields <<< "$line" splits a line on colons into the array fields, and the IFS=: applies only to that read.

for a middle

Explain the per-command assignment scope, what -a and -r each do, and the whitespace-versus-non-whitespace splitting rule that makes x::y three fields while x y is two.

for a senior

Show judgment about where the mechanism stops: no multi-character delimiters, no quoted-field CSV, and a per-line shell loop being the wrong tool for large inputs compared with a single awk pass.

for a principal

Own the boundary decision — when ad-hoc delimiter splitting in shell is acceptable glue and when the data contract demands a real parser or a different language, and how you stop that judgment from being relitigated in every script review.

## The per-command assignment `IFS=: read -ra fields <<< "$line"` is one command with a variable assignment prefixed to it. In bash, an assignment placed before a command name applies to that command's execution only. Because `read` is a builtin, bash arranges for `IFS` to hold the old value again as soon as it returns: ```bash IFS=, read -r a b <<< "x,y" printf '%q\n' "$IFS" # $' \t\n' - still the default ``` This matters because `IFS` is one of the most far-reaching globals in the shell: it governs how *every* later unquoted expansion is split. A stray `IFS=,` on its own line changes behaviour hundreds of lines away. Writing the assignment as a prefix confines the blast radius to one command, and it is why the idiom is always spelled `IFS=: read` and never `IFS=:` then `read` on the next line. The same trick works for any command, and a subshell is the other containment option when several commands need the changed separator: `( IFS=,; ... )`. ## What `-a` and `-r` do `read` normally assigns to the variable names you list, with the **last** name receiving all remaining fields joined by the original separators: ```bash IFS=: read -r user rest <<< "root:x:0:0" # user=root rest=x:0:0 ``` With `-a NAME`, `read` instead assigns each field to a successive element of the indexed array `NAME`, replacing any previous contents, so you do not have to know the field count in advance: ```bash IFS=: read -ra fields <<< "root:x:0:0" echo "${#fields[@]}" # 4 printf '[%s]\n' "${fields[@]}" ``` `-r` is orthogonal and should always be present for data: it stops `read` from consuming backslashes as escapes. ## Whitespace vs non-whitespace separators This is the rule candidates most often get wrong, and it is the reason the same mechanism behaves differently for a log line and for a CSV row: - **IFS whitespace** (space, tab, newline): a *run* counts as one separator, and leading and trailing runs are discarded. So with the default `IFS`, `" x y "` splits into exactly two fields and empty fields never appear. - **IFS non-whitespace** (`:`, `,`, `|`): each occurrence delimits a field of its own. Repeats therefore produce empty fields, and a leading separator produces an empty first field. ```bash IFS=, read -ra a <<< "x,,y" echo "${#a[@]}" # 3 -> x, empty, y read -ra b <<< " x y " echo "${#b[@]}" # 2 -> x, y ``` Empty fields surviving is a feature: a database export where column three is blank must still line up with column four. Note one asymmetry worth remembering rather than deriving: a *trailing* separator does not add a final empty element in `read`, because the remainder after the last separator is empty and there is nothing left to assign. ## Practical uses and limits Common shapes you will write: ```bash while IFS=: read -r user _ uid _; do [[ $uid -ge 1000 ]] && printf '%s\n' "$user" done < /etc/passwd IFS='|' read -ra cols <<< "$row" ``` The limits are real. This splits on a delimiter; it does not *parse*. Genuine CSV allows quoted fields containing the delimiter, escaped quotes and embedded newlines, none of which IFS splitting understands — reach for a real parser there. And splitting a whole file this way is per-line work in the shell, so for large inputs a single `awk` invocation is usually the better engineering call. A last detail: `IFS` also acts as a *joiner* in a few expansions, so changing it affects more than splitting. Keeping the assignment scoped to one command sidesteps all of that.

  • Why prefix the assignment instead of setting IFS on its own line and restoring it later?
    A prefix is exception-safe and local: `IFS` is restored automatically when the command returns, with no chance of an early `return`, `exit` or error path skipping your restore line. Manual save-and-restore also breaks under `set -e` and in functions that exit mid-way. Where several commands need it, use `local IFS` in a function or a subshell instead.
  • What happens if the line has more fields than the variable names you give read?
    The last named variable receives everything left over, separators included: `IFS=: read -r user rest <<< "root:x:0:0"` leaves `rest` as `x:0:0`. That is convenient for "first field and the tail" splits, but if you want independent columns you need either enough names or `-a` to capture them into an array.
  • How do you split on a multi-character delimiter such as `::`?
    You cannot with IFS: every character in IFS is an independent single-character delimiter, so `IFS='::'` is identical to `IFS=':'`. Either substitute the delimiter for a single unused character first with parameter expansion or `sed`, or use a tool built for it such as `awk -F'::'`.

saying these in an interview costs you the question

  • Thinks IFS=: before read changes IFS for the whole script
  • Expects repeated colons to collapse like spaces do
  • Believes IFS can hold a multi-character delimiter
  • Says read -a appends to the array rather than replacing it
  • Treats IFS splitting as a CSV parser that understands quoted fields

context