skip to content

In R, how do NA, NULL and NaN differ, and why does mean(c(4, NA, 10)) return NA?

level: juniorimportance: must knowfreq 54%

answer

  1. missing value versus no value
  2. NA spreads through arithmetic
  3. one of them is a number
  4. na.rm and is.na()
  5. comparisons with NA are NA

basics

~20 s

In R, NA is a missing value of any type, NaN is an undefined numeric result like 0/0, and NULL is the absence of an object. NA propagates through arithmetic, so mean() returns NA unless na.rm = TRUE.

solid answer

~40 s

`NA` means **a value exists but is unknown**. It has a place in the vector, has length 1, and has typed variants such as `NA_real_` and `NA_character_`. `NaN` is a special **double** meaning "not a number", produced by things like `0 / 0`. `NULL` is **no object at all**: `length(NULL)` is 0 and `c(1, NULL, 3)` is just `c(1, 3)`. Any arithmetic involving an unknown gives an unknown, so `mean(c(4, NA, 10))` is `NA`. `mean(c(4, NA, 10), na.rm = TRUE)` gives `7`. Test with `is.na()`, never `x == NA`, which is itself `NA`. Note that `is.na(NaN)` is `TRUE`, but `is.nan(NA)` is `FALSE`. A classic filter trap: `x[x > 5]` keeps an `NA` element for every missing value.

code

r · 18 lines
r
x <- c(4, NA, 10)
mean(x)
#> [1] NA
mean(x, na.rm = TRUE)
#> [1] 7
x[x > 5]
#> [1] NA 10
x[which(x > 5)]
#> [1] 10

is.na(NaN); is.nan(NA)
#> [1] TRUE
#> [1] FALSE
length(NULL); length(NA)
#> [1] 0
#> [1] 1
c(1, NULL, 3)
#> [1] 1 3

go deeper

for a junior

Recall the three meanings: NA missing, NaN undefined number, NULL no object. Know that mean() returns NA unless na.rm = TRUE, and that you test with is.na().

for a middle

Explain propagation rules, including NA & FALSE, and why x[x > 5] keeps an NA slot. Know which() and filter() as fixes and the empty-vector NaN edge case.

for a senior

Show how you audit missing values before summarising, count them with sum(is.na()), and catch filters that silently add NA rows to a report.

for a principal

Set team conventions for missing values: when na.rm is allowed, how missing counts are reported beside totals, and how read-time NA handling is specified.

## Three different kinds of "nothing" | | `NA` | `NaN` | `NULL` | |---|---|---|---| | Meaning | a value that exists but is **missing or unknown** | an **undefined number**, e.g. `0 / 0` | **no object**, an empty placeholder | | Type | any atomic type (`NA` is logical; `NA_integer_`, `NA_real_`, `NA_character_`) | double only | its own type, `"NULL"` | | `length()` | 1 | 1 | 0 | | Inside `c()` | occupies a position | occupies a position | disappears: `c(1, NULL, 3)` is `c(1, 3)` | | Test | `is.na(x)` | `is.nan(x)` | `is.null(x)` | Two asymmetries are worth remembering: - `is.na(NaN)` is `TRUE`, because R treats NaN as a kind of missing value. - `is.nan(NA)` is `FALSE`, because a plain `NA` is not a number at all. `NULL` usually appears as "nothing here": a list element that does not exist, a function that returns nothing useful, or an optional argument left unset. ## NA propagates R's rule for `NA` is that an unknown input gives an unknown output: ```r x <- c(4, NA, 10) x + 1 # 5 NA 11 sum(x) # NA mean(x) # NA mean(x, na.rm = TRUE) # 7 x == NA # NA NA NA is.na(x) # FALSE TRUE FALSE ``` `x == NA` returns `NA` for every element, because comparing anything with an unknown gives an unknown. That is why `is.na()` exists. Many summary functions accept **`na.rm = TRUE`**: `sum()`, `mean()`, `min()`, `max()`, `median()`, `sd()` and others. It drops the `NA` values before computing. Use it deliberately, not by reflex. Dropping missing values changes the question being answered, and `sum(is.na(x))` tells you how many you are dropping. Logical operators can still give a definite answer when the unknown does not matter: `NA & FALSE` is `FALSE`, and `NA | TRUE` is `TRUE`. ## The filtering trap Logical indexing with an `NA` in the condition returns an `NA` element: ```r x <- c(4, NA, 10) x > 5 # FALSE NA TRUE x[x > 5] # NA 10 x[which(x > 5)] # 10 ``` `[` cannot decide whether the missing element passes, so it returns `NA` in that slot. On a data frame, `df[df$revenue > 100, ]` produces a row of `NA`s for every row where `revenue` is missing. Two common fixes: 1. Wrap the condition in `which()`, which returns only the positions that are `TRUE`. 2. Use `subset()` or dplyr's `filter()`, both of which drop rows where the condition is `NA`. ## Edge cases that surprise people - `mean(c(NA_real_, NA_real_), na.rm = TRUE)` is `NaN`: after removing the `NA`s nothing is left, and the mean of an empty vector is 0/0. - `length(NA)` is 1, but `length(NULL)` is 0. Code that checks `length(x) == 0` to mean "missing" will miss an `NA`. - Assigning `NULL` to a list element **removes** it: `lst$a <- NULL` deletes `a`. Assigning `NA` keeps a slot holding a missing value. - Reading a CSV turns empty fields into `NA` in numeric columns. In character columns an empty field may arrive as `""` unless you pass `na.strings` to `read.csv()`. ## What a strong answer adds Say what the missing values mean before choosing `na.rm`. Count them with `sum(is.na(x))`, filter with `which()` or `filter()`, and never compare to `NA` with `==`. How to handle missing data statistically, whether to impute, drop or model it, is a separate question from R's semantics.

  • In R, what does mean(c(NA_real_, NA_real_), na.rm = TRUE) return, and why?
    It returns `NaN`. With `na.rm = TRUE` both values are removed, leaving an empty numeric vector, and the mean of nothing is 0 divided by 0, which is `NaN`. Code that later checks only `is.na()` will still catch it, because `is.na(NaN)` is `TRUE`.
  • In R, why does df[df$revenue > 100, ] return rows full of NA?
    Where `revenue` is `NA`, the condition is `NA`, and `[` returns a row of `NA`s for that position rather than dropping it. Use `df[which(df$revenue > 100), ]`, `subset(df, revenue > 100)` or dplyr's `filter()`, which all drop rows where the condition is `NA`.

saying these in an interview costs you the question

  • In R, x == NA is the way to find missing values.
  • NA and NULL are two names for the same missing value.
  • na.rm = TRUE should be added everywhere by default.
  • is.nan() also catches every NA.
  • R silently drops NA rows when you filter with [.