skip to content

R

Statistical computing in R: vectors and data frames, vectorized operations, the apply family, factors, and the package ecosystem for analysis and visualization. Interviewers ask in data and research roles, usually to check that you think in vectors and data frames instead of writing element-by-element loops.

on this pageshow

questions

6

In R, how do NA, NULL and NaN differ, and why does mean(c(4, NA, 10)) return NA?

level: juniorimportance: must knowfreq 54%

answer

  1. missing value versus no value
  2. NA spreads through arithmetic
  3. one of them is a number
  4. na.rm and is.na()
  5. comparisons with NA are NA

basics

~20 s

In R, NA is a missing value of any type, NaN is an undefined numeric result like 0/0, and NULL is the absence of an object. NA propagates through arithmetic, so mean() returns NA unless na.rm = TRUE.

solid answer

~40 s

`NA` means **a value exists but is unknown**. It has a place in the vector, has length 1, and has typed variants such as `NA_real_` and `NA_character_`. `NaN` is a special **double** meaning "not a number", produced by things like `0 / 0`. `NULL` is **no object at all**: `length(NULL)` is 0 and `c(1, NULL, 3)` is just `c(1, 3)`. Any arithmetic involving an unknown gives an unknown, so `mean(c(4, NA, 10))` is `NA`. `mean(c(4, NA, 10), na.rm = TRUE)` gives `7`. Test with `is.na()`, never `x == NA`, which is itself `NA`. Note that `is.na(NaN)` is `TRUE`, but `is.nan(NA)` is `FALSE`. A classic filter trap: `x[x > 5]` keeps an `NA` element for every missing value.

code

r · 18 lines
r
x <- c(4, NA, 10)
mean(x)
#> [1] NA
mean(x, na.rm = TRUE)
#> [1] 7
x[x > 5]
#> [1] NA 10
x[which(x > 5)]
#> [1] 10

is.na(NaN); is.nan(NA)
#> [1] TRUE
#> [1] FALSE
length(NULL); length(NA)
#> [1] 0
#> [1] 1
c(1, NULL, 3)
#> [1] 1 3

go deeper

for a junior

Recall the three meanings: NA missing, NaN undefined number, NULL no object. Know that mean() returns NA unless na.rm = TRUE, and that you test with is.na().

for a middle

Explain propagation rules, including NA & FALSE, and why x[x > 5] keeps an NA slot. Know which() and filter() as fixes and the empty-vector NaN edge case.

for a senior

Show how you audit missing values before summarising, count them with sum(is.na()), and catch filters that silently add NA rows to a report.

for a principal

Set team conventions for missing values: when na.rm is allowed, how missing counts are reported beside totals, and how read-time NA handling is specified.

## Three different kinds of "nothing" | | `NA` | `NaN` | `NULL` | |---|---|---|---| | Meaning | a value that exists but is **missing or unknown** | an **undefined number**, e.g. `0 / 0` | **no object**, an empty placeholder | | Type | any atomic type (`NA` is logical; `NA_integer_`, `NA_real_`, `NA_character_`) | double only | its own type, `"NULL"` | | `length()` | 1 | 1 | 0 | | Inside `c()` | occupies a position | occupies a position | disappears: `c(1, NULL, 3)` is `c(1, 3)` | | Test | `is.na(x)` | `is.nan(x)` | `is.null(x)` | Two asymmetries are worth remembering: - `is.na(NaN)` is `TRUE`, because R treats NaN as a kind of missing value. - `is.nan(NA)` is `FALSE`, because a plain `NA` is not a number at all. `NULL` usually appears as "nothing here": a list element that does not exist, a function that returns nothing useful, or an optional argument left unset. ## NA propagates R's rule for `NA` is that an unknown input gives an unknown output: ```r x <- c(4, NA, 10) x + 1 # 5 NA 11 sum(x) # NA mean(x) # NA mean(x, na.rm = TRUE) # 7 x == NA # NA NA NA is.na(x) # FALSE TRUE FALSE ``` `x == NA` returns `NA` for every element, because comparing anything with an unknown gives an unknown. That is why `is.na()` exists. Many summary functions accept **`na.rm = TRUE`**: `sum()`, `mean()`, `min()`, `max()`, `median()`, `sd()` and others. It drops the `NA` values before computing. Use it deliberately, not by reflex. Dropping missing values changes the question being answered, and `sum(is.na(x))` tells you how many you are dropping. Logical operators can still give a definite answer when the unknown does not matter: `NA & FALSE` is `FALSE`, and `NA | TRUE` is `TRUE`. ## The filtering trap Logical indexing with an `NA` in the condition returns an `NA` element: ```r x <- c(4, NA, 10) x > 5 # FALSE NA TRUE x[x > 5] # NA 10 x[which(x > 5)] # 10 ``` `[` cannot decide whether the missing element passes, so it returns `NA` in that slot. On a data frame, `df[df$revenue > 100, ]` produces a row of `NA`s for every row where `revenue` is missing. Two common fixes: 1. Wrap the condition in `which()`, which returns only the positions that are `TRUE`. 2. Use `subset()` or dplyr's `filter()`, both of which drop rows where the condition is `NA`. ## Edge cases that surprise people - `mean(c(NA_real_, NA_real_), na.rm = TRUE)` is `NaN`: after removing the `NA`s nothing is left, and the mean of an empty vector is 0/0. - `length(NA)` is 1, but `length(NULL)` is 0. Code that checks `length(x) == 0` to mean "missing" will miss an `NA`. - Assigning `NULL` to a list element **removes** it: `lst$a <- NULL` deletes `a`. Assigning `NA` keeps a slot holding a missing value. - Reading a CSV turns empty fields into `NA` in numeric columns. In character columns an empty field may arrive as `""` unless you pass `na.strings` to `read.csv()`. ## What a strong answer adds Say what the missing values mean before choosing `na.rm`. Count them with `sum(is.na(x))`, filter with `which()` or `filter()`, and never compare to `NA` with `==`. How to handle missing data statistically, whether to impute, drop or model it, is a separate question from R's semantics.

  • In R, what does mean(c(NA_real_, NA_real_), na.rm = TRUE) return, and why?
    It returns `NaN`. With `na.rm = TRUE` both values are removed, leaving an empty numeric vector, and the mean of nothing is 0 divided by 0, which is `NaN`. Code that later checks only `is.na()` will still catch it, because `is.na(NaN)` is `TRUE`.
  • In R, why does df[df$revenue > 100, ] return rows full of NA?
    Where `revenue` is `NA`, the condition is `NA`, and `[` returns a row of `NA`s for that position rather than dropping it. Use `df[which(df$revenue > 100), ]`, `subset(df, revenue > 100)` or dplyr's `filter()`, which all drop rows where the condition is `NA`.

saying these in an interview costs you the question

  • In R, x == NA is the way to find missing values.
  • NA and NULL are two names for the same missing value.
  • na.rm = TRUE should be added everywhere by default.
  • is.nan() also catches every NA.
  • R silently drops NA rows when you filter with [.
open as a page

In R, what does c(1, 2, 3, 4) + c(10, 20) return, and why is vectorised code preferred over an element-by-element loop?

level: juniorimportance: must knowfreq 56%

basics

~20 s

In R, it returns 11 22 13 24: the shorter vector is recycled to the longer one's length. Vectorised arithmetic runs its loop in compiled code, while growing a result inside an R loop copies it every iteration.

open as a page

In R, what do lapply(), sapply() and vapply() each return, and when is vapply() the safer choice?

level: middleimportance: should knowfreq 38%

basics

~20 s

In R, lapply() always returns a list; sapply() simplifies to a vector or matrix when results line up, else a list; vapply() requires a declared type and length and errors otherwise, making it predictable in functions.

open as a page

In R, what do df["revenue"], df[["revenue"]], df$revenue and df[, "revenue"] each return for a data frame df?

level: middleimportance: should knowfreq 46%

basics

~10 s

In R, df["revenue"] returns a one-column data frame; df[["revenue"]] and df$revenue return the column vector; df[, "revenue"] also returns a vector because drop = TRUE, unless drop = FALSE is given.

open as a page

In R, why does as.numeric() on a factor such as factor(c("10", "20", "5")) return 1 2 3, and how do you convert it correctly?

level: middleimportance: should knowfreq 40%

basics

~10 s

In R, a factor stores integer codes pointing into its levels, so as.numeric() returns the codes. Convert with as.numeric(as.character(f)) or as.numeric(levels(f))[f]. Remove unused levels with droplevels().

open as a page

In R, revenue totals by region differ between aggregate(revenue ~ region, ...) and dplyr's group_by() |> summarise(); how do you diagnose which is right?

level: seniorimportance: should knowfreq 36%

basics

~20 s

In R, aggregate()'s formula method drops rows with NA by default, while dplyr's sum() gives NA for any group with a missing value and keeps an NA region as a group. Count the NAs, then choose explicitly.

open as a page