skip to content

In R, what do df["revenue"], df[["revenue"]], df$revenue and df[, "revenue"] each return for a data frame df?

level: middleimportance: should knowfreq 46%

answer

  1. a data frame is a list
  2. single bracket keeps the container
  3. double bracket extracts
  4. drop = TRUE bites one column
  5. tibbles never drop

basics

~10 s

In R, df["revenue"] returns a one-column data frame; df[["revenue"]] and df$revenue return the column vector; df[, "revenue"] also returns a vector because drop = TRUE, unless drop = FALSE is given.

solid answer

~40 s

A data frame is a **list of equal-length columns**, so list rules apply. `[` with one argument, `df["revenue"]`, **subsets the list** and returns a data frame with one column. `[[`, as in `df[["revenue"]]`, **extracts** one element, the column vector itself. `df$revenue` does the same as `[[`, but it allows partial name matching, silently by default. The matrix-style `df[, "revenue"]` selects rows and columns, and when the result has **one column** it defaults to `drop = TRUE` and returns a plain vector. `df[, "revenue", drop = FALSE]` keeps a data frame. Code that picks columns with a variable and then calls `nrow()` or `names()` on the result breaks when a single column is chosen. A **tibble**, dplyr's data frame, never drops: `tb[, "revenue"]` stays a tibble.

code

r · 20 lines
r
sales <- data.frame(region = c("North", "South", "West"),
                    revenue = c(120, 80, 150))

class(sales["revenue"])
#> [1] "data.frame"
sales[["revenue"]]
#> [1] 120  80 150
sales[, "revenue"]
#> [1] 120  80 150
sales[, "revenue", drop = FALSE]
#>   revenue
#> 1     120
#> 2      80
#> 3     150

cols <- "revenue"
nrow(sales[, cols])
#> NULL
nrow(sales[, cols, drop = FALSE])
#> [1] 3

go deeper

for a junior

Recall the four forms and what each returns: [ keeps a data frame, [[ and $ give the vector, and the row-column form drops a single column to a vector.

for a middle

Explain the list semantics behind [ and [[, the drop argument and its default, $ partial matching, and why a tibble never drops.

for a senior

Diagnose a helper that fails only when one column is selected, and write indexing that behaves the same for base data frames and tibbles.

for a principal

Set conventions for shared analysis code, such as [[ for vectors, drop = FALSE or select() for frames, and no $ with computed names, so shape bugs cannot recur.

## A data frame is a list A base R **data frame** is a list whose elements are columns (vectors of equal length), with `names` for the columns and `row.names` for the rows. Most indexing behaviour follows from list semantics plus a matrix-like two-argument form. ```r sales <- data.frame( region = c("North", "South", "West"), revenue = c(120, 80, 150) ) ``` ## The four forms side by side | Expression | Returns | Why | |---|---|---| | `sales["revenue"]` | data frame, 3 x 1 | `[` on a list returns a smaller list of the same kind | | `sales[["revenue"]]` | numeric vector `120 80 150` | `[[` extracts one element | | `sales$revenue` | numeric vector `120 80 150` | shorthand for `[[` with a literal name | | `sales[, "revenue"]` | numeric vector `120 80 150` | matrix-style `[` with `drop = TRUE` by default | | `sales[, "revenue", drop = FALSE]` | data frame, 3 x 1 | dropping switched off | | `sales[, c("region", "revenue")]` | data frame, 3 x 2 | several columns never drop to a vector | The single-bracket, single-argument form and the double bracket are the list rules: **`[` keeps the container, `[[` takes the thing out**. The row-column form adds the `drop` argument, which simplifies a one-column result to a vector. ## Where drop = TRUE causes bugs The dangerous case is code where the column list is a variable: ```r top_cols <- c("revenue") subset_df <- sales[, top_cols] # a vector, not a data frame nrow(subset_df) # NULL names(subset_df) # NULL ``` It works while `top_cols` has two or more names and breaks the day someone passes one. Fixes: 1. Write `sales[, top_cols, drop = FALSE]` whenever the result must stay a data frame. 2. Or use the single-argument form, `sales[top_cols]`, which never drops. 3. Or use dplyr: `select(sales, all_of(top_cols))` always returns a data frame. Rows behave differently. `sales[2, ]` returns a one-row data frame, because dropping applies to the columns dimension when a single column is selected. ## $ and partial matching `$` is convenient at the console but has two limits: - It takes a **literal name**. `sales$col_name` looks for a column literally called `col_name`. To use a variable, write `sales[[col_name]]`. - It does **partial matching**. `sales$rev` finds `revenue`. By default R does this silently; it warns only if you set `options(warnPartialMatchDollar = TRUE)`. A typo can therefore read the wrong column. `[[` matches exactly by default. A missing column gives `NULL` with `$` or `[[`, not an error, so check `"revenue" %in% names(sales)` when the name comes from input. ## Tibbles at a glance A **tibble** is the data frame variant dplyr returns and creates with `tibble()` or `as_tibble()`. It keeps the `data.frame` class but is stricter: - `tb[, "revenue"]` **stays a tibble**. There is no dropping, so the bug above cannot happen. - `tb$rev` does **not** partially match. It returns `NULL` with a warning about an unknown column. - It prints only the first rows and shows column types. So the same bracket expression can return a vector for a base data frame and a tibble for a tibble. That matters when a helper function receives "a data frame" from either world. Use `[[` for a vector and `[` with one argument for a data frame, and the code behaves the same for both. ## Rows and conditions - `sales[sales$revenue > 100, ]` selects rows by condition. The trailing comma means "all columns". - `sales[order(-sales$revenue), ]` sorts rows. - Where the condition contains `NA`, base `[` returns `NA` rows. `subset()` or dplyr's `filter()` drop them.

  • In R, how do you pull a column whose name is stored in a variable?
    Use `df[[col]]` for the vector or `df[col]` for a one-column data frame. `df$col` looks for a column literally named `col` and returns `NULL` if there is none. With dplyr, `select(df, all_of(col))` returns a one-column data frame.
  • In R, why can the same helper return different shapes for a data frame and a tibble?
    Because `x[, "revenue"]` drops to a vector on a base data frame but stays a tibble on a tibble. A helper that relies on the result's shape behaves differently depending on its input. Writing `x[["revenue"]]` or `x["revenue"]` gives the same shape for both.

saying these in an interview costs you the question

  • df["col"] and df[["col"]] return the same thing in R.
  • df[, "col"] always returns a data frame.
  • df$name works with a column name stored in a variable.
  • $ on a data frame only matches exact column names.
  • A tibble drops a single selected column to a vector like base R.