In R, why does as.numeric() on a factor such as factor(c("10", "20", "5")) return 1 2 3, and how do you convert it correctly?
answer
- integer codes behind labels
- levels sort as text
- convert labels, not codes
- unused levels linger
- stringsAsFactors changed in 4.0
basics
~10 sIn R, a factor stores integer codes pointing into its levels, so as.numeric() returns the codes. Convert with as.numeric(as.character(f)) or as.numeric(levels(f))[f]. Remove unused levels with droplevels().
solid answer
~40 sA **factor** is R's type for categorical data: an integer vector of **codes** plus a `levels` attribute holding the labels. `factor(c("10", "20", "5"))` sorts its levels as text, giving `"10" "20" "5"`, and stores the codes `1 2 3`. `as.numeric(f)` returns those **codes**, not the labels, so the numbers are silently wrong. Convert through the labels: `as.numeric(as.character(f))`, or the more efficient `as.numeric(levels(f))[f]`. Two related traps: subsetting keeps **unused levels**, so `table()` still shows them with zero counts until you call `droplevels()`, and `factor(x, levels = ...)` turns any value not listed into `NA`. Since **R 4.0.0**, `data.frame()` and `read.csv()` default to `stringsAsFactors = FALSE`, so text columns are no longer factors unless you ask.
code
r · 20 linessize <- factor(c("10", "20", "5", "10"))
levels(size)
#> [1] "10" "20" "5"
as.numeric(size)
#> [1] 1 2 3 1
as.numeric(as.character(size))
#> [1] 10 20 5 10
as.numeric(levels(size))[size]
#> [1] 10 20 5 10
survey <- data.frame(answer = factor(c("yes", "no", "maybe", "yes")))
kept <- survey[survey$answer != "maybe", , drop = FALSE]
table(kept$answer)
#>
#> maybe no yes
#> 0 1 2
table(droplevels(kept$answer))
#>
#> no yes
#> 1 2go deeper
Recall that a factor stores integer codes plus labels, so as.numeric() returns codes. Know the as.numeric(as.character(f)) fix.
Explain text sorting of levels, why indexing by a factor uses its codes, droplevels() for unused levels, and the silent NA for values outside levels.
Trace a wrong numeric summary back to a column that became a factor at import, and show the checks that stop it: level counts, NA counts, explicit column types.
Decide conventions for categorical data in shared analysis code: when factors are created, who fixes level order, and how saved objects from older R versions are handled.
## What a factor is A **factor** represents a categorical variable: region, survey answer, product tier. Internally it is: - an **integer vector of codes**, one per element; - a **`levels`** attribute: the character vector of distinct labels; - the class `"factor"` (or `c("ordered", "factor")` for an ordered factor). ```r answer <- factor(c("yes", "no", "maybe", "yes")) levels(answer) # "maybe" "no" "yes" as.integer(answer) # 3 2 1 3 ``` By default `factor()` sorts the levels. For character input that is **alphabetical text order**, which is not numeric order. ## The as.numeric trap ```r size <- factor(c("10", "20", "5", "10")) levels(size) # "10" "20" "5" as.numeric(size) # 1 2 3 1 as.numeric(as.character(size)) # 10 20 5 10 as.numeric(levels(size))[size] # 10 20 5 10 ``` Step by step: 1. As text, `"10" < "20" < "5"`, because text compares character by character and `"1" < "2" < "5"`. 2. The codes are therefore `10 -> 1`, `20 -> 2`, `5 -> 3`. 3. `as.numeric(size)` returns the codes `1 2 3 1`. There is no error and no warning, just wrong numbers. The fixes both go through the **labels**: | Expression | How it works | Note | |---|---|---| | `as.numeric(as.character(f))` | turns every element into its label, then parses | easiest to read | | `as.numeric(levels(f))[f]` | parses each level once, then indexes by the codes | faster on long vectors with few levels | Indexing with a factor, as in `[f]`, uses its integer codes. That is why the second form works. This bug usually starts earlier. A numeric column became a factor because one entry was not a clean number, such as `"n/a"` or `"1,200"`. That was common when text columns defaulted to factors. ## Levels outlive the data Levels are an attribute. Subsetting the data does not change them: ```r survey <- data.frame(answer = factor(c("yes", "no", "maybe", "yes"))) kept <- survey[survey$answer != "maybe", , drop = FALSE] levels(kept$answer) # "maybe" "no" "yes" table(kept$answer) # maybe 0, no 1, yes 2 kept$answer <- droplevels(kept$answer) levels(kept$answer) # "no" "yes" ``` Unused levels show up as zero-count rows in `table()`, as empty groups in some summaries, and as empty categories in plots. `droplevels()` works on a factor or on a whole data frame. ## Controlling levels on purpose - **Order:** `factor(x, levels = c("low", "medium", "high"))` fixes the order used by `table()`, sorting and plots. - **Ordered factors:** `factor(x, levels = ..., ordered = TRUE)` makes comparisons such as `x > "low"` meaningful. - **Values not in `levels`** become `NA`. A misspelled `"Hgh"` silently turns into a missing value, so check `sum(is.na(f))` after building a factor. - **Relabelling:** assigning to `levels(f)` renames labels by position. Renaming in the wrong order relabels your data. ## The R 4.0.0 change Before R 4.0.0, `data.frame()` and `read.csv()` defaulted to `stringsAsFactors = TRUE`, so every text column became a factor. Since R 4.0.0 the default is **`FALSE`**. Modern code meets factors mainly when it creates them on purpose, for modelling, ordering categories or plotting. Older scripts and saved `.rds` files can still carry them. ## Why interviewers ask The question tests whether you know the difference between **what a value looks like** and **what R stores**. A strong answer names the codes-versus-labels split, gives both fixes, and mentions `droplevels()` and the silent `NA` for values outside `levels`.
- In R, what happens to a value that is not listed in factor(x, levels = c("low", "high"))?It becomes `NA`, silently. A value such as `"Low"` or `"medium"` that is not in `levels` has no code to map to. Check `sum(is.na(f))` against `sum(is.na(x))` after building the factor to catch labels that were dropped this way.
- In R, when is as.numeric(levels(f))[f] preferable to as.numeric(as.character(f))?On long vectors with few distinct levels. `as.character(f)` builds a full character vector and parses every element, while `levels(f)` parses each distinct label once and then indexes by the integer codes. Both return the same numbers.
A factor is like a coat-check: you hold numbered tickets, and the rack holds the coats. Asking for the ticket numbers gets you 1, 2, 3, never the coats, and if the rack is ordered by label, ticket 3 can hold the smallest coat.
saying these in an interview costs you the question
- as.numeric() on a factor returns the numbers shown in the labels.
- Factor levels are sorted numerically when the labels look like numbers.
- Filtering rows out of a data frame also removes their factor levels.
- Text columns still become factors by default in read.csv() in R 4.x.
- A value missing from the levels argument causes an error.