In R, what does c(1, 2, 3, 4) + c(10, 20) return, and why is vectorised code preferred over an element-by-element loop?
answer
- everything is a vector
- the shorter one repeats
- warning on uneven lengths
- growing a vector copies it
- the loop runs in C
basics
~20 sIn R, it returns 11 22 13 24: the shorter vector is recycled to the longer one's length. Vectorised arithmetic runs its loop in compiled code, while growing a result inside an R loop copies it every iteration.
solid answer
~50 sR has no scalars: `5` is a numeric vector of length 1, and arithmetic operators work **element-wise** on whole vectors. When lengths differ, R **recycles** the shorter vector. `c(1, 2, 3, 4) + c(10, 20)` becomes `c(1+10, 2+20, 3+10, 4+20)`, which is `11 22 13 24`. If the longer length is not a multiple of the shorter, R still recycles but warns that *longer object length is not a multiple of shorter object length*. Vectorised code such as `x * 2` or `ifelse(x > 100, "hit", "miss")` runs its loop inside compiled C, so it is short and fast. An R `for` loop is not slow in itself if you **preallocate** the result, for example with `numeric(n)`. The real trap is growing a vector with `out <- c(out, value)`, which copies the whole vector on every pass.
code
r · 21 linesc(1, 2, 3, 4) + c(10, 20)
#> [1] 11 22 13 24
c(1, 2, 3) + c(10, 20)
#> [1] 11 22 13
#> Warning: longer object length is not a multiple of shorter object length
x <- runif(1e4)
grown <- c()
for (i in seq_along(x)) grown <- c(grown, x[i] * 2) # copies on every pass
filled <- numeric(length(x))
for (i in seq_along(x)) filled[i] <- x[i] * 2 # preallocated
vectorised <- x * 2
identical(grown, vectorised)
#> [1] TRUE
identical(filled, vectorised)
#> [1] TRUEgo deeper
Recall that R values are vectors, that operators work element-wise, and that the shorter vector is recycled. Work out 1 to 4 plus c(10, 20) by hand.
Explain the three loop shapes: growing, preallocated and vectorised. Say why growing a vector costs quadratic copying, and why a recycling warning should stop an analysis.
Show how you find a quietly wrong result from uneven recycling, and when a preallocated loop is the honest choice because each step depends on the last.
Weigh readability against speed for an analyst team: which vectorised idioms to standardise, and when rewriting a hot loop in compiled code is justified.
## Everything is a vector R's basic data type is the **atomic vector**: an ordered sequence of values of one type (`logical`, `integer`, `double`, `character`, and a few rarer ones). There is no separate scalar type. `42` is a double vector of length 1, and `"North"` is a character vector of length 1. Indexing is **1-based**, so `x[1]` is the first element. Because values are vectors, most operators and many functions are **vectorised**: they apply element by element across the whole vector in one call. ```r sales <- c(120, 80, 150, 95) target <- c(100, 90, 100, 90) sales - target # 20 -10 50 5 sales >= target # TRUE FALSE TRUE TRUE ifelse(sales >= 100, "hit", "miss") # "hit" "miss" "hit" "miss" ``` ## Recycling: when lengths differ If two vectors in an element-wise operation have different lengths, R **recycles** the shorter one, repeating it from the start until it matches the longer one. | Expression | Recycled as | Result | |---|---|---| | `c(1, 2, 3, 4) + c(10, 20)` | `c(10, 20, 10, 20)` | `11 22 13 24` | | `c(1, 2, 3, 4) * 2` | `c(2, 2, 2, 2)` | `2 4 6 8` | | `c(1, 2, 3) + c(10, 20)` | `c(10, 20, 10)` | `11 22 13`, **with a warning** | The length-1 case is how `x * 2` or `x > 100` work at all. The multiple-length case is legitimate and silent. The uneven case still produces a result, and R only **warns**: *longer object length is not a multiple of shorter object length*. A warning scrolls past easily in a long script, so an uneven-length operation is a common source of quietly wrong numbers. Recycling is **positional**. R pairs element 1 with element 1 and so on. It does not look at names or keys to decide what goes together. ## Why vectorised code beats a naive loop Three different things get called "a loop in R", and they perform very differently: 1. **Growing a result:** `out <- c(); for (i in seq_along(x)) out <- c(out, x[i] * 2)`. Every `c()` call allocates a new vector and copies everything collected so far. The total work grows with the **square** of the length. This is the pattern that makes R "slow". 2. **Preallocated loop:** `out <- numeric(length(x)); for (i in seq_along(x)) out[i] <- x[i] * 2`. The result is allocated once and filled in place. Since R 3.4 loops are byte-compiled by default, and this form is often acceptable. 3. **Vectorised call:** `out <- x * 2`. One call, and the element loop runs in compiled C inside R. It is usually the fastest option and always the shortest. The same reasoning applies to the apply family. `sapply(x, function(v) v * 2)` is still an R-level loop that calls your function once per element. It is tidier than a `for` loop, but it is not vectorisation. ## Habits interviewers look for - Use `seq_along(x)` rather than `1:length(x)`. When `x` is empty, `1:length(x)` is `1:0`, which is `c(1, 0)`, and the loop runs twice on nothing. - Preallocate with `numeric(n)`, `character(n)` or `vector("list", n)` when a loop is genuinely needed, for example when each step depends on the previous result. - Reach first for vectorised building blocks: arithmetic, comparisons, `ifelse()`, `cumsum()`, `pmax()`/`pmin()`, `rowSums()`/`colSums()`, `paste()`. - Treat a recycling warning as an error in analysis code: check lengths before combining vectors. ## A worked comparison Suppose a budget line needs a 10% uplift on every month's spend, `spend`, of length 12: - Vectorised: `spend * 1.1`, one expression. - Preallocated loop: four lines, correct, somewhat slower. - Growing loop: correct result, but it copies an ever-larger vector twelve times. For twelve values nobody notices. For a million rows the difference is dramatic.
- In R, why is seq_along(x) safer than 1:length(x) in a for loop?When `x` has length 0, `1:length(x)` evaluates to `1:0`, which is `c(1, 0)`. The loop then runs twice and indexes elements that do not exist. `seq_along(x)` returns an empty integer vector for an empty `x`, so the loop body never runs.
- In R, is sapply() a vectorised alternative to a for loop?Not in the performance sense. `sapply()` calls the R function once per element, so it does the same R-level work as a preallocated `for` loop. It is shorter and handles the result for you. True vectorisation means one call whose element loop runs in compiled code, such as `x * 2` or `cumsum(x)`.
saying these in an interview costs you the question
- Adding vectors of different lengths in R always throws an error.
- R pads the shorter vector with zeros or NA.
- Recycling matches elements by name, like a join.
- Every for loop in R is slow, so all loops must go.
- sapply() is vectorised, so it is as fast as x * 2.