Data & Analysis Libraries
The libraries that do the actual data work in code — pandas and NumPy in Python, dplyr and ggplot2 in R, plus the plotting layers on top. Interviewers ask about them because they reveal whether you think in vectorized operations or in loops.
on this pageshowhide
explore
- Pandasempty
- Series and DataFrameempty
- Time Series and I/Oempty
- NumPyempty
- ndarray and Dtypesempty
- Indexing and Slicingempty
- Broadcastingempty
- Linear Algebraempty
- Matplotlibempty
- Seabornempty
- dplyrempty
- ggplot2empty
- Kotlin DataFrameempty
- Working With Data in One Process295 questions
- Tables and Arrays24 questions
- Column Types and What They Cost20 questions
- Whole Columns at a Time27 questions
- Getting At Rows and Columns26 questions
- Absent Values22 questions
- Split, Apply, Combine25 questions
- Putting Two Tables Together25 questions
- Long and Wide20 questions
- Time-Indexed Data23 questions
- Getting Data In and Out22 questions
- Saying It With a Chart22 questions
- When One Machine Is Not Enough18 questions
- Trusting the Result21 questions
questions
295 · 1 sectionTwo columns of 1,000 values each, taken from different tables, are added together — what decides which value pairs with which?
basics
~10 sThe tool decides, not the expression. Designs that carry a per-row identifier pair the two operands by that identifier first; designs that carry none pair strictly by position and refuse operands of different lengths.
A match of two tables on their key columns returns zero rows, though the same codes print on both sides. Why?
basics
~20 sEquality compares stored values, not the rendering. Two codes that print alike can differ by letter case, leading or trailing spaces, an invisible character, a lost leading zero, or one side holding digits as text and the other as numbers.
Matching 1,000 orders to a customer list on customer id returns 940 rows: which match shape was used, and what do the others keep?
basics
~20 sLosing 60 rows means the match kept only rows partnered on both sides. The three alternatives keep the orders table whole, the customer list whole, or both whole, and unpartnered rows then survive with the other side's columns holding the absent-value marker.
Twelve monthly pieces are stacked into one longer table - how does that differ from gluing two tables side by side?
basics
~20 sStacking adds records, so the result is longer and the hazard is two pieces whose column sets do not agree. Gluing side by side adds attributes, so the result is wider and the hazard is rows paired in an order nobody checked.
A match of 1,000 order rows against a customer reference table returns 1,340 rows - how is that possible?
basics
~20 sA key match returns, for each key value, the first side's occurrence count multiplied by the second side's. If a customer identifier appears more than once in the reference table, the orders carrying it are copied and the total grows.