Statistical Computing in R
Learn R through vectors, lists and data frames, then build probability calculations, descriptive summaries, tests, regression, resampling and time-series tools. Assemble an analysis pipeline with explicit data policies and honest statistical interpretation.
11 projects, 275 hands-on levels, run in your browser.
Syllabus
- Foundations: code through statistics: Begin with R functions, values, variables, choices, loops and vectors. Small numerical tasks lead to a composed peak-and-category summary, preparing you for the later statistical projects.
- R Foundations: Build fluency with R vectors: elementwise arithmetic, recycling, positional and named indexing, missing-value policies, character values and factors. Learn how output type and shape affect later analysis.
- Vectors, Lists & the apply Family: Use lists and the apply family to express iteration with deliberate return types. Practice Reduce, Filter, Map, matrices and grouped summaries, while distinguishing concise interfaces from performance guarantees.
- Data Frames from Scratch: Work with aligned data-frame columns through selection, filtering, derived values, sorting, aggregation and joins. Compose a transaction-value pipeline with explicit row and key behavior.
- Probability & Distributions: Compute discrete and continuous probabilities, moments and standardized values, then connect formulas with R distribution functions. Seeded simulations illustrate sampling variation, large-sample limits and their assumptions.
- Descriptive Statistics & EDA: Describe center, spread, quantiles, asymmetry and paired relationships. Compare robust summaries and outlier rules while keeping estimator conventions and diagnostic limitations explicit.
- Hypothesis Testing: Build test statistics, p-values and confidence intervals for means and counts, then compare with R test procedures. Examine test assumptions, Type I/II errors, power and multiplicity without equating significance with practical importance.
- Linear Regression from Scratch: Derive intercept least-squares fits and matrix formulations, measure residual and in-sample fit quantities, and evaluate predictions. Introduce logistic probabilities and a log-likelihood gradient-ascent update, with identifiability and diagnostic limits.
- Resampling & Simulation: Construct bootstrap and permutation experiments with explicit sampling rules, estimate held-out error through cross-validation, and connect leave-one-out resampling with the jackknife. Distinguish Monte Carlo noise from original-data uncertainty and respect dependence assumptions.
- Time Series & Smoothing: Study equally spaced time-ordered data through smoothing, lag correlations, differences and a simple additive decomposition. Fit an intercept AR relationship and compare forecast baselines, preserving chronology and avoiding universal stationarity claims.
- Capstone: A Statistical Analysis Engine: Assemble taught cleaning, exploratory, regression, inference and reporting stages into an inspectable analysis result. Drop missing targets, impute remaining predictor gaps and cap modeling values under an explicit teaching policy. Keep original observed outcomes for the mean test and bootstrap interval; report training fit without claiming held-out validation or causation.
Key concepts
- aggregate: aggregate(value ~ group, data=df, FUN=mean) summarizes observed groups. The formula interface normally omits rows missing either the value or grouping variable…
- ANOVA: Analysis of variance: tests whether several group means differ, using the F-statistic (between-group variance over within-group). R's aov fits it from a fo…
- apply over margins: apply(m, 1, f) applies f to each ROW, apply(m, 2, f) to each COLUMN. Margin 1 is rows, margin 2 is columns.
- Atomic types: R atomic vector types include logical, integer, double, complex, character and raw. An atomic vector has one storage type; ordinary combinations of logical, nu…
- Autocorrelation: Autocorrelation describes linear association with lagged values. White noise has zero population correlation at nonzero lags, while lag zero is one for positiv…
- Autoregressive model: An intercept AR(1) relationship is x[t]=c+phi x[t-1]+error. Covariance-over-past-variance estimates the intercept-model slope; recover c from lag-pair means an…
- Bootstrap: An ordinary bootstrap draws same-size samples of observed positions with replacement and recomputes a statistic. Replicate spread and quantiles can estimate un…
- CDF: The cumulative distribution function P(X <= x) , the running total of the PMF or the area under the PDF to the left of x. pnorm , pbinom , pexp compute it.
- Center: Three measures of the typical value: the mean (balance point), the median (middle value, robust to outliers), and the mode (most frequent). They diverge under…
- Central Limit Theorem: For IID observations with finite positive variance, the centered sample mean scaled by sigma/sqrt(n) converges in distribution to a standard normal. Finite-sam…
- Chi-square test: A test for COUNTS: it compares observed cell counts to those expected under a null, summing (O - E)^2 / E . Used for goodness-of-fit and for independence in a…
- Coercion: Automatic type conversion. Logicals become numbers (TRUE is 1, FALSE is 0), so sum(x > 3) counts how many elements exceed 3. Combining types promotes to the…
- Complete cases: Rows or values with no missing data. x[!is.na(x)] keeps them; imputation (filling NA with a value like the mean) is the alternative to dropping.
- Confidence interval: A 95% confidence procedure covers the fixed target parameter in 95% of repeated samples under its assumptions. Inverting a matching family of tests gives test/…
- Contingency table: A cross-tabulation of two categorical variables. chisq.test on it tests whether the variables are independent; Cramer's V scales the association to [0, 1].
- Correlation: Pearson correlation is covariance divided by the product of positive standard deviations and measures linear association. Spearman correlation applies Pearson…
- Covariance: Measures whether two variables rise and fall together (positive) or oppose (negative). cov(x, y) . Its scale depends on the units, which is why correlation nor…
- Cross-validation: Fit each fold complement and evaluate its held-out rows, including all learned preprocessing within the training part. Pooled error averages observations, whic…
- Data frame: R's table: a LIST of equal-length columns with a class attribute, so is.list(df) is TRUE. Columns can differ in type; df$col or df[['col']] reads o…
- Decomposition: Splitting a time series into trend, seasonal, and remainder components. Once separated, each piece can be modeled or removed; the parts sum back to the origina…
- Degrees of freedom: The number of independent pieces of information left after estimating parameters, e.g. n-1 for a sample variance. It sets the shape of the t and chi-square dis…
- Design matrix: The matrix X of predictors used in regression, with a leading column of 1s for the intercept. cbind(1, x) builds the simple-regression version.
- Differencing: diff(x) forms consecutive changes. An exact linear trajectory has constant first differences; an exact quadratic has constant second differences on equally spa…
- Distribution: The rule assigning probabilities to outcomes. R names them with the d/p/q/r prefixes: dnorm (density), pnorm (cumulative), qnorm (quantile), rnorm (random draw…
- Expectation: The probability-weighted average of outcomes, sum(values * probs) , the distribution's center of mass. For a binomial it is n*p .
- F-statistic: In ordinary one-way ANOVA, F divides between-group mean square by within-group mean square. Its null calibration requires appropriate independent, equal-varian…
- Factor: A categorical vector stores integer codes and an explicit level set. Default levels are sorted under the current collation, but custom order and unused levels…
- Forecasting: Predicting future values of a series. Baselines include the naive forecast (carry the last value) and the drift method (extend the average trend); accuracy is…
- Formula: R's y ~ x syntax describing a relationship: response on the left, predictors on the right. lm , aov , glm , and aggregate all read formulas.
- Functionals: Functions that take functions: Reduce folds a vector to one value, Filter keeps matching elements, Map applies in parallel, do.call calls a function with a lis…
- Gradient descent: Gradient descent subtracts learning-rate times a loss gradient. Gradient ascent adds learning-rate times a score/log-likelihood gradient. The logistic lesson m…
- Hat matrix & leverage: For a full-rank linear design, H=X(X-transpose X)^-1 X-transpose maps responses to fitted values. Diagonal leverages reflect predictor geometry; trace(H) equal…
- Hypothesis test: A hypothesis test compares a statistic with its distribution under a specified null and other assumptions. A calibrated rejection rule controls Type I error; t…
- Imputation: Imputation fills missing entries using a declared model or rule. Mean filling preserves row alignment but can distort variation and relationships; filled value…
- IQR: The interquartile range Q3 - Q1 , the spread of the middle half of the data. Robust to outliers, and the basis of the boxplot and the 1.5-IQR outlier fence.
- Jackknife: A resampling method that leaves out one observation at a time to estimate a statistic's bias and variance. The simpler ancestor of the bootstrap.
- lapply / sapply: lapply(x, f) applies f to each element and returns a LIST; sapply is the same but simplifies the result to a vector or matrix when it can.
- Law of Large Numbers: Under suitable conditions, such as IID observations with finite expectation, sample averages converge to the population mean. This is a limiting result, not a…
- Least squares: Least squares minimizes the sum of squared residuals over the specified model. The inverse normal-equation formula requires full column rank; QR or SVD methods…
- Levels: levels(f) returns the declared categories of a factor, including unused levels. This count can differ from the number of categories actually observed.
- Linear regression: Fitting a line (or plane) by LEAST SQUARES, minimizing the sum of squared residuals. lm(y ~ x) fits it; the slope is cov(x,y)/var(x) , the intercept passes thr…
- List: R's heterogeneous container: it can hold values of different types, and other lists, at once. [[ extracts a single element (the bare value); [ returns a su…
- lm: R's linear-model function: lm(y ~ x, data = df) . coef extracts coefficients, fitted / resid the fitted values and residuals, predict makes new predictions…
- Logical indexing: Subsetting with TRUE/FALSE values keeps matching positions. NA in a logical index can create missing output entries; use an explicit missingness policy such as…
- Logistic regression: Regression for a yes/no outcome: it models the probability through the sigmoid 1/(1+exp(-z)) , fit by maximizing the log-likelihood (R's glm(y ~ x, family…
- Matrix: A vector with two dimensions, filled COLUMN by column. Indexed m[i, j] (1-based), with fast whole-matrix summaries rowSums / colMeans and margin-wise apply(m,…
- merge (join): merge(a, b, by = 'id') joins two data frames on a shared key (an inner join by default; all.x = TRUE makes it a left join). The relational join in base…
- Monte Carlo: Monte Carlo estimates quantities by random simulation. For independent finite-variance contributions, standard error typically scales as 1/sqrt(n). Realized er…
- Moving average & smoothing: Averaging a sliding window to see through noise. Exponential smoothing weights recent points more: s[t] = alpha*x[t] + (1-alpha)*s[t-1] .
- Multiple comparisons: For m independent true-null tests each rejecting with probability alpha, familywise error is 1-(1-alpha)^m. Bonferroni testing each at alpha/m bounds familywis…
- Mutate: Adding or transforming a column by assignment: df$total <- df$price * df$qty . The verb for derived variables.
- NA (missing value): NA marks missing information; is.na detects NA and NaN. Many operations propagate it, but there are exceptions such as NA^0. Do not compare with == NA to detec…
- na.rm: Many summaries accept na.rm=TRUE to omit missing entries. The remaining sample still needs enough observations: mean(numeric(0)) is NaN and sd of one value is…
- Named vector: A vector whose elements carry names, looked up with v['key'] (a sub-vector) or v[['key']] (the bare value). The lightweight key-value store of…
- Normal distribution: A normal distribution is specified by its mean and positive sd. Approximately 68%, 95% and 99.7% lie within one, two and three sd. The standard-normal 97.5th p…
- Normal equations: For full-column-rank X, least-squares coefficients satisfy (X-transpose X) beta = X-transpose y. Explicit inversion is a teaching formula; lm uses QR-based met…
- Null hypothesis: A null specifies a parameter restriction or probability model, often but not always no effect. Rejection is conditional on the test assumptions; failure to rej…
- One-based indexing: R indexes from 1, not 0: x[1] is the first element. A negative index DROPS elements ( x[-1] is everything but the first), unlike most languages.
- Outlier: A value far from the rest. Tukey's rule flags points beyond Q1 - 1.5*IQR or Q3 + 1.5*IQR ; the z-score rule flags points more than k sd from the mean. Wins…
- Overfitting: Overfitting captures sample-specific patterns that fail to generalize. A train/test error gap can be a symptom, but a positive gap alone is not proof; sampling…
- p-value: The probability, IF the null is true, of a result at least as extreme as the one observed. Small p casts doubt on the null. It is not the probability the null…
- Permutation test: A permutation test rearranges exchangeable labels or observations under a justified null/design. This track uses a seeded raw fraction of simulated statistics…
- PMF and PDF: A probability MASS function gives the probability of each discrete outcome (and sums to 1); a probability DENSITY function describes a continuous distribution…
- Power: Power is the probability of rejecting a specified false null under a specified alternative and design. Larger effects, samples or alpha often increase power wi…
- predict: predict(model,newdata) applies a fitted relationship using compatible input names and preprocessing. Held-out assessment must respect the intended population,…
- Probability: A number in [0, 1] measuring how likely an event is. Independent events multiply; complements subtract from 1. The foundation under every statistical method.
- Quantile: A quantile is a cutoff associated with a probability level. Sample definitions differ; R quantile defaults to type 7 interpolation. The 25th percentile means p…
- R-squared: For a nonconstant response in ordinary intercept regression, training R-squared is 1-SSE/SST and simple-regression R-squared equals squared Pearson correlation…
- Recycling: In ordinary vector arithmetic, a shorter nonempty operand repeats to align with a longer one. A nonmultiple length normally warns, and zero-length arithmetic r…
- Reduce (fold): Reduce(f, x) combines elements left to right into a single value, with an optional init seed that also handles the empty case. Sum, product, and running aggreg…
- Resampling: Estimating uncertainty by re-drawing from the data: WITH replacement (the bootstrap) or by shuffling labels (permutation). When a formula is hard, let the comp…
- Residual: A residual is observed minus fitted response. Ordinary unweighted least-squares residuals sum approximately to zero when an intercept is fitted. That identity…
- RMSE / MAE: RMSE is sqrt(mean(error^2)); MAE is mean(abs(error)). State whether errors are training residuals or held-out predictions. Both have response units; RMSE weigh…
- Robust statistics: A robust statistic limits sensitivity to certain departures or contamination. Median, IQR and MAD resist extreme magnitudes more than mean and sd, but finite-s…
- Select & filter: Use df[cols] or df[, cols, drop=FALSE] to select columns while retaining a data frame. A shared row mask preserves alignment; define how missing mask values ar…
- set.seed: set.seed initializes R random generation. Reproduction requires the same RNG settings and draw sequence, with compatible software behavior; a seed alone is not…
- Sigmoid & logit: The sigmoid 1/(1+exp(-z)) squashes any real number to a probability in (0,1); its inverse, the logit log(p/(1-p)) , is the log-odds. The link between linear pr…
- Significance level (alpha): A significance level alpha specifies a target upper bound on null rejection probability for a valid test. Actual error depends on calibration and assumptions;…
- Simulation: Generating random data from a model to study its behavior. replicate(n, expr) runs an experiment n times; sample , runif , and rnorm provide the randomness.
- Skewness: The third standardized population-style moment used here is mean((x-mean(x))^3)/mean((x-mean(x))^2)^(3/2), requiring positive scale. Mean greater than median i…
- Split-apply-combine: The pattern at the heart of data analysis: split data into groups, apply a summary to each, combine the results. tapply , split , and aggregate do it in one li…
- Standard error: Standard error is the sampling sd of a statistic. For an IID finite-variance sample mean, sample sd divided by sqrt(n) estimates it. Dependence, unequal design…
- Standardize (z-score): For at least two finite observations with positive sample sd, (x-mean(x))/sd(x) gives mean-zero, sample-sd-one values. It changes units without forcing normali…
- Stationarity: Weak stationarity requires constant finite mean and variance and covariance depending only on lag. Strict stationarity concerns shift-invariant joint distribut…
- Subsetting: Extracting parts of a vector by position ( x[c(1,3)] ), by a logical mask ( x[x > 0] ), or by name ( x['a'] ). The most-used operation in R.
- t-test: t procedures address means with estimated variability. One-sample and paired tests use n-1 degrees of freedom for observations or differences; default two-samp…
- table: table(x) tallies how many times each distinct value appears, returning a named count vector; table(a, b) builds a contingency table for two variables.
- tapply: tapply(x, group, f) applies f to each group of x defined by group , returning a named result, the canonical split-apply-combine call.
- The apply family: lapply, sapply, vapply and mapply express iteration over structures. lapply returns a list; sapply may simplify; vapply checks a supplied result template. Thes…
- The d/p/q/r family: R's four functions per distribution: d for the density or mass, p for the cumulative probability (left tail), q for the quantile (inverse of p), r for rand…
- Time series: A time series is ordered by time and may have temporal dependence. Some series are independent noise. Analysis must consider sampling intervals, dependence, ch…
- Type I and II error: A type I error rejects a true null (false positive, rate alpha); a type II error fails to reject a false null (false negative). Power is 1 minus the type II ra…
- unlist: Flattens a list of values into a plain vector, so sum(unlist(l)) totals a list of numbers. The bridge from list-land back to vectors.
- vapply: vapply applies a function with an explicit expected result template, checking length and compatible type. For scalar numeric results, numeric(1) gives predicta…
- Variance & sd: Population variance is expected squared deviation from the mean; sd is its square root. Sample var and sd in R use n-1. Sample variance is unbiased for an IID…
- Vector: R's atomic data structure: an ordered sequence of values of ONE type. Even a single number is a length-1 vector, which is why so much of R is vectorized.
- Vectorization: Applying an operation across a vector through its interface, such as x*2. This often expresses intent concisely and may use optimized internals, but speed depe…
- which: which returns positions that are TRUE, omitting FALSE and NA. which.max and which.min return the first extremum position; they are not membership tests.