warming up your workspace

Machine Learning in R

Learn predictive modeling in R through feature preparation, nearest neighbors, naive Bayes, trees, ensembles, clustering, PCA and evaluation. Build teaching-scale algorithms and compose fitted projects, using base R library routines for selected operations and independent comparisons. Finish with validation-based model selection and a reserved test evaluation.

11 projects, 275 hands-on levels, run in your browser.

Syllabus

  • Foundations: code through machine learning: Never written code before? Start here. You will learn the basics of R, functions, variables, types, decisions, loops, and vectors, with simple machine-learning examples. By the end you are ready for Project 1.
  • Features & the Model Matrix: Build a reusable model-matrix preparation workflow: reserve test rows, fit numeric imputation and scaling, encode categories with saved dictionaries, and apply transformations to new rows in the same feature order. Inspect shape, unknown-category and constant-feature policies. Fit learned preparation inside each training fold so validation and final test information stay outside fitting.
  • k-Nearest Neighbors: Build distances, neighbor ranking, deterministic votes and weighted classification/regression in R. Store training data and fitted preprocessing for reusable prediction, then choose neighbor counts using held-out validation. Explore the effects of scaling, duplicate observations and dimension without assuming one distance rule suits every dataset.
  • Naive Bayes: Build Gaussian and multinomial naive Bayes from class priors, fitted feature distributions, fixed vocabularies and add-one smoothing. Use direct log scores to limit underflow and keep class, variance and unknown-token policies explicit. Compose trained continuous-feature and spam-style count classifiers with reusable prediction.
  • Decision Trees: Build a teaching-scale CART-like numeric classification tree from Gini/entropy, valid threshold search, weighted child impurity and recursive growth. Store feature indices and leaf predictions, support depth and minimum-leaf stops, and perform structural cost-complexity pruning with comparable loss units. Inspect fitted predictions and validate complexity choices on held-out data.
  • Ensembles: Compose fitted bagged trees and a feature-subsampled forest with bootstrap identities and eligible out-of-bag predictions. Build weighted binary AdaBoost and feature-dependent squared-loss gradient boosting, then explore prediction blending and out-of-fold stacking. Track seeds, stopping policies, signed permutation changes and coverage; ensemble improvement is evaluated rather than guaranteed.
  • k-Means & Partitional Clustering: Build capped Lloyd assignment/update fitting with explicit empty-cluster policies, stored centers and diagnostics, and multiple starts. Distinguish deterministic farthest-first helpers from randomized k-means++ seeding. Evaluate inertia and silhouette and compare compatible fitted results with R kmeans; initialization and local optima limit universal monotonicity claims.
  • Hierarchical & Density Clustering: Build distance/linkage calculations, an agglomerative merge sequence and cuts, and complete DBSCAN core-led expansion with border/noise policies. Compare compatible hierarchical results with R dist/hclust and inspect pairwise partition agreement. These methods offer different geometries, with scale and parameter assumptions rather than guarantees about every cluster shape.
  • PCA & Dimensionality Reduction: Build PCA from training-fitted centering, optional scaling, sample covariance and eigendecomposition. Store a retained basis for projection and inverse transformation of new rows, inspect reduced reconstruction error and variance retention, and compare aligned results with prcomp. Respect global sign ambiguity, repeated-eigenvalue subspaces, rank deficiency and wide-matrix dimensions.
  • Model Evaluation & Selection: Build aligned confusion counts, defined metric conventions, complete threshold ROC records and pairwise/trapezoidal AUC checks. Run validation with fold-local fitted preparation and models, preserve candidate identities and select from comparable held-out scores. Refit the chosen candidate before reserved test evaluation and report sample size, class counts and limitations.
  • Capstone: A Predictive Modeling Pipeline: Assemble a complete predictive modeling pipeline with training-fitted imputation and scaling, real kNN/Gaussian naive Bayes/Gini-stump candidates, fold-local fitting and validation-based hyperparameter/model selection. Retain per-candidate held-out predictions and scores, refit the winner on all training rows, and evaluate once on reserved test data with counts, metrics and a training-fitted baseline. Expose fitted state and prediction APIs; retain the one-feature report as a compatible helper rather than the whole system.

Key concepts

  • Accuracy: The fraction of aligned nonempty predictions that match truth. Interpret it with class frequencies, error types, sample size and the selected baseline.
  • AdaBoost: Binary boosting that weights learners using their weighted error and reweights observations after each round. Alpha is half the log odds of correctness; zero e…
  • AUC: Area under a complete ROC curve, equal to positive-over-negative ranking probability plus half the tie probability. It measures ranking, not calibration; finit…
  • Bagging: Fitting models on bootstrap resamples and averaging or voting their predictions. Its variance effect depends on member variability and correlation; bias is not…
  • Baseline: A declared reference procedure, such as a training-fitted constant-class predictor. Evaluate it on the same held-out rows; beating one accuracy baseline is nei…
  • Bayes' theorem: Posterior probability is proportional to prior probability times likelihood, with normalization over the possible classes when the total evidence is positive.
  • Bias-variance tradeoff: Under squared-error assumptions, expected error decomposes into squared bias, estimator variance and irreducible noise. This is not a universal decomposition o…
  • Binning: Replacing continuous values with interval categories. Boundaries, endpoint inclusion and out-of-range behavior must be explicit; learned cut points are fitted…
  • Boosting: Sequentially fitting models to improve a combined predictor according to a chosen objective. AdaBoost changes sample weights; gradient boosting fits negative-l…
  • Bootstrap sample: Sampling n observation identities with replacement from n rows. The expected omitted fraction is (1-1/n)^n, approaching about 0.368; an individual sample need…
  • CART: Classification and Regression Trees use greedy recursive binary splitting with terminal predictions. This track implements a teaching-scale numeric classificat…
  • Class imbalance: Unequal class frequencies, sometimes making overall accuracy conceal important minority-class errors. Choose metrics and baselines to match the intended use; a…
  • Classification: Supervised prediction of a discrete class. A declared positive class, error costs and evaluation design determine which performance metrics are informative.
  • Clustering: Grouping observations using a declared similarity or density notion. k-means favors squared-distance compactness, hierarchy uses linkage, and DBSCAN uses densi…
  • Confusion matrix: Counts of actual and predicted class combinations with explicitly declared axes. The binary matrix supports accuracy, precision, recall and F1 at one threshold…
  • Core, border, noise: In the DBSCAN convention here, core means at least min_pts observations within inclusive eps, including self. A non-core point near a core is border; otherwise…
  • Cosine similarity: The dot product divided by the product of nonzero vector lengths. It compares orientation rather than magnitude; 1 minus this similarity is the cosine dissimil…
  • Covariance matrix: A matrix of feature variances on the diagonal and covariances off it. Complete-data sample covariance is symmetric positive semidefinite; symmetry gives real e…
  • Cross-validation: Repeated fitting on training-fold complements and scoring held-out folds. Include learned preprocessing inside each fit. Fold scores depend on split design and…
  • Curse of dimensionality: Challenges that can arise as dimension increases, including sparse coverage and reduced contrast between near and far distances under some data distributions.…
  • Data leakage: Use of information unavailable to the intended training or prediction stage, such as held-out labels, future data, shared entities or all-data preprocessing. T…
  • DBSCAN: Density clustering with inclusive radius eps and minimum neighborhood size including self here. Core points expand connected regions; border points can join wi…
  • Decision threshold: A cutoff converting an oriented score into a label. This track uses score>=threshold for positive. Scores may be uncalibrated; precision need not increase m…
  • Decision tree: A model that routes observations through nested split rules to terminal predictions. Small trees can be easy to inspect; interpretability depends on size and c…
  • Dendrogram: A tree recording hierarchical merges and their dissimilarity heights. Single-linkage heights are nondecreasing, with ties allowed; a chosen count or height def…
  • Distance metric: A distance obeying nonnegativity, identity, symmetry and the triangle inequality. Euclidean and Manhattan distances are metrics; 1 minus cosine similarity gene…
  • Dummy encoding: Indicator encoding with one reference level omitted, typically leaving k-1 columns. With an intercept, this removes the full-indicator linear dependence; absen…
  • Eigenvectors & eigenvalues: For a covariance matrix, unit eigenvectors define principal directions and eigenvalues give fitted sample variances. A distinct direction is unique only up to…
  • Encoding: Representing categories using codes or indicator columns. Save the training category mapping, define unknown-category behavior and avoid implying an unsupporte…
  • Ensemble: A predictor combining multiple fitted models. Averaging, voting and learned combinations can improve performance, but improvement over every member is not guar…
  • Entropy: Negative sum of p*log2(p) over positive class proportions, measured in bits. Zero-probability terms contribute zero by convention.
  • Euclidean distance: The square root of the sum of squared coordinate differences. Feature units and fitted scaling determine the geometry being measured.
  • Explained variance: A component variance divided by the positive total variance. Cumulative fractions guide retention; high retained variance does not guarantee useful prediction…
  • F1 score: The harmonic mean of finite precision and recall, or 2TP/(2TP+FP+FN) from counts. This curriculum uses zero for the all-zero count denominator. F-beta changes…
  • Feature: An input variable used by a model. Its units, availability at prediction time and relationship to other inputs affect how useful it is.
  • Feature engineering: Constructing model inputs through transformations, encoding, interactions and missing-value handling. Learned transformation parameters must be fitted only on…
  • Feature importance: A model- and evaluation-specific measure of feature contribution. Permutation importance compares performance after shuffling a feature on held-out or eligible…
  • Feature scaling: Transforming feature magnitudes to define their relative influence. Standardization, min-max scaling and robust scaling use different fitted statistics and hav…
  • Gaussian naive Bayes: Naive Bayes with a per-class normal model for each numeric feature. A fitted mean and variance, a positive-variance policy and consistent sample or population…
  • Generalization: Performance on new data from the intended prediction setting. A train/test score gap is descriptive and may reflect overfitting, sampling variation or distribu…
  • Gini impurity: One minus the sum of squared class proportions. It is the disagreement probability for two independent draws from that empirical distribution, with replacement.
  • Gradient boosting: Adding fitted weak learners to a running model along a negative-loss-gradient direction. For squared-error loss, residuals provide the fitting targets; the fit…
  • Grid search: Evaluating a declared Cartesian grid of parameter candidates on validation or CV evidence, preserving candidate identities and a tie rule. Reserve final test d…
  • Hierarchical clustering: Agglomerative clustering repeatedly merges groups according to a linkage rule. A fitted hierarchy can be cut at a chosen count or height; a flat result still n…
  • Hyperparameter: A modeling setting chosen outside the ordinary parameter fit, such as neighbor count or tree depth. It may be selected from training-side validation evidence a…
  • Impurity: How mixed a node class distribution is. Gini and entropy are zero for a pure node and maximal at a uniform distribution over the classes under consideration.
  • Imputation: Replacing missing values using a declared rule. Means or medians are estimated on training data and reused for new rows. Filling gaps changes the distribution…
  • Inertia (within-cluster SS): The sum of squared distances to assigned cluster centers. The globally optimal value cannot increase when adding allowed centers, but separately initialized lo…
  • Information gain: The reduction in impurity after a split, usually parent entropy minus size-weighted child entropy. This track also uses analogous Gini reductions with the impu…
  • k-Means: Clustering to reduce within-cluster squared Euclidean distance. Lloyd assignment/update iterations can reach a local solution; initialization, empty-cluster po…
  • k-means++: Randomized seeding: select an initial point, then sample each next point proportional to its squared distance to its nearest existing seed. Deterministically s…
  • k-Nearest Neighbors: Predicting from nearby stored training observations using votes or target averages. Fitting stores observations and preprocessing state; distance, feature scal…
  • Laplace smoothing: Adding one pseudocount to each entry in a fixed category vocabulary and adding its size to the total. It gives positive probabilities for that vocabulary, not…
  • Learning rate: A multiplier on each boosting update. A smaller rate often requires more rounds; it does not by itself guarantee improved generalization or lower loss on every…
  • Likelihood: The probability mass or density of observed features under a model, viewed as a function of model or class. A continuous density is not probability mass at one…
  • Linkage: A rule for cross-cluster dissimilarity: minimum pair distance for single, maximum for complete, or mean over all cross pairs for average linkage. Different rul…
  • Lloyd's algorithm: Alternating nearest-center assignment with cluster-mean updates. Use explicit tie and empty-cluster policies, a convergence criterion and an iteration cap.
  • Log-probabilities: Representing products of probabilities as sums of logarithms. Compute log densities directly where possible; taking log after a product has underflowed cannot…
  • Manhattan distance: The sum of absolute coordinate differences. It aggregates coordinate gaps linearly, in contrast with the squared aggregation inside Euclidean distance.
  • Min-max scaling: Subtracting the training minimum and dividing by its positive range. Training values lie in [0,1]; new values transformed with those same bounds may fall outsi…
  • Model matrix: A numeric matrix with observations in rows and features in columns. Fitted models require consistent feature order and preprocessing for new rows.
  • mtry: The number of candidate features considered at each forest split. Floor(sqrt(p)), bounded below by one, is a common classification default; its effect depends…
  • Naive Bayes: A classifier using Bayes scores with features modeled as conditionally independent given class. Its usefulness depends on data and modeling assumptions; it doe…
  • Noise points: Observations left unassigned under a clustering method and its parameters, labeled zero in this track DBSCAN implementation. Noise status is not a proof of mea…
  • One-hot encoding: One indicator column per category. The fitted category vocabulary and column order must be retained when transforming new observations.
  • Ordinal encoding: Mapping genuinely ordered categories to codes using a declared dictionary. Numeric codes express an order but do not automatically establish meaningful equal s…
  • Out-of-bag (OOB) error: Out-of-bag predictions for a row use only ensemble members whose bootstrap samples omitted that row. Coverage must be reported; preprocessing, tuning and repea…
  • Overfitting: Learning sample-specific structure that does not transfer to the intended new-data setting. Strong training performance with weaker held-out results is a warni…
  • Pipeline: An executable sequence with an explicit information boundary: reserve test data, fit preparation and candidate models inside CV, select using validation, refit…
  • Posterior: The conditional class probability after accounting for observed features under the assumed model. Normalize prior-times-likelihood scores when their total is p…
  • Precision: TP/(TP+FP), the true fraction among positive predictions. It is undefined with no predicted positives and is reported as NA here; a high fraction need not mean…
  • Principal component analysis: A linear transformation using orthogonal directions of greatest fitted variance. It can be computed from covariance eigendecomposition or centered-data SVD; op…
  • Prior: A class probability before the query features are considered, often estimated from training class fractions. Prior times likelihood is an unnormalized class sc…
  • Projection & reconstruction: Multiplying consistently centered/scaled rows by fitted component directions yields scores. Multiplying scores by transposed directions reconstructs fitted coo…
  • Pruning: Replacing subtrees with terminal nodes or restricting growth. A cost-complexity comparison must use comparable loss units and a declared leaf penalty; its tuni…
  • Rand index: For at least two observations, the fraction of unordered pairs whose together/apart relationship agrees between partitions. This unadjusted index differs from…
  • Random forest: An ensemble of bootstrap-trained trees that considers sampled feature subsets at each split. Feature subsampling encourages diversity but does not guarantee in…
  • Recall: TP/(TP+FN), the fraction of actual positives detected. It is undefined with no actual positives and is reported as NA here. Raising a fixed-score threshold can…
  • Regression (ML): Supervised prediction of a numeric outcome. This track includes neighbor averaging and fitted squared-loss boosting; evaluation must use held-out outcomes.
  • Regularization: Adding a nonnegative weighted penalty to a compatible loss or otherwise constraining fitting. The penalty and its weight influence the chosen model; better gen…
  • Residual: Observed target minus current prediction. Residuals are the weak-learner targets for squared-error boosting, while other losses use their own gradient targets.
  • ROC curve: A curve of TPR vertically against FPR horizontally over common score thresholds, with both true classes present. Scores need not be calibrated probabilities; p…
  • Scree plot: A plot of ordered eigenvalues used to inspect possible retention cutoffs. Largest adjacent drop is one heuristic; the Kaiser >1 rule has its intended interp…
  • Silhouette score: For a non-singleton point, (b-a)/max(a,b), where a is mean distance to other own-cluster members and b is the minimum mean distance to any other cluster. Singl…
  • Stacking: Fitting a meta-model on base-model predictions. Out-of-fold predictions help keep meta-training targets separate from the fitting of the corresponding base pre…
  • Standardize (z-score): Subtracting a fitted mean and dividing by a fitted standard deviation. Training values have mean zero and sample sd one when sd is positive; new data uses the…
  • Supervised learning: Learning a relationship between inputs and known targets from labeled training examples. Classification predicts categories; regression predicts numeric outcom…
  • The elbow method: A heuristic cut where additional components or clusters contribute less improvement. A numerical largest-gap or second-difference rule is one convention, not a…
  • Train/test split: Separating data used for fitting and model selection from a reserved final evaluation set. Group and time structure can require more careful splitting than ran…
  • Underfitting: Failing to capture useful structure in the problem, often because the representation or model is too restrictive. Poor scores can also result from noise, unava…
  • Unsupervised learning: Finding structure without using a target label for fitting. Clustering and dimensionality reduction can be evaluated using geometry, stability, domain knowledg…