Machine Learning in R
Learn predictive modeling in R through feature preparation, nearest neighbors, naive Bayes, trees, ensembles, clustering, PCA and evaluation. Build teaching-scale algorithms and compose fitted projects, using base R library routines for selected operations and independent comparisons. Finish with validation-based model selection and a reserved test evaluation.
11 projects, 275 hands-on levels, run in your browser.
Syllabus
- Foundations: code through machine learning: Never written code before? Start here. You will learn the basics of R, functions, variables, types, decisions, loops, and vectors, with simple machine-learning examples. By the end you are ready for Project 1.
- Features & the Model Matrix: Build a reusable model-matrix preparation workflow: reserve test rows, fit numeric imputation and scaling, encode categories with saved dictionaries, and apply transformations to new rows in the same feature order. Inspect shape, unknown-category and constant-feature policies. Fit learned preparation inside each training fold so validation and final test information stay outside fitting.
- k-Nearest Neighbors: Build distances, neighbor ranking, deterministic votes and weighted classification/regression in R. Store training data and fitted preprocessing for reusable prediction, then choose neighbor counts using held-out validation. Explore the effects of scaling, duplicate observations and dimension without assuming one distance rule suits every dataset.
- Naive Bayes: Build Gaussian and multinomial naive Bayes from class priors, fitted feature distributions, fixed vocabularies and add-one smoothing. Use direct log scores to limit underflow and keep class, variance and unknown-token policies explicit. Compose trained continuous-feature and spam-style count classifiers with reusable prediction.
- Decision Trees: Build a teaching-scale CART-like numeric classification tree from Gini/entropy, valid threshold search, weighted child impurity and recursive growth. Store feature indices and leaf predictions, support depth and minimum-leaf stops, and perform structural cost-complexity pruning with comparable loss units. Inspect fitted predictions and validate complexity choices on held-out data.
- Ensembles: Compose fitted bagged trees and a feature-subsampled forest with bootstrap identities and eligible out-of-bag predictions. Build weighted binary AdaBoost and feature-dependent squared-loss gradient boosting, then explore prediction blending and out-of-fold stacking. Track seeds, stopping policies, signed permutation changes and coverage; ensemble improvement is evaluated rather than guaranteed.
- k-Means & Partitional Clustering: Build capped Lloyd assignment/update fitting with explicit empty-cluster policies, stored centers and diagnostics, and multiple starts. Distinguish deterministic farthest-first helpers from randomized k-means++ seeding. Evaluate inertia and silhouette and compare compatible fitted results with R kmeans; initialization and local optima limit universal monotonicity claims.
- Hierarchical & Density Clustering: Build distance/linkage calculations, an agglomerative merge sequence and cuts, and complete DBSCAN core-led expansion with border/noise policies. Compare compatible hierarchical results with R dist/hclust and inspect pairwise partition agreement. These methods offer different geometries, with scale and parameter assumptions rather than guarantees about every cluster shape.
- PCA & Dimensionality Reduction: Build PCA from training-fitted centering, optional scaling, sample covariance and eigendecomposition. Store a retained basis for projection and inverse transformation of new rows, inspect reduced reconstruction error and variance retention, and compare aligned results with prcomp. Respect global sign ambiguity, repeated-eigenvalue subspaces, rank deficiency and wide-matrix dimensions.
- Model Evaluation & Selection: Build aligned confusion counts, defined metric conventions, complete threshold ROC records and pairwise/trapezoidal AUC checks. Run validation with fold-local fitted preparation and models, preserve candidate identities and select from comparable held-out scores. Refit the chosen candidate before reserved test evaluation and report sample size, class counts and limitations.
- Capstone: A Predictive Modeling Pipeline: Assemble a complete predictive modeling pipeline with training-fitted imputation and scaling, real kNN/Gaussian naive Bayes/Gini-stump candidates, fold-local fitting and validation-based hyperparameter/model selection. Retain per-candidate held-out predictions and scores, refit the winner on all training rows, and evaluate once on reserved test data with counts, metrics and a training-fitted baseline. Expose fitted state and prediction APIs; retain the one-feature report as a compatible helper rather than the whole system.
Key concepts
- Accuracy: The fraction of aligned nonempty predictions that match truth. Interpret it with class frequencies, error types, sample size and the selected baseline.
- AdaBoost: Binary boosting that weights learners using their weighted error and reweights observations after each round. Alpha is half the log odds of correctness; zero e…
- AUC: Area under a complete ROC curve, equal to positive-over-negative ranking probability plus half the tie probability. It measures ranking, not calibration; finit…
- Bagging: Fitting models on bootstrap resamples and averaging or voting their predictions. Its variance effect depends on member variability and correlation; bias is not…
- Baseline: A declared reference procedure, such as a training-fitted constant-class predictor. Evaluate it on the same held-out rows; beating one accuracy baseline is nei…
- Bayes' theorem: Posterior probability is proportional to prior probability times likelihood, with normalization over the possible classes when the total evidence is positive.
- Bias-variance tradeoff: Under squared-error assumptions, expected error decomposes into squared bias, estimator variance and irreducible noise. This is not a universal decomposition o…
- Binning: Replacing continuous values with interval categories. Boundaries, endpoint inclusion and out-of-range behavior must be explicit; learned cut points are fitted…
- Boosting: Sequentially fitting models to improve a combined predictor according to a chosen objective. AdaBoost changes sample weights; gradient boosting fits negative-l…
- Bootstrap sample: Sampling n observation identities with replacement from n rows. The expected omitted fraction is (1-1/n)^n, approaching about 0.368; an individual sample need…
- CART: Classification and Regression Trees use greedy recursive binary splitting with terminal predictions. This track implements a teaching-scale numeric classificat…
- Class imbalance: Unequal class frequencies, sometimes making overall accuracy conceal important minority-class errors. Choose metrics and baselines to match the intended use; a…
- Classification: Supervised prediction of a discrete class. A declared positive class, error costs and evaluation design determine which performance metrics are informative.
- Clustering: Grouping observations using a declared similarity or density notion. k-means favors squared-distance compactness, hierarchy uses linkage, and DBSCAN uses densi…
- Confusion matrix: Counts of actual and predicted class combinations with explicitly declared axes. The binary matrix supports accuracy, precision, recall and F1 at one threshold…
- Core, border, noise: In the DBSCAN convention here, core means at least min_pts observations within inclusive eps, including self. A non-core point near a core is border; otherwise…
- Cosine similarity: The dot product divided by the product of nonzero vector lengths. It compares orientation rather than magnitude; 1 minus this similarity is the cosine dissimil…
- Covariance matrix: A matrix of feature variances on the diagonal and covariances off it. Complete-data sample covariance is symmetric positive semidefinite; symmetry gives real e…
- Cross-validation: Repeated fitting on training-fold complements and scoring held-out folds. Include learned preprocessing inside each fit. Fold scores depend on split design and…
- Curse of dimensionality: Challenges that can arise as dimension increases, including sparse coverage and reduced contrast between near and far distances under some data distributions.…
- Data leakage: Use of information unavailable to the intended training or prediction stage, such as held-out labels, future data, shared entities or all-data preprocessing. T…
- DBSCAN: Density clustering with inclusive radius eps and minimum neighborhood size including self here. Core points expand connected regions; border points can join wi…
- Decision threshold: A cutoff converting an oriented score into a label. This track uses score>=threshold for positive. Scores may be uncalibrated; precision need not increase m…
- Decision tree: A model that routes observations through nested split rules to terminal predictions. Small trees can be easy to inspect; interpretability depends on size and c…
- Dendrogram: A tree recording hierarchical merges and their dissimilarity heights. Single-linkage heights are nondecreasing, with ties allowed; a chosen count or height def…
- Distance metric: A distance obeying nonnegativity, identity, symmetry and the triangle inequality. Euclidean and Manhattan distances are metrics; 1 minus cosine similarity gene…
- Dummy encoding: Indicator encoding with one reference level omitted, typically leaving k-1 columns. With an intercept, this removes the full-indicator linear dependence; absen…
- Eigenvectors & eigenvalues: For a covariance matrix, unit eigenvectors define principal directions and eigenvalues give fitted sample variances. A distinct direction is unique only up to…
- Encoding: Representing categories using codes or indicator columns. Save the training category mapping, define unknown-category behavior and avoid implying an unsupporte…
- Ensemble: A predictor combining multiple fitted models. Averaging, voting and learned combinations can improve performance, but improvement over every member is not guar…
- Entropy: Negative sum of p*log2(p) over positive class proportions, measured in bits. Zero-probability terms contribute zero by convention.
- Euclidean distance: The square root of the sum of squared coordinate differences. Feature units and fitted scaling determine the geometry being measured.
- Explained variance: A component variance divided by the positive total variance. Cumulative fractions guide retention; high retained variance does not guarantee useful prediction…
- F1 score: The harmonic mean of finite precision and recall, or 2TP/(2TP+FP+FN) from counts. This curriculum uses zero for the all-zero count denominator. F-beta changes…
- Feature: An input variable used by a model. Its units, availability at prediction time and relationship to other inputs affect how useful it is.
- Feature engineering: Constructing model inputs through transformations, encoding, interactions and missing-value handling. Learned transformation parameters must be fitted only on…
- Feature importance: A model- and evaluation-specific measure of feature contribution. Permutation importance compares performance after shuffling a feature on held-out or eligible…
- Feature scaling: Transforming feature magnitudes to define their relative influence. Standardization, min-max scaling and robust scaling use different fitted statistics and hav…
- Gaussian naive Bayes: Naive Bayes with a per-class normal model for each numeric feature. A fitted mean and variance, a positive-variance policy and consistent sample or population…
- Generalization: Performance on new data from the intended prediction setting. A train/test score gap is descriptive and may reflect overfitting, sampling variation or distribu…
- Gini impurity: One minus the sum of squared class proportions. It is the disagreement probability for two independent draws from that empirical distribution, with replacement.
- Gradient boosting: Adding fitted weak learners to a running model along a negative-loss-gradient direction. For squared-error loss, residuals provide the fitting targets; the fit…
- Grid search: Evaluating a declared Cartesian grid of parameter candidates on validation or CV evidence, preserving candidate identities and a tie rule. Reserve final test d…
- Hierarchical clustering: Agglomerative clustering repeatedly merges groups according to a linkage rule. A fitted hierarchy can be cut at a chosen count or height; a flat result still n…
- Hyperparameter: A modeling setting chosen outside the ordinary parameter fit, such as neighbor count or tree depth. It may be selected from training-side validation evidence a…
- Impurity: How mixed a node class distribution is. Gini and entropy are zero for a pure node and maximal at a uniform distribution over the classes under consideration.
- Imputation: Replacing missing values using a declared rule. Means or medians are estimated on training data and reused for new rows. Filling gaps changes the distribution…
- Inertia (within-cluster SS): The sum of squared distances to assigned cluster centers. The globally optimal value cannot increase when adding allowed centers, but separately initialized lo…
- Information gain: The reduction in impurity after a split, usually parent entropy minus size-weighted child entropy. This track also uses analogous Gini reductions with the impu…
- k-Means: Clustering to reduce within-cluster squared Euclidean distance. Lloyd assignment/update iterations can reach a local solution; initialization, empty-cluster po…
- k-means++: Randomized seeding: select an initial point, then sample each next point proportional to its squared distance to its nearest existing seed. Deterministically s…
- k-Nearest Neighbors: Predicting from nearby stored training observations using votes or target averages. Fitting stores observations and preprocessing state; distance, feature scal…
- Laplace smoothing: Adding one pseudocount to each entry in a fixed category vocabulary and adding its size to the total. It gives positive probabilities for that vocabulary, not…
- Learning rate: A multiplier on each boosting update. A smaller rate often requires more rounds; it does not by itself guarantee improved generalization or lower loss on every…
- Likelihood: The probability mass or density of observed features under a model, viewed as a function of model or class. A continuous density is not probability mass at one…
- Linkage: A rule for cross-cluster dissimilarity: minimum pair distance for single, maximum for complete, or mean over all cross pairs for average linkage. Different rul…
- Lloyd's algorithm: Alternating nearest-center assignment with cluster-mean updates. Use explicit tie and empty-cluster policies, a convergence criterion and an iteration cap.
- Log-probabilities: Representing products of probabilities as sums of logarithms. Compute log densities directly where possible; taking log after a product has underflowed cannot…
- Manhattan distance: The sum of absolute coordinate differences. It aggregates coordinate gaps linearly, in contrast with the squared aggregation inside Euclidean distance.
- Min-max scaling: Subtracting the training minimum and dividing by its positive range. Training values lie in [0,1]; new values transformed with those same bounds may fall outsi…
- Model matrix: A numeric matrix with observations in rows and features in columns. Fitted models require consistent feature order and preprocessing for new rows.
- mtry: The number of candidate features considered at each forest split. Floor(sqrt(p)), bounded below by one, is a common classification default; its effect depends…
- Naive Bayes: A classifier using Bayes scores with features modeled as conditionally independent given class. Its usefulness depends on data and modeling assumptions; it doe…
- Noise points: Observations left unassigned under a clustering method and its parameters, labeled zero in this track DBSCAN implementation. Noise status is not a proof of mea…
- One-hot encoding: One indicator column per category. The fitted category vocabulary and column order must be retained when transforming new observations.
- Ordinal encoding: Mapping genuinely ordered categories to codes using a declared dictionary. Numeric codes express an order but do not automatically establish meaningful equal s…
- Out-of-bag (OOB) error: Out-of-bag predictions for a row use only ensemble members whose bootstrap samples omitted that row. Coverage must be reported; preprocessing, tuning and repea…
- Overfitting: Learning sample-specific structure that does not transfer to the intended new-data setting. Strong training performance with weaker held-out results is a warni…
- Pipeline: An executable sequence with an explicit information boundary: reserve test data, fit preparation and candidate models inside CV, select using validation, refit…
- Posterior: The conditional class probability after accounting for observed features under the assumed model. Normalize prior-times-likelihood scores when their total is p…
- Precision: TP/(TP+FP), the true fraction among positive predictions. It is undefined with no predicted positives and is reported as NA here; a high fraction need not mean…
- Principal component analysis: A linear transformation using orthogonal directions of greatest fitted variance. It can be computed from covariance eigendecomposition or centered-data SVD; op…
- Prior: A class probability before the query features are considered, often estimated from training class fractions. Prior times likelihood is an unnormalized class sc…
- Projection & reconstruction: Multiplying consistently centered/scaled rows by fitted component directions yields scores. Multiplying scores by transposed directions reconstructs fitted coo…
- Pruning: Replacing subtrees with terminal nodes or restricting growth. A cost-complexity comparison must use comparable loss units and a declared leaf penalty; its tuni…
- Rand index: For at least two observations, the fraction of unordered pairs whose together/apart relationship agrees between partitions. This unadjusted index differs from…
- Random forest: An ensemble of bootstrap-trained trees that considers sampled feature subsets at each split. Feature subsampling encourages diversity but does not guarantee in…
- Recall: TP/(TP+FN), the fraction of actual positives detected. It is undefined with no actual positives and is reported as NA here. Raising a fixed-score threshold can…
- Regression (ML): Supervised prediction of a numeric outcome. This track includes neighbor averaging and fitted squared-loss boosting; evaluation must use held-out outcomes.
- Regularization: Adding a nonnegative weighted penalty to a compatible loss or otherwise constraining fitting. The penalty and its weight influence the chosen model; better gen…
- Residual: Observed target minus current prediction. Residuals are the weak-learner targets for squared-error boosting, while other losses use their own gradient targets.
- ROC curve: A curve of TPR vertically against FPR horizontally over common score thresholds, with both true classes present. Scores need not be calibrated probabilities; p…
- Scree plot: A plot of ordered eigenvalues used to inspect possible retention cutoffs. Largest adjacent drop is one heuristic; the Kaiser >1 rule has its intended interp…
- Silhouette score: For a non-singleton point, (b-a)/max(a,b), where a is mean distance to other own-cluster members and b is the minimum mean distance to any other cluster. Singl…
- Stacking: Fitting a meta-model on base-model predictions. Out-of-fold predictions help keep meta-training targets separate from the fitting of the corresponding base pre…
- Standardize (z-score): Subtracting a fitted mean and dividing by a fitted standard deviation. Training values have mean zero and sample sd one when sd is positive; new data uses the…
- Supervised learning: Learning a relationship between inputs and known targets from labeled training examples. Classification predicts categories; regression predicts numeric outcom…
- The elbow method: A heuristic cut where additional components or clusters contribute less improvement. A numerical largest-gap or second-difference rule is one convention, not a…
- Train/test split: Separating data used for fitting and model selection from a reserved final evaluation set. Group and time structure can require more careful splitting than ran…
- Underfitting: Failing to capture useful structure in the problem, often because the representation or model is too restrictive. Poor scores can also result from noise, unava…
- Unsupervised learning: Finding structure without using a target label for fitting. Clustering and dimensionality reduction can be evaluated using geometry, stability, domain knowledg…