warming up your workspace

Deep Learning with Python

Build modern AI from scratch: neural networks, backprop, CNNs, transformers, a tiny GPT.

11 projects, 275 hands-on levels, run in your browser.

Syllabus

  • Foundations: code through AI: Never written code before? Start here. You will learn the absolute basics of Python, output, variables, types, decisions, loops, and functions, using neural networks, weights, and predictions as your playground. By the end you are ready for Project 1.
  • Foundations: From Regression to the Neuron: Start with a model that predicts numbers and gradient descent that adjusts its parameters. Fit linear regression, work with multiple features, and build weighted sums and activation functions. A sigmoid neuron with a binary classification loss gives logistic regression. Finish by training and evaluating a neuron on the AND truth table.
  • Neural Networks: Combine neurons into dense layers and layers into a multilayer perceptron. Write matrix-based forward passes, use softmax for class probabilities, and construct a hidden-layer network for XOR, which a single linear decision boundary cannot separate. Finish with reusable initialization, prediction and accuracy reporting; training these networks comes next.
  • Backpropagation and Autograd: Backpropagation computes loss gradients by applying the chain rule backward through a computation graph; an optimizer then uses those gradients to update parameters. Differentiate a neuron and a two-layer network, compare derivatives with finite differences, build a scalar autograd engine, and train an XOR network while retaining its weights and loss history.
  • Training Neural Networks: Build losses that score predictions and optimizers that update parameters: SGD, momentum, RMSprop and Adam. Combine forward and backward passes into mini-batch training with persistent optimizer state and epoch histories. Train a small regression network on synthetic data and measure its progress; a chosen update rule does not guarantee improvement on every step.
  • Deep Classification: Train a multi-class network on synthetic clusters and distinguish training accuracy from performance on held-out data. Implement stable softmax cross-entropy, L2 regularization, dropout, validation splits, early stopping and confusion-based metrics. Assemble a classifier that retains the best validation parameters and reports measured results; regularization benefits must be evaluated.
  • Convolutional Networks: Convolutional networks build in local connectivity and shared filters for spatial inputs. Implement the unflipped sliding-kernel convention used in CNNs, inspect hand-designed filters, pool feature maps and differentiate a convolutional layer. Train a tiny CNN on synthetic vertical and horizontal bars, inspect its learned filters and evaluate held-out images.
  • Sequences and Recurrent Networks: Process ordered inputs one step at a time with a recurrent hidden state, a compressed representation rather than a perfect memory. Build the RNN cell and sequence forward pass, derive backpropagation through time, examine gradient decay and clipping, and train a character model. Finish with reusable training, reporting, greedy generation and temperature sampling.
  • Tokenization and Embeddings: Turn raw text into stable token IDs and learn dense vectors for those IDs. Build character, word and UTF-8 byte tokenizers, embedding lookup and accumulated gradients, cosine similarity and vector analogy calculations. Train a small bag-of-embeddings sentiment classifier and inspect its predictions and neighbors; useful semantic geometry depends on the data and objective, not on lookup alone.
  • Attention and Transformers: Build scaled dot-product attention, causal masks, sinusoidal position information and multiple attention heads. Combine output projections, residual connections, layer normalization and feed-forward layers into stacked transformer blocks and a reusable next-token forward model. Full attention compares token pairs and has quadratic sequence cost at fixed width; generation still proceeds one token at a time.
  • Capstone: A Tiny GPT: Build and train an educational character-level autoregressive transformer. Compose embeddings and positions, strict causal attention, residual feed-forward layers and an output head; differentiate all eleven parameter arrays and check them numerically. Train on a tiny corpus with Adam, retain optimizer state and history, and provide reporting, greedy generation and temperature sampling. This single-block training model omits layer normalization and does not claim the architecture or capabilities of a production GPT.

Key concepts

  • Activation function: A function applied to a weighted sum plus bias. Nonlinear choices such as ReLU, sigmoid and tanh let stacked layers represent nonlinear relationships; affine l…
  • Attention: A weighted combination of value vectors using query-key scores. Masks restrict available positions. Attention weights describe this combination, but are not by…
  • Backpropagation: Reverse application of the chain rule to compute loss gradients for parameters and intermediate values. Contributions from every path to a reused value must be…
  • Batch: A group of examples processed together. Gradients may be summed or averaged before an update, according to the objective. Batch size affects memory, computatio…
  • Convolution (CNN): The CNN sliding-kernel operation used here is cross-correlation: multiply each input patch by the kernel without reversing it, then sum. Shared local filters b…
  • Cross-entropy: A classification loss that penalizes low probability assigned to the correct class. It is commonly paired with softmax because softmax turns raw scores into cl…
  • Embedding: A dense vector representation of an item, often a row in a learned lookup table. Training can organize useful relationships, but random initialization, vector…
  • Epoch: One full pass through the training dataset. Multiple epochs let the model refine its parameters, but too many can cause overfitting if validation performance s…
  • Gradient descent: Subtracting a learning-rate multiple of the gradient from the parameters. The negative gradient is the steepest local direction in Euclidean geometry at a diff…
  • Learning rate: The step size used when updating parameters from gradients. Too small trains slowly; too large can overshoot, diverge, or bounce around the minimum.
  • Loss function: A numerical training objective that measures prediction error, sometimes with a regularization term. Its units, sum or mean reduction, and definition must matc…
  • Neuron: A unit that computes a weighted sum of inputs plus a bias, then applies an activation function; networks stack many.
  • Overfitting: When a model learns the training examples too specifically and performs worse on new data. It often appears as training loss improving while validation loss st…
  • Parameter: A learnable number inside a model, such as a weight or bias. Training changes parameters; hyperparameters like learning rate or batch size are chosen outside t…
  • Random seed: A value used to initialize pseudorandom choices such as weight initialization, shuffling, or sampling. Setting a seed makes experiments easier to reproduce, th…
  • Recurrent network (RNN): A network with a hidden state carried across a sequence, letting it model order and context.
  • Softmax: Normalizes finite scores using exponentials to produce class probabilities. Exact probabilities are positive and sum to one; floating-point computation can und…
  • Tensor: A multidimensional numeric array used to hold inputs, activations, parameters, gradients, and batches. A vector of token ids, a batch of images, and a weight m…
  • Tokenization: Splitting text into units and mapping those units to IDs in a consistent vocabulary. Character, whitespace-word and byte schemes have different round-trip and…
  • Transformer: A family of models built from attention, position information, feed-forward layers and residual connections, usually with normalization. Autoregressive variant…
  • Weights and bias: Learnable coefficients and offsets in a model. An optimizer adjusts them using a training objective; an individual update can increase loss if its step is unsu…