warming up your workspace

GPU Computing and Graphics with Python

Learn GPU programming concepts through CPU Python and NumPy models, then assemble triangle rendering, sphere ray tracing, image filters and particle simulations. Exercises do not execute CUDA kernels or measure GPU speedups.

11 projects, 275 hands-on levels, run in your browser.

Syllabus

  • Foundations: Python for GPU models: Learn Python output, variables, types, decisions, loops and functions using small work-count and array examples. Printed device labels are practice text; the exercises run CPU Python.
  • The GPU Programming Model: Model per-index callbacks, blocks, grids, guards and array operations in CPU Python/NumPy. Assemble an AXPY launcher and compare its values with a reference expression. This is a partial CUDA indexing model, without GPU scheduling or hardware execution.
  • Data-Parallel Patterns: Compute maps, reductions, scans, gathers and scatters in NumPy. Distinguish ideal parallel dependency depth from execution time, then compose an energy-share pipeline from squaring, summing and division.
  • Memory Models and Tiling: Use explicit latency, transaction and traffic models to reason about locality. Implement a CPU blur with copied tile halos, verify its values and report a padded logical-read ratio. Hardware latency and acceleration are not measured here.
  • Linear Algebra and GPU Models: Build vector operations, matrix-vector products and scalar/blocked matrix multiplication in CPU NumPy. Verify partial tiles and matrix shapes, then report the result with a nominal full-tile traffic factor. The code does not allocate GPU shared memory.
  • Vectors and Transforms: Use vectors, matrices and homogeneous coordinates to compose model, fixed-view camera and perspective transforms. The final output is normalized device coordinates; clipping and viewport mapping are separate steps.
  • The Rasterization Pipeline: Build a CPU screen-space triangle renderer with closed-edge integer samples, barycentric color interpolation and a depth buffer. Assemble actual overlapping triangles and display the resulting image.
  • Ray Tracing: Trace primary camera rays against one sphere, select the nearest forward intersection and apply local diffuse/ambient shading. Assemble and display a CPU grayscale image. Reflections, cast shadows and GPU ray-tracing hardware are outside this implementation.
  • Shader Concepts and Image Processing: Use CPU arrays for coordinate patterns, color maps and neighborhood filters. Distinguish unflipped cross-correlation from mathematical convolution, then display a grayscale/blur/Sobel-x pipeline with explicit cropping and magnitude mapping.
  • Particles and Physics Models: Represent particle state in NumPy, compare integration rules and calculate gravity, drag, springs and direct all-pairs forces. Assemble a seeded CPU fireworks simulation with post-step positions, matching life values and a fading trajectory display.
  • Capstone: A CPU Triangle Renderer: Assemble an outward tetrahedron, camera transforms, perspective depth, flat lighting and depth-tested rasterization into a displayed image. The renderer uses CPU Python/NumPy. It does not combine ray tracing or particles, clip crossing polygons, or execute GPU kernels.

Key concepts

  • Arithmetic intensity: Floating-point operations divided by bytes transferred at a specified memory boundary. Its value depends on which reads/writes and reuse the model counts; it i…
  • Kernel: On a GPU, a function executed by threads in a device launch. In this course, Python callbacks model selected indexing behavior on the CPU; using a kernel name…
  • Memory coalescing: Combining memory requests from a group of threads into transactions according to addresses, access sizes, alignment and hardware rules. Consecutive addresses c…
  • N-body: Computing interactions among N bodies. Direct all-pairs force evaluation has quadratic pair count; whether an implementation is compute- or memory-limited depe…
  • Rasterization: Generating covered samples from geometric primitives and interpolating attributes. This course rasterizes triangles into CPU arrays and uses a depth buffer to…
  • Ray tracing: Following mathematical rays, intersecting scene geometry and evaluating a hit or background. The course traces primary rays against a sphere and applies local…
  • Roofline model: An ideal performance ceiling: the minimum of peak compute throughput and bandwidth times arithmetic intensity, using consistent units. It omits other limits su…
  • Shared memory / tiling: Programmer-managed memory accessible to threads in a block on relevant GPU architectures. Tiling stages reusable data there and requires appropriate synchroniz…
  • SIMT: Single Instruction, Multiple Threads: a GPU execution model that groups logical threads for instruction execution. Divergent control flow and scheduling matter…
  • Thread, block, grid: In the CUDA model, a launch groups logical threads into blocks and blocks into a grid. Indices identify work items; they do not imply every logical thread runs…