warming up your workspace

Bioinformatics with Python

Learn Python through DNA strings, candidate translations, alignment, sequence differences, parsing, population summaries and bounded assembly experiments.

11 projects, 275 hands-on levels, run in your browser.

Syllabus

  • Foundations: code through bioinformatics: Never written code before? Start here. You will learn the absolute basics of Python, output, variables, types, decisions, loops, and functions, using DNA bases, sequences, codons, and GC content as your playground. By the end you are ready for Project 1.
  • DNA Basics: Represent canonical DNA as strings: validate the declared alphabet, count bases, distinguish complements and reverse complements, compute composition and find overlapping motif positions. Sequence patterns are observations, not proof of function.
  • Transcription & Translation: Model the RNA counterpart of a coding strand, use the standard genetic code, translate declared frames and scan complete ATG-to-stop candidates on both strands. Retain coordinates and distinguish candidate coding regions from verified genes.
  • Sequence Composition & Motifs: Characterize a sequence statistically and find the patterns in it: nucleotide and dinucleotide frequencies, CpG sites, melting temperature, k-mer spectra, exact and approximate motif matching, and building a consensus from a set of related sequences.
  • Pairwise Alignment: Build edit-distance and scored global/local dynamic-programming matrices, recover deterministic traceback paths, and report identity with an explicit denominator. An optimal score depends on its scoring model and does not prove homology.
  • Mutations & Variants: Compare synthetic prealigned references and samples. Report substitutions, individual and joint codon contexts, partial codons and stop changes, and model explicitly described indels. Keep sequence effects separate from functional or clinical claims.
  • FASTA & FASTQ Parsing: Real sequence data arrives in FASTA and FASTQ files. Parse FASTA records (header + sequence, possibly multi-line), handle many records at once, parse FASTQ reads with their quality strings, decode Phred quality scores, and filter low-quality reads, the daily bread of working with sequencing data.
  • Phylogenetics: Measure prealigned sequence differences, handle the finite JC model domain, and build auditable UPGMA trees with cluster-size weighting and branch lengths. A tree is a model-based hypothesis; the algorithm does not validate a molecular clock.
  • Population Genetics: Count diploid genotypes and allele copies, keep missing data explicit, compare observed and expected diversity, and report HWE approximation applicability. Compare consistently weighted populations while retaining undefined monomorphic cases.
  • Genome Search & Assembly: Build positional k-mer indexes, verify whole reads after seeding, and assemble bounded forward-only error-free examples. Unsupported joins remain separate contigs, and each resulting contig is checked against the reference.
  • Capstone: Characterize an Unknown Sample: Build a source-bound dossier for a fictional DNA sample: input validation, composition, complete candidate ORFs, restriction recognition and first-strand cuts, and an explicitly prealigned reference comparison. Preserve inputs and limitations without claiming organism identification or biological function.

Key concepts

  • Codon: A triplet interpreted in a declared reading frame under a genetic code. Sense codons specify amino acids; stop codons signal termination. DNA coding-strand tri…
  • Distance matrix: A symmetric table of pairwise distances between all sequences, the input to tree-building.
  • GC content: The fraction of analyzed bases that are G or C, with alphabet and denominator stated. It is a composition descriptor, not an organism identifier or sufficient…
  • Hamming distance: The number of positions at which two equal-length sequences differ, the simplest measure of sequence dissimilarity.
  • Motif: A short sequence pattern used in a search or statistical model. Some motifs have demonstrated biological roles, but a string match alone does not show binding,…
  • Nucleotide: A building block of DNA/RNA, represented by a letter: A, C, G, T (or U in RNA). A sequence is a string of these.
  • Phylogenetic tree: A branching hypothesis of evolutionary relationships estimated from data under a method and assumptions. Distance-based tree building is one approach; uncertai…
  • Reading frame: Where you start grouping a sequence into codons; the three possible offsets give three frames (six with the reverse strand).
  • Scoring matrix: A table of match/mismatch scores used to evaluate an alignment, rewarding biologically likely substitutions.
  • Sequence alignment: An arrangement of sequences, allowing gaps, optimized under a stated scoring rule. Similarity can support a homology hypothesis with suitable context; an align…
  • UPGMA: Average-linkage clustering: merge the closest clusters and weight updated distances by cluster sizes. Heights yield an ultrametric rooted tree; evolutionary in…