Data Science with Python
Learn data science and machine learning by doing: pandas, analysis, visualization, statistics, and ML algorithms built from scratch.
11 projects, 275 hands-on levels, run in your browser.
Syllabus
- Foundations: code through data science: Never written code before? Start here. You will learn the absolute basics of Python, output, variables, types, decisions, loops, and functions, using datasets, samples, and averages as your playground. By the end you are ready for Project 1.
- Data with pandas: pandas is the foundation of data work in Python. Master the two core structures: the Series (a labeled 1D array) and the DataFrame (a labeled table). Load data, inspect it, select rows and columns, filter, and compute summary statistics, the everyday vocabulary of every data scientist.
- Data Cleaning & Wrangling: Real data is messy. Detect and handle missing values, find and drop duplicates, fix wrong data types, standardize inconsistent text, and transform columns with map, apply, and binning. Cleaning is where data scientists spend most of their time, and getting it right is what makes every later analysis trustworthy.
- Exploratory Data Analysis: Before modeling, you explore. Compute summary statistics (center and spread), aggregate by group, examine distributions with value counts, and measure relationships between variables with correlation. EDA is how you build intuition for a dataset and discover the patterns worth investigating.
- Data Visualization: A chart often reveals what a table of numbers hides. Build the four workhorse plots with matplotlib: line charts for trends, bar charts for comparing categories, histograms for distributions, and scatter plots for relationships. Label them clearly and choose the right chart for the question.
- Statistics & Inference: Calculate descriptive statistics, sampling uncertainty, illustrative confidence intervals and hypothesis tests. State the assumptions behind each method and distinguish observed effects, statistical evidence and practical decisions.
- Linear Regression: Apply a linear model, fit coefficients with least squares, evaluate squared errors and R-squared, and learn parameters through finite gradient-descent updates. Combine these steps in a house-price teaching example with explicit units and validation limits.
- Classification: Predict categories, not numbers. Build two classifiers from scratch: k-nearest neighbors (classify by the majority vote of the closest points) and logistic regression (a sigmoid model trained with gradient descent). Then evaluate them properly with the confusion matrix, accuracy, precision, recall, and F1.
- Model Evaluation: Build train/test partitions and cross-validation, compare model flexibility, and learn the statistical meaning of bias and variance. Finish with training-only candidate selection and a calculated final report, while understanding evaluation limits.
- Unsupervised Learning: Find structure in data that has no labels. Build k-means clustering from scratch (group points by similarity), learn to choose the number of clusters, and implement PCA (principal component analysis) to reduce dimensions while keeping the most variance, the two pillars of unsupervised learning.
- Capstone: An End-to-End ML Project: Work through a small synthetic study-hours dataset: prepare paired observations, explore them, fit on the first six cleaned rows, evaluate reserved rows and return a calculated model and prediction report.
Key concepts
- Bias-variance tradeoff: Too simple a model underfits (high bias); too complex overfits (high variance). Good models balance the two.
- Classification: A supervised learning task where the target is a category, such as churn/not churn or species name. Classifiers are evaluated with confusion matrices, precisio…
- Confusion matrix: A table of true/false positives and negatives summarizing a classifier's mistakes, the basis of precision and recall.
- Data leakage: When training accidentally uses information that would not be available at prediction time. Leakage can make evaluation look excellent while the real deployed…
- DataFrame: A rectangular table with labeled rows and columns, usually handled with pandas. DataFrames are the main workspace for cleaning, joining, aggregating, plotting,…
- F1 score: The harmonic mean of precision and recall. It is useful when positive cases are rare and you need one number that punishes both false positives and false negat…
- Feature and label: A feature is an input variable describing an example; the label is the target you predict. Supervised learning maps features to labels.
- Gradient descent: Minimizing a loss by repeatedly stepping the parameters downhill along the negative gradient, the workhorse of model fitting.
- Hypothesis test: A procedure that weighs evidence against a null hypothesis, deciding whether an effect is statistically significant.
- k-means clustering: An unsupervised method that partitions points into k groups by alternately assigning points to the nearest center and moving centers to the mean.
- Loss function: A number measuring how wrong a model's predictions are; training minimizes it.
- Missing value: A blank or unknown value, often represented as NaN , None , or null. Missingness can mean measurement failure, non-applicability, or omitted data, so the right…
- Outlier: A value far from the main pattern of the data. Outliers can be real important events, measurement errors, or data-entry mistakes, so they should be inspected b…
- Overfitting: When a model memorizes training noise and fails on new data; the gap between train and test performance reveals it.
- p-value: The probability of seeing data at least this extreme if the null hypothesis were true; small p casts doubt on the null.
- Precision and recall: Precision is the fraction of positive predictions that are correct; recall is the fraction of actual positives found. They trade off; F1 combines them.
- R-squared: The fraction of variance in the target explained by a regression model; 1 is perfect, 0 is no better than the mean.
- Regression: A supervised learning task where the target is numeric, such as price, temperature, or demand. Regression models are often evaluated with errors, residuals, an…
- Standardization: Rescaling a feature by subtracting its mean and dividing by its standard deviation. It helps distance-based and gradient-based models treat features on differe…
- Train/test split: Holding out part of the data to evaluate a model on examples it never trained on, the basic guard against fooling yourself.