"You can't understand data you can't describe."
A colorful, beginner-friendly guide to describing any dataset with numbers and charts. 📊
Imagine you have exam scores for 100 students. You can't read all 100 numbers and understand anything useful. Descriptive statistics lets you summarise that entire dataset into a handful of numbers — each one telling a different story.
This notebook walks you through every core measure from scratch, with plain-English analogies before every formula, hands-on code, and vibrant colorful visualizations. No prior statistics or math background needed.
| # | Topic | What You'll Learn |
|---|---|---|
| 1 | Mean | The average — and why outliers break it |
| 2 | Median | The middle value — resistant to outliers |
| 3 | Mode | Most frequent value — works on categories too |
| 4 | Measures of Spread | Range, variance, std deviation, IQR |
| 5 | Percentiles & Quartiles | Ranking values — Q1, Q2, Q3, ECDF |
| 6 | Skewness & Kurtosis | The shape of a distribution |
| 7 | Full Workflow | Everything together on a real multi-column DataFrame |
Every section uses the same 60-student exam score dataset so you see how each measure tells a different story about the same numbers.
📊 Histogram — score distribution with mean & median lines
📦 Box plot — IQR, quartiles, and outliers visualised
🎻 Violin plot — distribution shape per subject
📈 ECDF curve — cumulative percentile ranking
🌈 Quartile bar — colourful Q1/Q2/Q3/Q4 breakdown
↔️ Std dev comparison — same mean, three different spreads
↙↗ Skewness shapes — right, symmetric, and left-skewed side by side
- Go to kaggle.com/code → New Notebook
- File → Import Notebook → upload
statistics_descriptive_stats.ipynb - Run All
▶️ — all charts render instantly, no installs needed
- Go to colab.research.google.com
- File → Upload notebook
- Runtime → Run All
numpy>=1.21.0
pandas>=1.3.0
matplotlib>=3.5.0
scipy>=1.7.0pip install -r requirements.txt✅ All pre-installed on Kaggle and Google Colab.
| Measure | Code | When to use |
|---|---|---|
| Mean | np.mean(data) |
Symmetric data, no extreme outliers |
| Median | np.median(data) |
Skewed data, income, house prices |
| Mode | scipy.stats.mode(data) |
Categorical data, shoe sizes, votes |
| Std Dev | np.std(data, ddof=1) |
Spread in original units |
| IQR | np.percentile(d,75) - np.percentile(d,25) |
Spread without outlier influence |
| Percentile | np.percentile(data, p) |
Ranking a value against others |
| Skewness | pd.Series(data).skew() |
Checking distribution asymmetry |
| Full summary | df.describe() |
First step on any new dataset |
After finishing the notebook, try these:
- Load the Titanic dataset and run
.describe()onAgeandFare - Compare mean vs median of
Fare— which is more representative? Why? - Find outliers in
Ageusing the 1.5×IQR rule - Plot a violin plot comparing
Farefor survivors vs non-survivors - Check skewness of
Age— does it match what you see in the histogram?
| Topic | What you'll learn |
|---|---|
| Distributions | Normal, binomial, Poisson — shapes of data |
| Hypothesis Testing | Is a result real or just random chance? |
| Correlation | How strongly do two variables move together? |
| Regression | Predict one variable from another |
Found a bug or want to add a section?
- Fork the repository
- Create a branch:
git checkout -b feature/add-kurtosis-section - Commit your changes
- Open a Pull Request
Open source under the MIT License.