Descriptive Statistics

Before visualising, let’s look at the “Bird’s Eye View” of our data. The describe() function gives us the range, mean, and standard deviation for all numeric columns.

# Get a statistical summary
stats = df_clean.describe()
display(stats)
Table 1: Descriptive Statistics for Numeric Variables
bill_length_mm bill_depth_mm flipper_length_mm body_mass_g
count 333.00 333.00 333.00 333.00
mean 43.99 17.16 200.97 4,207.06
std 5.47 1.97 14.02 805.22
min 32.10 13.10 172.00 2,700.00
25% 39.50 15.60 190.00 3,550.00
50% 44.50 17.30 197.00 4,050.00
75% 48.60 18.70 213.00 4,775.00
max 59.60 21.50 231.00 6,300.00

Grouping by Species

The statistics above are for all penguins mixed together. Let’s group them by species to see if there are distinct size differences.

Question: Which species has the highest average body mass?

# Corrected: Group by species, select columns, THEN calculate mean
species_means = df_clean.groupby("species")[
    ["bill_length_mm", "bill_depth_mm", "flipper_length_mm", "body_mass_g"]
].mean()

display(species_means)
Table 2: Average Measurements by Species
bill_length_mm bill_depth_mm flipper_length_mm body_mass_g
species
Adelie 38.82 18.35 190.10 3,706.16
Chinstrap 48.83 18.42 195.82 3,733.09
Gentoo 47.57 15.00 217.24 5,092.44