Before visualising, let’s look at the “Bird’s Eye View” of our data. The describe() function gives us the range, mean, and standard deviation for all numeric columns.
# Get a statistical summarystats = df_clean.describe()display(stats)
Table 1: Descriptive Statistics for Numeric Variables
bill_length_mm
bill_depth_mm
flipper_length_mm
body_mass_g
count
333.00
333.00
333.00
333.00
mean
43.99
17.16
200.97
4,207.06
std
5.47
1.97
14.02
805.22
min
32.10
13.10
172.00
2,700.00
25%
39.50
15.60
190.00
3,550.00
50%
44.50
17.30
197.00
4,050.00
75%
48.60
18.70
213.00
4,775.00
max
59.60
21.50
231.00
6,300.00
Grouping by Species
The statistics above are for all penguins mixed together. Let’s group them by species to see if there are distinct size differences.
Question: Which species has the highest average body mass?
# Corrected: Group by species, select columns, THEN calculate meanspecies_means = df_clean.groupby("species")[ ["bill_length_mm", "bill_depth_mm", "flipper_length_mm", "body_mass_g"]].mean()display(species_means)