Introduction
柱状图比较离散类别对应的数值。类别之间没有连续过程,柱与柱之间的空隙正是在表达这一点;如果横轴是日期、年龄或剂量,通常该先考虑折线或点图。
一根柱最适合回答"谁高、谁低、差多少"。加入第二个分类后有两条路线:并排柱保留共同零基线,适合比较各组绝对值;堆叠柱把各部分加成总量,适合看构成,但只有最底层有共同基线。
它也不是连续数据分布的替代品。每组有很多原始观测时,只画均值柱会藏掉离散程度、样本量和异常值——那种情况该用点图、箱线图、小提琴图或点估计加误差棒。
读的时候按这个顺序:
- 先确认零基线。 柱长靠长度编码,截断纵轴会夸大差异。
- 比较柱端位置。 单柱和并排柱共享基线,柱端越容易对齐,比较越可靠。
- 堆叠图先看总长。 整根柱回答总量,各区块回答组成。
- 中间区块只作粗略比较。 它们上下边界都在移动,没有共同起点。
- 百分比柱只表达构成。 每根都是 100%,看不出样本量大小。
Example Data
柱状图要的输入有两种形状,对应两个不同的图层——区别不在外观,而在谁负责算柱高:
| 输入 | 用什么 | 柱高来自 |
|---|---|---|
| 一行一个原始观测 | geom_bar()(默认 stat = "count") |
图层现场数行数 |
| 每个类别已有一个汇总值 | geom_col()(等价于 stat = "identity") |
数据里的 y |
这件事错了不会报错。hair_eye_color 的 Freq 已经是频数,直接写 geom_bar() 的话,R 数到的是每种组合有几行(永远是 1),而不是学生人数——图照样生成,含义完全错了。所以下面第一条用 mtcars 演示原始行计数,其余全部用 geom_col()。
频数表也可以写 geom_bar(aes(weight = Freq)),统计层会按类别累加权重。临时探索很方便,但正式配方更倾向先把汇总表算出来再用 geom_col():中间数据能检查、能复用,缺组也更容易发现。
hair_eye_color 是一个完整的 Hair × Eye × Sex × Freq 长表。
library(dplyr)
library(ggplot2)
library(ggpubr)
library(scales)
library(biopalette)
# 公开地址,和数据集页上「下载 CSV」给的是同一个 —— 不写仓库相对路径:
# 那个目录不进仓库,读者 clone 下来也没有这个文件,这段代码就跑不了
d <- read.csv("https://assets.evanzhou.org/tessera/csv/hair_eye_color.csv")
cars <- read.csv("https://assets.evanzhou.org/tessera/csv/mtcars.csv")
hair_levels <- c("Black", "Brown", "Red", "Blond")
eye_levels <- c("Brown", "Blue", "Hazel", "Green")
sex_levels <- c("Male", "Female")
# 4 × 4 × 2 必须齐全。缺一格的话,后面百分比堆叠那张图照样凑满 100%,
# 而且看不出少了什么 —— 所以检查放在归一化之前
if (nrow(d) != length(hair_levels) * length(eye_levels) * length(sex_levels) ||
any(count(d, Hair, Eye, Sex)$n != 1)) {
stop("Hair × Eye × Sex 组合不完整")
}
# 显式定 levels:发色和眼色都有熟悉的原始顺序,交给字母序会乱
d <- d |>
mutate(
Hair = factor(Hair, levels = hair_levels),
Eye = factor(Eye, levels = eye_levels),
Sex = factor(Sex, levels = sex_levels)
)
# 三张汇总表,各对应一种柱图
hair_total <- d |>
summarise(Freq = sum(Freq), .by = Hair)
hair_sex <- d |>
summarise(Freq = sum(Freq), .by = c(Hair, Sex))
hair_eye <- d |>
summarise(Freq = sum(Freq), .by = c(Hair, Eye))
Palettes
柱子的 fill 表达类别,所以用定性色板 walter_white2。单序列只要一个颜色;并排或堆叠时,每个水平固定一个颜色,图例顺序必须和分组或堆叠顺序一致——不一致的话,每查一次颜色都要在脑子里翻转一遍。
相邻色块比相隔的散点更考验颜色边界。堆叠柱用细白线分隔区块,但白线只负责断开边界,补救不了本身就难区分的配色。
色板少于类别数时不应循环颜色:减少重点类别、合并有业务意义的类别,或者明确把其余类别当作中性上下文。
category_colors <- setNames(
get_palette("walter_white2", type = "qualitative")[1:4],
eye_levels
)
# 单序列那三张只要一个色,取色板第一个
SOLO <- unname(category_colors[1])
bar_theme <- theme_pubr(base_size = 13, legend = "right") +
theme(
# 只留水平网格线:柱端要和 y 轴刻度对齐,竖线在这里没有用处
panel.grid.major.y = element_line(color = "#E5E3DC", linewidth = 0.35),
panel.grid.major.x = element_blank(),
axis.line = element_line(color = "#333330", linewidth = 0.45),
axis.ticks = element_line(color = "#333330", linewidth = 0.4),
plot.title = element_text(face = "bold", size = 15, hjust = 0),
plot.subtitle = element_text(color = "grey35", hjust = 0),
legend.title = element_text(face = "bold"),
plot.margin = margin(14, 18, 12, 12)
)
Recipe
| No. | Method | Input Data | Palettes |
|---|---|---|---|
| 1 | ggplot2 |
cars |
SOLO |
| 2 | ggplot2 |
hair_total |
SOLO |
| 3 | ggplot2 |
hair_total |
SOLO |
| 4 | ggplot2 |
hair_sex |
category_colors[1:2] |
| 5 | ggplot2 |
hair_eye |
category_colors |
| 6 | ggplot2 |
hair_percent |
category_colors |
1 · 从原始观测自动计数
mtcars 一行是一款汽车,所以 geom_bar() 可以直接统计 4、6、8 缸各有多少款。
这里没有 y:柱高来自绘图层算出的 count,标签也必须通过 after_stat(count) 读同一个统计结果。
#| fig: count
#| fig-width: 7.6
#| fig-height: 5.1
cars$cyl <- factor(cars$cyl, levels = c(4, 6, 8))
ggplot(cars, aes(cyl)) +
geom_bar(width = 0.68, fill = SOLO) +
geom_text(
stat = "count", # 和柱用同一个统计层
aes(label = after_stat(count)), # 直接写 count 会找不到这一列
vjust = -0.55,
size = 3.7,
fontface = "bold"
) +
scale_y_continuous(
limits = c(0, NA), # 从零起,柱长才是长度编码
breaks = seq(0, 15, 5),
expand = expansion(mult = c(0, 0.1)) # 底部不留白、顶部留给数值标签
) +
labs(
title = "Cars by cylinder count",
subtitle = "geom_bar() counts one-row-per-car observations",
x = "Cylinders",
y = "Cars"
) +
bar_theme




