Folioevanzhou.org
← 返回 Tessera
R6 张图2026-08-11

bar

Discrete categories compared by bar length from a shared zero, side by side or stacked into a total.

示例数据
mtcarshair_eye_color
配色
walter_white2
语言
R
成图预览另有 5 张在配方里
bar

Introduction

柱状图比较离散类别对应的数值。类别之间没有连续过程,柱与柱之间的空隙正是在表达这一点;如果横轴是日期、年龄或剂量,通常该先考虑折线或点图。

一根柱最适合回答"谁高、谁低、差多少"。加入第二个分类后有两条路线:并排柱保留共同零基线,适合比较各组绝对值;堆叠柱把各部分加成总量,适合看构成,但只有最底层有共同基线。

它也不是连续数据分布的替代品。每组有很多原始观测时,只画均值柱会藏掉离散程度、样本量和异常值——那种情况该用点图、箱线图、小提琴图或点估计加误差棒。

读的时候按这个顺序:

  1. 先确认零基线。 柱长靠长度编码,截断纵轴会夸大差异。
  2. 比较柱端位置。 单柱和并排柱共享基线,柱端越容易对齐,比较越可靠。
  3. 堆叠图先看总长。 整根柱回答总量,各区块回答组成。
  4. 中间区块只作粗略比较。 它们上下边界都在移动,没有共同起点。
  5. 百分比柱只表达构成。 每根都是 100%,看不出样本量大小。

Example Data

柱状图要的输入有两种形状,对应两个不同的图层——区别不在外观,而在谁负责算柱高:

输入 用什么 柱高来自
一行一个原始观测 geom_bar()(默认 stat = "count") 图层现场数行数
每个类别已有一个汇总值 geom_col()(等价于 stat = "identity") 数据里的 y

这件事错了不会报错。hair_eye_color 的 Freq 已经是频数,直接写 geom_bar() 的话,R 数到的是每种组合有几行(永远是 1),而不是学生人数——图照样生成,含义完全错了。所以下面第一条用 mtcars 演示原始行计数,其余全部用 geom_col()。

频数表也可以写 geom_bar(aes(weight = Freq)),统计层会按类别累加权重。临时探索很方便,但正式配方更倾向先把汇总表算出来再用 geom_col():中间数据能检查、能复用,缺组也更容易发现。

hair_eye_color 是一个完整的 Hair × Eye × Sex × Freq 长表。

library(dplyr)
library(ggplot2)
library(ggpubr)
library(scales)
library(biopalette)

# 公开地址,和数据集页上「下载 CSV」给的是同一个 —— 不写仓库相对路径:
# 那个目录不进仓库,读者 clone 下来也没有这个文件,这段代码就跑不了
d <- read.csv("https://assets.evanzhou.org/tessera/csv/hair_eye_color.csv")
cars <- read.csv("https://assets.evanzhou.org/tessera/csv/mtcars.csv")

hair_levels <- c("Black", "Brown", "Red", "Blond")
eye_levels <- c("Brown", "Blue", "Hazel", "Green")
sex_levels <- c("Male", "Female")

# 4 × 4 × 2 必须齐全。缺一格的话,后面百分比堆叠那张图照样凑满 100%,
# 而且看不出少了什么 —— 所以检查放在归一化之前
if (nrow(d) != length(hair_levels) * length(eye_levels) * length(sex_levels) ||
    any(count(d, Hair, Eye, Sex)$n != 1)) {
  stop("Hair × Eye × Sex 组合不完整")
}

# 显式定 levels:发色和眼色都有熟悉的原始顺序,交给字母序会乱
d <- d |>
  mutate(
    Hair = factor(Hair, levels = hair_levels),
    Eye = factor(Eye, levels = eye_levels),
    Sex = factor(Sex, levels = sex_levels)
  )

# 三张汇总表,各对应一种柱图
hair_total <- d |>
  summarise(Freq = sum(Freq), .by = Hair)

hair_sex <- d |>
  summarise(Freq = sum(Freq), .by = c(Hair, Sex))

hair_eye <- d |>
  summarise(Freq = sum(Freq), .by = c(Hair, Eye))

Palettes

柱子的 fill 表达类别,所以用定性色板 walter_white2。单序列只要一个颜色;并排或堆叠时,每个水平固定一个颜色,图例顺序必须和分组或堆叠顺序一致——不一致的话,每查一次颜色都要在脑子里翻转一遍。

相邻色块比相隔的散点更考验颜色边界。堆叠柱用细白线分隔区块,但白线只负责断开边界,补救不了本身就难区分的配色。

色板少于类别数时不应循环颜色:减少重点类别、合并有业务意义的类别,或者明确把其余类别当作中性上下文。

category_colors <- setNames(
  get_palette("walter_white2", type = "qualitative")[1:4],
  eye_levels
)

# 单序列那三张只要一个色,取色板第一个
SOLO <- unname(category_colors[1])

bar_theme <- theme_pubr(base_size = 13, legend = "right") +
  theme(
    # 只留水平网格线:柱端要和 y 轴刻度对齐,竖线在这里没有用处
    panel.grid.major.y = element_line(color = "#E5E3DC", linewidth = 0.35),
    panel.grid.major.x = element_blank(),
    axis.line = element_line(color = "#333330", linewidth = 0.45),
    axis.ticks = element_line(color = "#333330", linewidth = 0.4),
    plot.title = element_text(face = "bold", size = 15, hjust = 0),
    plot.subtitle = element_text(color = "grey35", hjust = 0),
    legend.title = element_text(face = "bold"),
    plot.margin = margin(14, 18, 12, 12)
  )

Recipe

No. Method Input Data Palettes
1 ggplot2 cars SOLO
2 ggplot2 hair_total SOLO
3 ggplot2 hair_total SOLO
4 ggplot2 hair_sex category_colors[1:2]
5 ggplot2 hair_eye category_colors
6 ggplot2 hair_percent category_colors

1 · 从原始观测自动计数

mtcars 一行是一款汽车,所以 geom_bar() 可以直接统计 4、6、8 缸各有多少款。

这里没有 y:柱高来自绘图层算出的 count,标签也必须通过 after_stat(count) 读同一个统计结果。

#| fig: count
#| fig-width: 7.6
#| fig-height: 5.1
cars$cyl <- factor(cars$cyl, levels = c(4, 6, 8))

ggplot(cars, aes(cyl)) +
  geom_bar(width = 0.68, fill = SOLO) +
  geom_text(
    stat = "count",                    # 和柱用同一个统计层
    aes(label = after_stat(count)),    # 直接写 count 会找不到这一列
    vjust = -0.55,
    size = 3.7,
    fontface = "bold"
  ) +
  scale_y_continuous(
    limits = c(0, NA),                        # 从零起,柱长才是长度编码
    breaks = seq(0, 15, 5),
    expand = expansion(mult = c(0, 0.1))      # 底部不留白、顶部留给数值标签
  ) +
  labs(
    title = "Cars by cylinder count",
    subtitle = "geom_bar() counts one-row-per-car observations",
    x = "Cylinders",
    y = "Cars"
  ) +
  bar_theme
bar — count
bar-count

2 · 从汇总值直接画

四种发色有明确、熟悉的原始顺序,这里不按人数擅自重排。geom_col() 直接读汇总后的 Freq。

#| fig: basic
#| fig-width: 7.6
#| fig-height: 5.1
ggplot(hair_total, aes(Hair, Freq)) +
  geom_col(width = 0.68, fill = SOLO) +
  geom_text(aes(label = Freq), vjust = -0.55, size = 3.7, fontface = "bold") +
  scale_y_continuous(
    limits = c(0, NA),
    expand = expansion(mult = c(0, 0.09))
  ) +
  labs(
    title = "Students by hair color",
    subtitle = "Original category order retained",
    x = "Hair color",
    y = "Students"
  ) +
  bar_theme
bar — basic
bar-basic

3 · 排序与横向

只有当问题明确变成"哪种发色人数最多"时,排序才有意义。类别名很长时横向也更好读。

#| fig: horizontal
#| fig-width: 7.6
#| fig-height: 5.1
# reorder() 按 Freq 重排因子水平。名义类别可以这么做,
# 时间、剂量、流程阶段这类有自然顺序的不行
ranked <- hair_total |>
  mutate(Hair = reorder(Hair, Freq))

ggplot(ranked, aes(Hair, Freq)) +
  geom_col(width = 0.64, fill = SOLO) +
  geom_text(aes(label = Freq), hjust = -0.35, size = 3.7, fontface = "bold") +
  scale_y_continuous(
    limits = c(0, NA),
    expand = expansion(mult = c(0, 0.1))
  ) +
  # clip = "off":翻转后标签落在绘图区外,不关掉裁剪就看不见
  coord_flip(clip = "off") +
  labs(
    title = "Students by hair color",
    subtitle = "Sorted only because the question is a ranking",
    x = NULL,
    y = "Students"
  ) +
  bar_theme
bar — horizontal
bar-horizontal

4 · 并排柱

Male 与 Female 只有两个水平,并排后仍能快速配对。每组四五根柱以后读者既找不到配对也读不动图例,那时该改用分面、点图或热图。

#| fig: grouped
#| fig-width: 8
#| fig-height: 5.2
# 同一个 position 对象同时给柱和标签 —— 分开写两个 dodge,
# 文字会落在另一根柱上
dodge <- position_dodge(width = 0.74)
sex_colors <- setNames(category_colors[1:2], sex_levels)

ggplot(hair_sex, aes(Hair, Freq, fill = Sex)) +
  geom_col(width = 0.66, position = dodge) +
  geom_text(
    aes(label = Freq),
    position = dodge,
    vjust = -0.5,
    size = 3.25,
    fontface = "bold"
  ) +
  scale_y_continuous(
    limits = c(0, NA),
    expand = expansion(mult = c(0, 0.11))
  ) +
  scale_fill_manual(values = sex_colors, breaks = sex_levels) +
  labs(
    title = "Hair color counts by sex",
    subtitle = "Side-by-side bars preserve a shared zero baseline",
    x = "Hair color",
    y = "Students",
    fill = "Sex"
  ) +
  bar_theme
bar — grouped
bar-grouped

5 · 堆叠柱

同时保留每种眼色的频数和每种发色的总人数。

数字只写在柱顶的总量上,不往每段里塞——中间段本来就没有共同起点,标了也不该拿来精确比较。

#| fig: stacked
#| fig-width: 8
#| fig-height: 5.2
ggplot(hair_eye, aes(Hair, Freq, fill = Eye)) +
  geom_col(
    width = 0.68,
    color = "white",       # 细白线断开相邻区块
    linewidth = 0.6,
    position = position_stack(reverse = TRUE)   # 堆叠顺序和图例顺序对齐
  ) +
  # 总量单独一层:数据换成 hair_total,且不继承 fill 映射
  geom_text(
    data = hair_total,
    aes(Hair, Freq, label = Freq),
    inherit.aes = FALSE,
    vjust = -0.55,
    size = 3.5,
    fontface = "bold"
  ) +
  scale_y_continuous(
    limits = c(0, NA),
    expand = expansion(mult = c(0, 0.09))
  ) +
  scale_fill_manual(values = category_colors, breaks = eye_levels) +
  labs(
    title = "Eye color counts",
    subtitle = "Stacked within hair color · totals shown above",
    x = "Hair color",
    y = "Students",
    fill = "Eye color"
  ) +
  bar_theme
bar — stacked
bar-stacked

6 · 100% 堆叠柱

主动放弃总量,只比较构成。

比例先在每种发色内部算好,不让绘图层隐式归一化——这样 Percent 能检查、能复用,也能直接用于标签。

#| fig: percentage
#| fig-width: 8
#| fig-height: 5.2
hair_percent <- hair_eye |>
  group_by(Hair) |>
  mutate(Percent = Freq / sum(Freq)) |>
  ungroup() |>
  # 翻转 levels:coord_flip() 之后,第一个水平会落到最下面
  mutate(Hair = factor(Hair, levels = rev(hair_levels)))

ggplot(hair_percent, aes(Hair, Percent, fill = Eye)) +
  geom_col(
    width = 0.68,
    color = "white",
    linewidth = 0.65,
    position = position_stack(reverse = TRUE)
  ) +
  geom_text(
    # 小于 8% 的区块放不下文字就不标 —— 挤成一团的数字比没有数字更糟
    aes(label = ifelse(Percent >= 0.08, percent(Percent, accuracy = 1), "")),
    position = position_stack(vjust = 0.5, reverse = TRUE),
    color = "#262622",
    size = 3.5,
    fontface = "bold"
  ) +
  scale_y_continuous(
    breaks = seq(0, 1, 0.25),
    labels = percent_format(accuracy = 1),
    expand = c(0, 0)
  ) +
  scale_fill_manual(values = category_colors, breaks = eye_levels) +
  coord_flip(clip = "off") +
  labs(
    title = "Eye color composition by hair color",
    subtitle = "Sexes combined · each bar sums to 100%",
    x = NULL,
    y = "Composition",
    fill = "Eye color"
  ) +
  bar_theme +
  theme(
    panel.grid.major.x = element_line(color = "#E5E3DC", linewidth = 0.35),
    panel.grid.major.y = element_blank(),
    axis.line.y = element_blank()
  )
bar — percentage
bar-percentage

Constraints

  • 柱必须从零开始。 柱长是长度编码,截断纵轴就没有共同起点,视觉差异会被人为放大。
  • 堆叠柱只有最底层有共同基线。 中间区块的上下边界都在移动,只能作粗略比较。
  • 百分比柱看不出样本量。 每根都是 100%,10 个人和 1000 个人长得一样。更麻烦的是它会掩盖缺组:少一类之后,剩下的部分仍会归一化成 100%,图上完全看不出来。
  • 均值柱藏掉了分布。 没有误差、样本量和形状的均值柱,往往比不画更容易误导。
  • 有自然顺序的类别不能为了好看重排。 时间、剂量、流程阶段和有等级的分类,顺序本身就是信息。

六个 recipe 怎么比

  1. 一行一个原始观测,要现场数类别 → recipe 1。
  2. 每类已有一个值,保留自然顺序 → recipe 2。
  3. 问题就是排名,或类别名很长 → recipe 3。
  4. 比较少数分组的绝对值 → recipe 4。
  5. 同时看部分与总量 → recipe 5。
  6. 只看构成比例 → recipe 6,同时另给总量,避免把比例变化误读成数量变化。
运行环境5 个包 · 2026-08-11T17:13:52.892+0800

R version 4.5.1 (2025-06-13 ucrt) · x86_64-w64-mingw32

  • biopalette 0.1.0
  • dplyr 1.2.1
  • ggplot2 4.0.3
  • ggpubr 0.6.2
  • scales 1.4.0