Folioevanzhou.org
← 返回 Tessera
R4 张图2026-08-11

box

Five summary marks per group - median, quartiles, whiskers and the points beyond them.

示例数据
plant_growth
配色
walter_white2
语言
R
成图预览另有 3 张在配方里
box

Introduction

箱线图用五个视觉部件压缩一组连续数据:中位数、上下四分位数、两条须,以及落在须外的点。它适合并排比较多个组的位置与离散程度,比均值柱保留的信息多,也比小提琴图更少依赖分布估计。

关心双峰、偏态尾部或完整密度形状时用小提琴图或密度图;只展示一个估计值和不确定性时用点估计加误差棒。箱线图不是所有"组间比较"的默认答案。

读的时候按这个顺序:

  1. 箱体下边缘是 Q1,上边缘是 Q3。 中间 50% 的数据落在箱体内。
  2. 箱体高度是 IQR = Q3 − Q1。 箱体越高,中间一半数据越分散。
  3. 箱内横线是中位数,不是均值。 偏态数据里两者可以相差很远。
  4. 须不是最小值和最大值。 下须延伸到不低于 Q1 − 1.5 × IQR 的最小观测,上须延伸到不高于 Q3 + 1.5 × IQR 的最大观测——围栏只是判定界限,须端落在真实观测上,不一定正好等于围栏位置。
  5. 须外的点是规则标出的潜在异常值。 它们不一定错误,也不能因为被标出就删掉。

Example Data

箱线图要的输入是:1 个分类变量 + 1 个连续数值变量,一行一个观测。不需要事先汇总——四分位和须都是绘图层现算的。

plant_growth 有一个对照组和两个处理组,每组恰好十株植物。每组只有十来个观测时,几个分位数会被单个点明显影响,所以这一页的主版本会把全部原始点叠上去。

library(ggplot2)
library(ggpubr)
library(biopalette)

# 公开地址,和数据集页上「下载 CSV」给的是同一个 —— 不写仓库相对路径:
# 那个目录不进仓库,读者 clone 下来也没有这个文件,这段代码就跑不了
d <- read.csv("https://assets.evanzhou.org/tessera/csv/plant_growth.csv")

# 显式定 levels:ctrl 必须排在两个处理组前面,交给字母序也对,但换套组名就不对了
group_levels <- c("ctrl", "trt1", "trt2")
d$group <- factor(d$group, levels = group_levels)

# 样本量写进 x 轴标签 —— 同样大小的箱体可能来自 10 个点或 1000 个点,
# 而箱线图本身不透露这件事
group_n <- table(d$group)
group_labels <- setNames(
  paste0(group_levels, "\n(n=", as.integer(group_n[group_levels]), ")"),
  group_levels
)

Palettes

处理组是无序类别,用定性色板 walter_white2。箱体用半透明填色配统一深色边框,原始点沿用同一组颜色和同一种圆形——颜色只解释组别一次。

箱线图里的色块面积比散点大,所以浅色很容易显得发白,深色又可能压住中位数线。保留深色轮廓加适度透明度,比单纯加深所有颜色更稳。

常规分组箱线图只用颜色。再给每组配一种点形,是增加一份没有新信息的图例负担。色板不足组数时不循环颜色,也不删数据:未编码的组用明确的中性色。

group_colors <- setNames(
  get_palette("walter_white2", type = "qualitative")[1:3],
  group_levels
)

# 四张图共用同一条 y 轴:箱体高低要横着比,坐标就不能各画各的
box_scale <- scale_y_continuous(
  breaks = seq(3.5, 6.5, 1),
  expand = expansion(mult = c(0.04, 0.06))
)

box_theme <- theme_pubr(base_size = 13, legend = "none") +
  theme(
    panel.grid.major.y = element_line(color = "#E5E3DC", linewidth = 0.35),
    panel.grid.major.x = element_blank(),
    axis.line = element_line(color = "#333330", linewidth = 0.45),
    axis.ticks = element_line(color = "#333330", linewidth = 0.4),
    plot.title = element_text(face = "bold", size = 15, hjust = 0),
    plot.subtitle = element_text(color = "grey35", hjust = 0),
    plot.margin = margin(14, 16, 12, 12)
  )

Recipe

No. Method Input Data Palettes
1 ggplot2 d —
2 ggplot2 d group_colors
3 ggplot2 d group_colors
4 ggplot2 d group_colors

1 · 基础分组箱线图

先只看统计结构。三个组已经由横轴位置区分,不需要图例,也不需要颜色。

#| fig: basic
#| fig-width: 7.5
#| fig-height: 5.2
ggplot(d, aes(group, weight)) +
  geom_boxplot(
    width = 0.56,
    fill = NA,          # 完全不填色,只留轮廓、中位数和须
    color = "#333330",
    linewidth = 0.7
  ) +
  scale_x_discrete(labels = group_labels) +   # 带 n 的标签
  box_scale +
  labs(
    title = "Plant weight by treatment",
    subtitle = "Outline only · no box fill",
    x = NULL,
    y = "Dried weight"
  ) +
  box_theme
box — basic
box-basic

2 · 箱体按组着色

处理组的颜色需要和同一份报告里的其它图保持一致时,把 fill 映射到组别。

边框、中位数和须仍统一用深色——三套视觉规则同时变化,读者要解的东西就多了一倍。

#| fig: color
#| fig-width: 7.5
#| fig-height: 5.2
ggplot(d, aes(group, weight, fill = group)) +
  geom_boxplot(
    width = 0.56,
    alpha = 0.72,       # 半透明:填色不能压过中位数线
    color = "#333330",
    linewidth = 0.7
  ) +
  scale_x_discrete(labels = group_labels) +
  scale_fill_manual(values = group_colors) +
  box_scale +
  labs(
    title = "Plant weight by treatment",
    subtitle = "One categorical color per treatment group",
    x = NULL,
    y = "Dried weight"
  ) +
  box_theme
box — color
box-color

3 · 叠加全部原始点

这是小样本时优先使用的版本:箱体给摘要,点给真实数据。

#| fig: points
#| fig-width: 7.5
#| fig-height: 5.2
ggplot(d, aes(group, weight, fill = group)) +
  geom_boxplot(
    width = 0.54,
    alpha = 0.58,
    color = "#333330",
    linewidth = 0.7,
    outlier.shape = NA   # 关掉默认的须外点,否则下一层会把它们再画一遍
  ) +
  geom_point(
    shape = 21,
    size = 2.8,
    stroke = 0.65,
    color = "#333330",
    # height = 0 只横向抖动:纵向抖动会篡改真实数值。
    # seed 绑在 position 上,重复导出时点的位置不会变
    position = position_jitter(width = 0.1, height = 0, seed = 2026)
  ) +
  scale_x_discrete(labels = group_labels) +
  scale_fill_manual(values = group_colors) +
  box_scale +
  labs(
    title = "Plant weight by treatment",
    subtitle = "All 30 observations retained",
    x = NULL,
    y = "Dried weight"
  ) +
  box_theme
box — points
box-points

4 · 箱体宽度与点大小

width 只改变箱体在分类轴上占多少空间,不改变 Q1、Q3、中位数或须;size 只改变原始点的视觉重量,不改变数据。两者都该按最终输出尺寸判断,不在大预览窗里凭感觉设。

窄箱体适合类别多、版面紧的图;宽箱体更有块面感,但容易挤满分类间距。小点让箱体摘要占主导,大点更强调每个真实观测。

#| fig: sizing
#| fig-width: 10
#| fig-height: 7.6
comparison_theme <- theme_pubr(base_size = 11, legend = "none") +
  theme(
    panel.grid.major.y = element_line(color = "#E5E3DC", linewidth = 0.3),
    panel.grid.major.x = element_blank(),
    axis.line = element_line(color = "#333330", linewidth = 0.4),
    plot.title = element_text(face = "bold", size = 12, hjust = 0),
    plot.margin = margin(8, 8, 8, 8)
  )

# 包成函数而不是复制四遍:除了要对比的那一个参数,四格必须完全相同
box_only <- function(box_width, title) {
  ggplot(d, aes(group, weight, color = group)) +
    geom_boxplot(
      width = box_width,
      fill = NA,          # 空心彩色箱体:轮廓和原始点共用组别色
      linewidth = 0.8,
      outlier.shape = NA
    ) +
    geom_point(
      size = 2.4,
      alpha = 0.82,
      position = position_jitter(width = 0.1, height = 0, seed = 2026)
    ) +
    scale_x_discrete(labels = group_levels) +
    scale_color_manual(values = group_colors) +
    box_scale +
    labs(title = title, x = NULL, y = "Dried weight") +
    comparison_theme
}

box_with_points <- function(point_size, title) {
  ggplot(d, aes(group, weight, fill = group)) +
    geom_boxplot(
      width = 0.54,
      alpha = 0.58,
      color = "#333330",
      linewidth = 0.7,
      outlier.shape = NA
    ) +
    geom_point(
      shape = 21,
      size = point_size,
      stroke = 0.65,
      color = "#333330",
      position = position_jitter(width = 0.1, height = 0, seed = 2026)
    ) +
    scale_x_discrete(labels = group_levels) +
    scale_fill_manual(values = group_colors) +
    box_scale +
    labs(title = title, x = NULL, y = "Dried weight") +
    comparison_theme
}

ggarrange(
  box_only(0.34, "Narrow boxes · width = 0.34"),
  box_only(0.76, "Wide boxes · width = 0.76"),
  box_with_points(1.7, "Small points · size = 1.7"),
  box_with_points(3.8, "Large points · size = 3.8"),
  ncol = 2,
  nrow = 2,
  align = "hv"    # 四格的绘图区严格对齐,y 轴才可比
)
box — sizing
box-sizing

Constraints

  • 须不是最小值和最大值。 ggplot2 默认用 1.5 × IQR 规则,须外仍可能有真实观测。
  • 样本太少只画箱体,会制造过度稳定的印象。 n = 5 也能算出四分位,一个整齐的箱子什么都看不出来。叠上原始点。
  • 同样大小的箱体可能来自 10 个点或 1000 个点。 箱线图不透露样本量,n 必须写进轴标签或图注。
  • 对称箱体不等于正态分布。 箱线图不足以判断分布假设。
  • 别用 scale_y_continuous(limits = ) 裁范围。 它会在统计之前删掉范围外的数据,四分位和须都会跟着变。只想放大视图用 coord_cartesian()。

四个 recipe 怎么比

  1. 只需要紧凑统计摘要 → recipe 1。
  2. 组别颜色需要跨图保持一致 → recipe 2。
  3. 样本量小、读者需要看到真实数据 → recipe 3。这是最诚实、也最适合正式使用的版本。
  4. 需要适配不同版面尺寸 → recipe 4。先定最终宽度,再调箱体 width 和散点 size。
运行环境3 个包 · 2026-08-11T17:37:58.343+0800

R version 4.5.1 (2025-06-13 ucrt) · x86_64-w64-mingw32

  • biopalette 0.1.0
  • ggplot2 4.0.3
  • ggpubr 0.6.2