Folioevanzhou.org
← 返回 Tessera
R2 张图2026-04-13

percentage_stacked

Every time point normalised to a full 100%, so the bands show only how the composition shifts.

示例数据
population_age_composition
配色
babel
语言
R
成图预览另有 1 张在配方里
percentage_stacked

Introduction

"总人口涨了"和"人口结构老了"是两件事,而且常常同时发生。百分比堆叠图把总量这一维主动扔掉,每个时间点都归一化成 100%,只留下构成。

这是它的全部价值,也是它最大的限制:看不出总量。所以它几乎总是要和一张总量图配着用,单独出现容易让人把"占比下降"读成"数量下降"。

读的时候按这个顺序:

  1. 高度不代表总量。 每一年都是满格 100%,图形的总高度是常数。
  2. 看色带厚度。 越厚占比越高。
  3. 沿时间追一条色带。 变厚是占比上升,变薄是下降。
  4. 注意中间的带更难读。 底部色带只有上边界在动,中间的上下边界都在动,眼睛很难判断它到底厚了没有。
  5. 回到绝对值。 占比降了,数量可能反而涨了。

Example Data

这张图要的输入是:时间 + 分类 + 数值,三列长表,一行是某个时间点某一类的数值。数值必须非负、可相加——归一化就是把同一时间点的各类除以它们的和。

population_age_composition 是 1950–2100 年 × 8 个年龄组的人口数。换成自己的数据,只要还是这三列就能直接跑。

library(dplyr)
library(ggplot2)
library(scales)
library(biopalette)

age_levels <- c("<5", "5-14", "15-24", "25-34", "35-44", "45-54", "55-64", ">65")
# 堆叠自下而上,所以要反过来:最年轻的一组落在底部
stack_levels <- rev(age_levels)

# 公开地址,和数据集页上「下载 CSV」给的是同一个 —— 不写仓库相对路径:
# 那个目录不进仓库,读者 clone 下来也没有这个文件,这段代码就跑不了
age_data <- read.csv("https://assets.evanzhou.org/tessera/csv/population_age_composition.csv") |>
  mutate(Group = factor(Group, levels = stack_levels, ordered = TRUE))

# 每个时间点必须恰好 8 行。缺行是这张图唯一会静默变形的地方:
# 归一化按定义就凑满 100%,所以"每年加起来等于 1"这个检查永远通过、什么也查不出,
# 而某年少一组时,剩下 7 条带照样填满 100%,图上看就是一条色带凭空消失又出现。
# 真写成 Value = NA 反而好办 —— sum() 会让那一整年变成 NA,直接开天窗,一眼就看见
rows_per_time <- age_data |> count(Time, name = "groups")
stopifnot(
  all(rows_per_time$groups == length(age_levels)),
  !anyNA(age_data[c("Time", "Group", "Value")]),
  all(age_data$Value >= 0)
)

Palettes

八个年龄组用 babel 的前 8 色。年龄顺序已经由堆叠位置稳定表达了,颜色只负责维持各条色带的身份。

babel 是定性色板,不把冷暖强行解释成年龄高低——如果用渐变色,读者会以为颜色深浅本身在编码什么。

两条顺序必须一致:图例顺序要跟堆叠顺序对齐,否则每看一次图例就要在脑子里翻一次;而最要紧的那一组应该放在底部,因为底部色带只有一条边界在动,是全图最好读的位置。

age_colors <- setNames(
  get_palette("babel", type = "qualitative")[seq_along(stack_levels)],
  stack_levels
)

Recipe

No. Method Input Data Palettes
1 ggplot2 age_percent age_colors
2 ggstream age_percent age_colors

1 · ggplot2

比例要自己先算一次。geom_area() 不会归一化,喂给它什么就画什么。

#| fig: area
#| fig-width: 10
#| fig-height: 6
# 按时间分组算占比 —— 这一步是"百分比堆叠"里的"百分比",图层不负责
age_percent <- age_data |>
  group_by(Time) |>
  mutate(Percent = Value / sum(Value)) |>
  ungroup()

ggplot(age_percent, aes(x = Time, y = Percent, fill = Group)) +
  geom_area(color = "white", linewidth = 0.1) +   # 白色细边把相邻色带断开
  scale_y_continuous(
    labels = percent_format(accuracy = 1),
    expand = c(0, 0)      # 去掉默认留白,让 0% 和 100% 贴住绘图区边界
  ) +
  scale_x_continuous(breaks = seq(1950, 2100, 25), expand = c(0, 0)) +
  scale_fill_manual(values = age_colors) +
  labs(
    title = "Population Age Composition",
    subtitle = "Tessera Toy: population_age_composition · percentage stacked area",
    x = "Year", y = "Proportion", fill = "Age Group"
  ) +
  theme_minimal(base_size = 13) +
  theme(
    legend.position = "bottom",
    panel.grid.minor = element_blank(),
    plot.title.position = "plot",   # 标题左对齐到整张图,不是对齐到绘图区
    plot.title = element_text(face = "bold", hjust = 0, size = 16)
  ) +
  guides(fill = guide_legend(nrow = 1))   # 图例排一行,和堆叠顺序对得上
percentage_stacked — area
percentage_stacked-area

2 · ggstream

同一份 age_percent,改用 type = "mirror" 把每个时间点的 100% 总厚度对称放到零线上下。

它的基线是浮动的——那正是它好看的原因,也是它读不出精确比例的原因:图上任何一条色带的厚度都得靠眼睛估,中间那几条尤其难估。所以它适合汇报开头那张图,不适合当结论图。

#| fig: stream
#| fig-width: 10
#| fig-height: 7
library(ggstream)

ggplot(age_percent, aes(x = Time, y = Percent, fill = Group)) +
  # Percent 已经归一化,mirror 只负责把总厚度居中。
  # type = "proportional" 也归一化,但仍从底部基线起,不是这里要的居中形态
  geom_stream(type = "mirror") +
  scale_y_continuous(
    breaks = seq(-0.5, 0.5, 0.25),
    labels = percent_format(accuracy = 1),
    expand = expansion(mult = 0.04)
  ) +
  scale_x_continuous(
    breaks = seq(1950, 2100, 25),
    expand = expansion(mult = 0.01)
  ) +
  scale_fill_manual(values = age_colors) +
  labs(
    title = "Population Age Composition",
    subtitle = "Tessera Toy: population_age_composition · centered yearly shares",
    x = "Year", y = "Centered share", fill = "Age Group"
  ) +
  guides(fill = guide_legend(nrow = 1)) +
  theme_minimal(base_size = 13) +
  theme(
    legend.position = "bottom",
    panel.grid.minor = element_blank(),
    plot.title.position = "plot",
    plot.title = element_text(face = "bold", hjust = 0, size = 16)
  )
percentage_stacked — stream
percentage_stacked-stream

Constraints

  • 占比下降不等于数量下降。 总量那一维已经被扔掉了,图上没有任何东西能反映它。这句话值得直接写在图注里,并配一张总量图。
  • 中间的色带最难读。 底部只有上边界在动,中间的上下边界同时在动,厚度变化基本靠猜。这是堆叠图的固有缺陷,不是画错了——所以最要紧的一组要放底部。
  • 类别多了就追不动。 超过七八条色带,读者没法跟住任何一条。

两个 recipe 怎么比

  1. 要让人读出数字 → recipe 1。固定的 0–100% 基线是能对齐的参照,stream 那条浮动基线不是。
  2. 要表达"结构在流动" → recipe 2。对称扩散的形态更能传达长期演变,适合做汇报开头那张图,但别拿它下结论。
运行环境5 个包 · 2026-08-13T11:58:18.204+0800

R version 4.5.1 (2025-06-13 ucrt) · x86_64-w64-mingw32

  • biopalette 0.1.0
  • dplyr 1.2.1
  • ggplot2 4.0.3
  • ggstream 0.1.0
  • scales 1.4.0