Folioevanzhou.org
← 返回 Tessera
R5 张图2026-08-11

density

A continuous variable smoothed into a curve, high where observations pile up and thin along the tails.

示例数据
airquality
配色
walter_white2
语言
R
成图预览另有 4 张在配方里
density

Introduction

密度图把一列连续观测平滑成一条分布曲线。曲线高的地方观测更集中,低而长的尾部表示少量观测延伸得更远。它回答"数据主要落在哪里""是否偏斜""有没有一个以上的峰",不回答某个精确观测值是多少。

它和箱线图互补:箱线图用中位数和四分位数给出稳定摘要,密度图保留完整形状,但形状会受带宽影响。小样本优先画原始点或箱线图;每组至少二三十个观测时,密度轮廓才开始有解释价值。

曲线是怎么算出来的:在每个观测上放一个小的平滑鼓包,再把所有鼓包相加(geom_density() 默认高斯核)。真正决定曲线长相的是带宽——带宽小保留更多局部起伏,也更容易把随机噪声画成多个峰;带宽大更稳定,但相邻的峰可能被抹成一个宽峰。adjust 是默认带宽的倍数,不是另一种统计方法。

读的时候按这个顺序:

  1. 先看位置。 整条曲线向右移动,说明这一组整体取值更高。
  2. 再看宽窄。 分布越宽,观测越分散。
  3. 看偏斜和尾部。 一侧拖得更长,说明极端方向不对称。
  4. 谨慎看多峰。 小鼓包可能只是样本少或带宽过小,不一定存在真实亚群。
  5. 分组图必须共用横轴。 每个面板自行缩放,会让不同月份看起来同样宽,失去比较基础。

Example Data

密度图要的输入是:1 个连续数值变量,可选再加 1 个分组变量。一行一个观测,不需要事先汇总。

airquality 是纽约 1973 年五到九月的逐日观测,153 行。这里用日最高气温 Temp——它没有缺失,五个月各有 30 或 31 天。

library(ggplot2)
library(ggpubr)
library(ggridges)
library(biopalette)

# 公开地址,和数据集页上「下载 CSV」给的是同一个 —— 不写仓库相对路径:
# 那个目录不进仓库,读者 clone 下来也没有这个文件,这段代码就跑不了
d <- read.csv("https://assets.evanzhou.org/tessera/csv/airquality.csv",
              na.strings = "NA")

month_levels <- 5:9
month_labels <- c("May", "June", "July", "August", "September")

# 显式固定 levels:月份有时间顺序,交给字母序会排成 August / July / June…,
# 分层图上曲线的上下顺序就既不合时间、也不合图例
d$month <- factor(
  d$Month,
  levels = month_levels,
  labels = month_labels,
  ordered = TRUE
)

Palettes

月份从五月到九月有自然顺序,但每个月仍是独立分组。用 walter_white2 的五色保持月份身份;分层图按时间顺序排列,让位置承担顺序信息,颜色只负责快速定位。

重叠密度需要透明度,而透明度会改变明度并制造色板里不存在的混合色——重叠区看到的不是原始色板。所以它适合分析,不适合单独用来判断配色。分面和分层版本让每组颜色完整铺开,更适合观察大面积填色、轮廓清晰度以及相邻颜色是否容易混淆。

色板装不下组数时不要循环用色:保留曲线,把未编码的组改成中性上下文。

month_colors <- setNames(
  get_palette("walter_white2", type = "qualitative")[1:5],
  month_labels
)

# 五张图共用同一条横轴:位置和宽度要横着比,坐标就不能各画各的
density_scale <- scale_x_continuous(
  breaks = seq(55, 95, 10),
  expand = expansion(mult = c(0.03, 0.04))
)

density_theme <- theme_pubr(base_size = 13) +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.major.x = element_line(color = "#E5E3DC", linewidth = 0.35),
    axis.line = element_line(color = "#333330", linewidth = 0.45),
    axis.ticks = element_line(color = "#333330", linewidth = 0.4),
    plot.title = element_text(face = "bold", size = 15, hjust = 0),
    plot.subtitle = element_text(color = "grey35", hjust = 0),
    legend.title = element_text(face = "bold"),
    plot.margin = margin(14, 16, 12, 12)
  )

Recipe

No. Method Input Data Palettes
1 ggplot2 d month_colors[1]
2 ggplot2 d month_colors
3 ggplot2 d month_colors
4 ggridges d month_colors
5 ggplot2 d month_colors[1]

1 · 单组密度

先忽略月份,只看 153 天温度的整体分布。

#| fig: basic
#| fig-width: 8
#| fig-height: 5.2
ggplot(d, aes(Temp)) +
  geom_density(
    adjust = 1,        # ggplot2 计算的默认带宽,1 倍
    # 填充只帮助看面积,不承担分组含义 —— 这张图里没有分组
    fill = unname(month_colors[1]),
    color = "#333330",
    alpha = 0.72,
    linewidth = 0.75
  ) +
  density_scale +
  labs(
    title = "Daily temperature distribution",
    subtitle = "New York · May–September 1973 · n = 153",   # n 写进副标题
    x = expression("Temperature (" * degree * "F)"),
    y = "Density"
  ) +
  density_theme
density — basic
density-basic

2 · 分组重叠

把月份映射给 fill 和 color,直接在共同坐标上比较位置差异。

五组已经接近这种画法的可读上限:透明度要降到 0.24 才看得见后画的曲线,而重叠区显示的已经不是原始色板颜色。

#| fig: overlay
#| fig-width: 8
#| fig-height: 5.6
ggplot(d, aes(Temp, fill = month, color = month)) +
  # 轮廓线用同色不降透明度:填充糊在一起时,至少边界还认得出是哪一组
  geom_density(adjust = 0.9, alpha = 0.24, linewidth = 0.8) +
  density_scale +
  scale_fill_manual(values = month_colors) +
  scale_color_manual(values = month_colors) +
  labs(
    title = "Overlaid monthly densities",
    subtitle = "Transparency reveals intersections but changes the palette",
    x = expression("Temperature (" * degree * "F)"),
    y = "Density",
    fill = "Month",
    color = "Month"
  ) +
  density_theme +
  theme(legend.position = "right")
density — overlay
density-overlay

3 · 共同坐标的分面

分面消除遮挡,同时保留相同的横轴与纵轴。

#| fig: facets
#| fig-width: 8.5
#| fig-height: 8.5
ggplot(d, aes(Temp, fill = month, color = month)) +
  # 没有重叠了,透明度可以调回来,颜色接近色板原值
  geom_density(adjust = 0.9, alpha = 0.76, linewidth = 0.75, show.legend = FALSE) +
  # scales = "fixed" 不能省:只有共同尺度下,位置、宽度和峰高才能横向比较。
  # 换成 "free" 的话每个月都会被拉满面板,五个月看起来一样宽
  facet_wrap(~ month, ncol = 1, scales = "fixed") +
  density_scale +
  scale_fill_manual(values = month_colors) +
  scale_color_manual(values = month_colors) +
  labs(
    title = "One density per month",
    subtitle = "Fixed axes preserve comparisons across panels",
    x = expression("Temperature (" * degree * "F)"),
    y = "Density"
  ) +
  density_theme +
  theme(
    strip.background = element_blank(),
    strip.text = element_text(face = "bold", hjust = 0)
  )
density — facets
density-facets

4 · 分层密度

月份有明确顺序时,分层密度比五条透明曲线更清楚:每组轮廓完整、颜色是色板原值,位置本身承担顺序。

#| fig: ridgeline
#| fig-width: 9
#| fig-height: 6.5
ggplot(d, aes(Temp, month, fill = month)) +
  geom_density_ridges(
    adjust = 0.9,
    scale = 0.86,          # < 1:相邻轮廓留出间隙,不发生颜色混合
    rel_min_height = 0.01, # 砍掉低于峰高 1% 的尾巴,免得每层拖一条长毛边
    color = "#333330",
    linewidth = 0.55,
    alpha = 0.92,
    show.legend = FALSE
  ) +
  density_scale +
  scale_fill_manual(values = month_colors) +
  labs(
    title = "Temperature shifts through the summer",
    subtitle = "Five independent density layers · no overlap mixing",
    x = expression("Temperature (" * degree * "F)"),
    y = NULL
  ) +
  theme_pubr(base_size = 13, legend = "none") +
  theme(
    panel.grid.major.x = element_line(color = "#E5E3DC", linewidth = 0.35),
    panel.grid.major.y = element_blank(),
    axis.line.y = element_blank(),     # y 轴只是分层的位置,不是数值轴
    axis.ticks.y = element_blank(),
    plot.title = element_text(face = "bold", size = 15, hjust = 0),
    plot.subtitle = element_text(color = "grey35", hjust = 0),
    plot.margin = margin(14, 16, 12, 12)
  )
density — ridgeline
density-ridgeline

5 · 带宽敏感性

固定数据与样式,只改 adjust。小值保留许多局部峰,大值把它们合并。

如果核心判断只在某一格成立,就不该把它写成稳定结论。

#| fig: bandwidth
#| fig-width: 11
#| fig-height: 4.2
# 包成函数而不是复制三遍:除了 adjust,三格的一切都必须完全相同,
# 否则比较的就不只是带宽了
bandwidth_y_scale <- scale_y_continuous(
  limits = c(0, 0.05),
  breaks = seq(0, 0.05, 0.01),
  expand = expansion(mult = c(0, 0.03))
)

bandwidth_plot <- function(adjust_value) {
  ggplot(d, aes(Temp)) +
    geom_density(
      adjust = adjust_value,
      fill = unname(month_colors[1]),
      color = "#333330",
      alpha = 0.72,
      linewidth = 0.65
    ) +
    density_scale +
    bandwidth_y_scale +  # 三格共用纵轴;不能让每个峰自动填满自己的面板
    labs(
      title = paste0("adjust = ", adjust_value),
      x = expression("Temperature (" * degree * "F)"),
      y = NULL
    ) +
    theme_pubr(base_size = 11, legend = "none") +
    theme(
      panel.grid.major.x = element_line(color = "#E5E3DC", linewidth = 0.3),
      panel.grid.major.y = element_blank(),
      plot.title = element_text(face = "bold", size = 12, hjust = 0),
      plot.margin = margin(8, 8, 8, 8)
    )
}

ggarrange(
  bandwidth_plot(0.55),
  bandwidth_plot(1),
  bandwidth_plot(1.6),
  ncol = 3,
  align = "hv"      # 三格的绘图区严格对齐,横轴才可比
)
density — bandwidth
density-bandwidth

Constraints

  • 曲线高度不是人数。 每条密度曲线下的面积都固定为 1,组间样本量不同也各自归一化。峰高表示相对集中程度,要表达样本量得另外标 n。
  • 带宽决定曲线长相。 一个看起来很明确的双峰,adjust 从 0.55 改到 1.6 可能就消失了。要解释峰数或分布细节,先跑一遍 recipe 5。
  • 小样本的形状不可信。 每组不到二三十个观测时,密度轮廓基本是在描噪声,该画原始点或箱线图。
  • 别用 scale_x_continuous(limits = ) 裁尾部。 它会先删掉范围外的观测再重新估计密度,得到的是另一条曲线。只想放大画面用 coord_cartesian()。

五个 recipe 怎么比

  1. 只看一个变量的整体形状 → recipe 1。
  2. 只有两三组且重叠不严重 → recipe 2。
  3. 需要最稳妥地逐组比较 → recipe 3。
  4. 组别有自然顺序、希望一张图紧凑展示 → recipe 4。
  5. 准备解释峰数或分布细节 → 先用 recipe 5 检查结论对带宽稳不稳。
运行环境4 个包 · 2026-09-02T12:11:17.646+0800

R version 4.5.1 (2025-06-13 ucrt) · x86_64-w64-mingw32

  • biopalette 0.2.2
  • ggplot2 4.0.3
  • ggpubr 0.6.2
  • ggridges 0.5.7