Folioevanzhou.org
← 返回 Tessera
模拟6,000 行 × 4 列236 KB无缺失数据 2026-08-20

gwas_manhattan

A deterministic simulated genome-wide association scan across 12 chromosomes, with fixed peak positions.

One row represents one marker.

Tessera · 收录于 2026-08-20

下载 CSV · 236 KB
数据预览前 6 行
MarkerChrPosP
rs1_11378,5260.458
rs1_21407,2080.247
rs1_31573,3570.164
rs1_41626,3320.383
rs1_51886,3540.683
rs1_611,182,7120.524
变量4 列
字符型1整数2双精度1
类型
#变量类型缺失统计
1MarkerMarker name, built from the chromosome number and the within-chromosome index.字符型—不同值 6000rs1_1 · rs1_2 · rs1_3 · rs1_4 · rs1_5 · rs1_6 · rs1_7 · rs1_8
2ChrChromosome number.整数—min 1q1 —中位 6.5均值 6.5q3 —max 12.0
3PosPosition within the chromosome, in base pairs.整数—min 4,491q1 —中位 4.9e+7均值 5.0e+7q3 —max 1.0e+8
4PkeyAssociation p-value.双精度—min 4.8e-10q1 —中位 0.497均值 0.498q3 —max 1
载入已写好列类型
library(readr)

gwas_manhattan <- read_csv(
  "https://assets.evanzhou.org/tessera/csv/gwas_manhattan.csv",
  col_types = cols(
    Marker = col_character(),
    Chr    = col_integer(),
    Pos    = col_integer(),
    P      = col_double()
  )
)
URLhttps://assets.evanzhou.org/tessera/csv/gwas_manhattan.csv

Source

这是一份由 Tessera manhattan recipe 固化的确定性模拟 GWAS 数据,不对应真实队列、表型或基因型。每条染色体包含 500 个标记,位置按染色体内坐标排序;背景 P 值来自固定随机种子,18 个候选信号使用固定标记名和固定 −log10(P)。

CSV 冻结后,矩形 Manhattan、圆形 Manhattan 与 Palette Lab 读取同一份数据,避免各处现场模拟产生不同峰位。

模拟时固化的三个显著性约定:

  • Suggestive threshold:P = 1e-6
  • Genome-wide threshold:P = 5e-8
  • Top hits:固定的 8 个示范标记,仅用于标签布局练习

Use cases

  • 矩形 Manhattan plot:精确比较染色体区段、阈值和候选位点
  • 圆形 Manhattan plot:在方形版面中概览全基因组信号分布
  • QQ plot:检查模拟 P 值背景与极端尾部(需另行计算期望分位数)
生成脚本R · 42 行
# 产物写到 ../csv/gwas_manhattan.csv —— 脚本和 CSV 是 content/tessera/data/ 下固定的兄弟目录,
# 所以按脚本自身定位,不依赖你在哪个目录敲这条命令。csv/ 不进仓库(见 .gitignore)。
#
# Rscript 时路径在 --file= 里,source() 时在 sys.frame()$ofile 里,两种都要认:
# 只取其中一种的话,另一种跑法会静默地把 CSV 写到当前目录去。
script_dir <- local({
  arg <- grep("^--file=", commandArgs(trailingOnly = FALSE), value = TRUE)
  path <- if (length(arg)) sub("^--file=", "", arg[[1L]]) else sys.frame(1)$ofile
  dirname(normalizePath(path, mustWork = TRUE))
})
out_csv <- file.path(script_dir, "..", "csv", "gwas_manhattan.csv")
dir.create(dirname(out_csv), recursive = TRUE, showWarnings = FALSE)

set.seed(2026)

n_chr <- 12
markers_per_chr <- 500

gwas_manhattan <- do.call(rbind, lapply(seq_len(n_chr), function(chr) {
  data.frame(
    Marker = paste0("rs", chr, "_", seq_len(markers_per_chr)),
    Chr = chr,
    Pos = sort(sample(seq(1, 100000000), markers_per_chr)),
    P = runif(markers_per_chr, min = 1e-4, max = 1),
    stringsAsFactors = FALSE
  )
}))

signals <- c(
  rs3_115 = 8.82, rs3_166 = 8.91,
  rs7_8 = 9.32, rs7_41 = 9.04, rs7_238 = 9.25,
  rs8_355 = 7.82, rs10_177 = 7.12, rs12_500 = 8.92,
  rs4_321 = 6.62, rs5_118 = 6.08, rs5_383 = 6.38,
  rs6_207 = 7.58, rs6_365 = 7.31, rs7_407 = 6.34,
  rs9_177 = 7.67, rs9_389 = 7.36, rs11_143 = 7.18, rs12_85 = 7.32
)

signal_rows <- match(names(signals), gwas_manhattan$Marker)
stopifnot(!anyNA(signal_rows))
gwas_manhattan$P[signal_rows] <- 10^(-unname(signals))

write.csv(gwas_manhattan, out_csv, row.names = FALSE)

表中统计由 scripts/profile_dataset.py 于 2026-08-20 数出。