Folioevanzhou.org
← 返回 Tessera
原样121.7 万 行 × 3 列24.6 MB无缺失数据 2026-04-22

hm3_snplist

The HapMap3 SNP list used by LDSC for merge-alleles filtering, with each SNP's two alleles.

One row represents one SNP.

Zenodo · 收录于 2026-04-22

下载 CSV · 24.6 MB
数据预览前 6 行
snpa1a2
rs3094315GA
rs3131972AG
rs3131969AG
rs1048488CT
rs3115850TC
rs2286139CT
变量3 列
字符型3
#变量类型缺失统计
1snpkeydbSNP rsID.字符型—不同值 1217311rs3094315 · rs3131972 · rs3131969 · rs1048488 · rs3115850 · rs2286139 · rs12562034 · rs4040617
2a1First allele, a single base. No indels.字符型—不同值 4T · A · C · G
3a2Second allele, a single base. No indels.字符型—不同值 4C · G · T · A
载入已写好列类型
library(readr)

hm3_snplist <- read_csv(
  "https://assets.evanzhou.org/tessera/csv/hm3_snplist.csv",
  col_types = cols(
    snp = col_character(),
    a1  = col_character(),
    a2  = col_character()
  )
)
URLhttps://assets.evanzhou.org/tessera/csv/hm3_snplist.csv

Source

LDSC 生态里那份标准的 HapMap3 SNP 清单,从 Zenodo(10.5281/zenodo.7773502,Alkes Group w_hm3.snplist)取 w_hm3.snplist.gz 解压转成 CSV。三列:位点的 rsID 和两个等位基因。

脚本读的时候把三列都强制成字符型(colClasses = "character"),并且顺手去掉了原文件可能带的表头行。等位基因必须当字符读——a1/a2 里没有数字,但 rsID 前缀去掉后是纯数字,任何"聪明"的类型推断都可能把它读坏。

它没有位置信息。 只有 rsID 和等位基因,没有染色体、没有坐标、没有频率。它的用途是做交集:把 GWAS 汇总统计过滤到 HapMap3 这个子集上,这是 LDSC 跑 LD score regression 前的标准一步。想要位置得另外 join 一份带坐标的参考。

Use cases

坦白说它不适合画图,收它是为了"拿来用",不是"拿来画"。

  • 条形图:a1 / a2 的碱基构成,各四档
  • 二维计数热图:a1 × a2 的 4×4 组合频次,能看出转换与颠换的比例
  • 读写性能对比:data.table::fread 对 read.csv、polars 对 pandas

那个 4×4 热图倒是个不错的小练习:转换(A↔G、C↔T)应该明显多于颠换,这是真实的分子生物学事实,能拿它验证自己的图有没有画反。

121.7 万行是这批数据里唯一超过百万的,拿它试性能是合理的——差距在这个量级上才看得出来。25MB 的 CSV 用 read.csv() 读会明显卡一下,读的时候给足类型规格会快得多,因为不用猜类型。

生成脚本R · 50 行
# 产物写到 ../csv/hm3_snplist.csv —— 脚本和 CSV 是 content/tessera/data/ 下固定的兄弟目录,
# 所以按脚本自身定位,不依赖你在哪个目录敲这条命令。csv/ 不进仓库(见 .gitignore)。
#
# Rscript 时路径在 --file= 里,source() 时在 sys.frame()$ofile 里,两种都要认:
# 只取其中一种的话,另一种跑法会静默地把 CSV 写到当前目录去。
script_dir <- local({
  arg <- grep("^--file=", commandArgs(trailingOnly = FALSE), value = TRUE)
  path <- if (length(arg)) sub("^--file=", "", arg[[1L]]) else sys.frame(1)$ofile
  dirname(normalizePath(path, mustWork = TRUE))
})
out_csv <- file.path(script_dir, "..", "csv", "hm3_snplist.csv")
dir.create(dirname(out_csv), recursive = TRUE, showWarnings = FALSE)

# Generate the HapMap3 SNP allele list dataset for Tessera.
# Rscript content/tessera/data/script/hm3_snplist.R

out_gz <- file.path(dirname(out_csv), "w_hm3.snplist.gz")

url <- "https://zenodo.org/records/7773502/files/w_hm3.snplist.gz?download=1"


evanverse::download_url(url, out_gz, overwrite = FALSE)

hm3 <- utils::read.table(
  gzfile(out_gz),
  header = FALSE,
  sep = "\t",
  col.names = c("snp", "a1", "a2"),
  colClasses = "character",
  quote = "",
  comment.char = ""
)

if (nrow(hm3) > 0L && identical(unname(as.character(hm3[1, ])), c("SNP", "A1", "A2"))) {
  hm3 <- hm3[-1L, , drop = FALSE]
}

expected_cols <- c("snp", "a1", "a2")
missing_cols <- setdiff(expected_cols, names(hm3))
if (length(missing_cols) > 0L) {
  stop(
    "HapMap3 SNP list is missing columns: ",
    paste(missing_cols, collapse = ", ")
  )
}

hm3 <- hm3[, expected_cols]
utils::write.csv(hm3, out_csv, row.names = FALSE, na = "")

message("Wrote ", out_csv, " with ", nrow(hm3), " rows and ", ncol(hm3), " columns.")

表中统计由 scripts/profile_dataset.py 于 2026-08-03 数出。