← 返回 Tessera▶数据预览前 6 行
| id | time_days | event | sex | age_years | ph_ecog |
|---|
| 1 | 306 | 1 | male | 74 | 1 |
| 2 | 455 | 1 | male | 68 | 0 |
| 3 | 1,010 | 0 | male | 56 | 0 |
| 4 | 210 | 1 | male | 57 | 1 |
| 5 | 883 | 1 | male | 60 | 0 |
| 6 | 1,022 | 0 | male | 74 | 1 |
变量6 列
| # | 变量 | 类型 | 缺失 | 统计 |
|---|
| 1 | idStable row number added by Tessera. | 整数 | — | min 1q1 —中位 114.5均值 114.5q3 —max 228.0 |
| 2 | time_dayskeyFollow-up time in days. | 整数 | — | min 5q1 —中位 255.5均值 305.2q3 —max 1,022 |
| 3 | event0 is censored. 1 is death. | 整数 | — | min 0q1 —中位 1均值 0.724q3 —max 1 |
| 4 | sexSex of the patient. | 字符型 | — | 不同值 2male · female |
| 5 | age_yearsAge at enrolment, in years. | 整数 | — | min 39.0q1 —中位 63.0均值 62.4q3 —max 82.0 |
| 6 | ph_ecogECOG performance score rated by the physician. 0 is fully active, higher is worse. | 整数 | 1 | min 0q1 —中位 1均值 0.952q3 —max 3 |
载入已写好列类型
library(readr)
lung_survival <- read_csv(
"https://assets.evanzhou.org/tessera/csv/lung_survival.csv",
col_types = cols(
id = col_integer(),
time_days = col_integer(),
event = col_integer(),
sex = col_character(),
age_years = col_integer(),
ph_ecog = col_integer()
)
)
URLhttps://assets.evanzhou.org/tessera/csv/lung_survival.csv
Source
R survival 包自带的 lung 数据,记录 North Central Cancer Treatment Group 的 228 例晚期肺癌患者。Tessera 保留绘制生存曲线所需的随访时间、结局、性别、年龄和 ECOG 体能状态,并增加稳定的行号 id。
原数据的隐式编码容易误用:status 是 1 = 删失、2 = 死亡,sex 是 1 = male、2 = female。这里已经明确转换为 event = 0/1 和有名字的 sex;转换过程完整保留在生成脚本中。
Use cases
- Kaplan–Meier 曲线:比较两组随访期间的生存概率,并配合风险表判断尾部稳定性
- 累计风险曲线:从风险累积的角度查看两组差异
- 固定时间点生存率:报告 1 年、2 年等预设时间点的估计值和置信区间
- Cox 模型:练习风险比估计,并用年龄或 ECOG 评分作调整
▶生成脚本R · 31 行
# 产物写到 ../csv/lung_survival.csv —— 脚本和 CSV 是 content/tessera/data/ 下固定的兄弟目录,
# 所以按脚本自身定位,不依赖你在哪个目录敲这条命令。csv/ 不进仓库(见 .gitignore)。
#
# Rscript 时路径在 --file= 里,source() 时在 sys.frame()$ofile 里,两种都要认:
# 只取其中一种的话,另一种跑法会静默地把 CSV 写到当前目录去。
script_dir <- local({
arg <- grep("^--file=", commandArgs(trailingOnly = FALSE), value = TRUE)
path <- if (length(arg)) sub("^--file=", "", arg[[1L]]) else sys.frame(1)$ofile
dirname(normalizePath(path, mustWork = TRUE))
})
out_csv <- file.path(script_dir, "..", "csv", "lung_survival.csv")
dir.create(dirname(out_csv), recursive = TRUE, showWarnings = FALSE)
lung <- survival::lung
# survival::lung uses 1 = censored / 2 = death and 1 = male / 2 = female.
# Freeze those implicit codes into explicit analysis-ready columns for Tessera.
lung_survival <- data.frame(
id = seq_len(nrow(lung)),
time_days = lung$time,
event = ifelse(lung$status == 2L, 1L, 0L),
sex = ifelse(lung$sex == 1L, "male", "female"),
age_years = lung$age,
ph_ecog = lung$ph.ecog
)
write.csv(
lung_survival,
out_csv,
row.names = FALSE
)
表中统计由 scripts/profile_dataset.py 于 2026-08-13 数出。