SkillAgentSearch skills...

scikit-survival-analysis

Time-to-event modeling with scikit-survival: Cox PH (elastic net), Random Survival Forests, Boosting, SVMs for censored data. C-index, Brier, time-dependent AUC; Kaplan-Meier, Nelson-Aalen, competing risks. Pipeline/GridSearchCV compatible.

Install / Use

npx skills add jaechang-hits/SciAgent-Skills --skill scikit-survival-analysis

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

91/100

Category

Automation

Supported Platforms

Universal

Our assessment of scikit-survival-analysis

scikit-survival-analysis scores 91/100 on our quality scale, 1088th of 2,866 Automation skills we index (top 38%).

Its SKILL.md is 27 KB long, well organised into 79 sections with 19 code examples: a thorough specification that gives an agent plenty to work with.

It has 367 GitHub stars, a meaningful sign that others use it.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
11/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 37 days ago, so scikit-survival-analysis is actively maintained.
  • No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
  • Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

scikit-survival-analysis compared with similar skills

All 4 of these similar skills score higher than scikit-survival-analysis; compare them before choosing.

SkillScoreStarsUpdatedFormat
scikit-survival-analysis (this skill)by jaechang-hits9136737d agoSKILL.md
Agent-Reachby Panniantong10090.8k19d agoCLAUDE.md
headroomby headroomlabs-ai10074.4ktodayCLAUDE.md
Scraplingby D4Vinci10085.7ktodayMCP Server
crawl4aiby unclecode10084.8k9d agoMCP Server

Frequently asked questions

How do I install scikit-survival-analysis?
Run npx skills add jaechang-hits/SciAgent-Skills --skill scikit-survival-analysis. The install tabs above show the steps for each supported agent.
Which AI agents does scikit-survival-analysis work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is scikit-survival-analysis safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is scikit-survival-analysis still maintained?
The repository was last updated 37 days ago, so scikit-survival-analysis is actively maintained.

name: scikit-survival-analysis description: "Time-to-event modeling with scikit-survival: Cox PH (elastic net), Random Survival Forests, Boosting, SVMs for censored data. C-index, Brier, time-dependent AUC; Kaplan-Meier, Nelson-Aalen, competing risks. Pipeline/GridSearchCV compatible. Use statsmodels for frequentist, pymc for Bayesian, lifelines for parametric." license: GPL-3.0

scikit-survival -- Survival Analysis

Overview

scikit-survival is a Python library for time-to-event analysis built on scikit-learn. It handles right-censored data (observations where the event has not yet occurred) using Cox models, ensemble methods, survival SVMs, and non-parametric estimators. All models follow the scikit-learn fit/predict API and integrate with Pipelines, cross-validation, and GridSearchCV.

When to Use

  • Modeling time-to-event outcomes with right-censored data (clinical trials, reliability)
  • Fitting Cox proportional hazards models (standard or elastic net penalized)
  • Building ensemble survival models (Random Survival Forest, Gradient Boosting)
  • Training survival SVMs for margin-based learning on medium-sized datasets
  • Evaluating survival predictions with censoring-aware metrics (C-index, Brier score, AUC)
  • Estimating non-parametric survival curves (Kaplan-Meier, Nelson-Aalen)
  • Analyzing competing risks with cumulative incidence functions
  • High-dimensional survival data with automatic feature selection (CoxNet L1/L2)
  • For simpler parametric models (Weibull, log-normal AFT) or statistical tests (log-rank), use lifelines
  • For deep learning survival models, use pycox or torchlife

Prerequisites

pip install scikit-survival scikit-learn pandas numpy matplotlib

Python: >= 3.9. Dependencies: scikit-learn, numpy, scipy, pandas, joblib, osqp (for some SVM solvers).

Data format: Survival outcomes are NumPy structured arrays with (event, time) fields. Events are boolean (True = event occurred, False = censored). Times are positive floats.

Pre-flight Interview

Settle these with the user before writing any analysis code.

decisions:
  - id: D1
    param: timeAndEventColumns
    kind: required
    source: data
    ask: "Which column holds follow-up time, which holds the event indicator, and does the indicator mark the event or censoring?"
    default: null

  - id: D2
    param: eventDefinition
    kind: required
    source: user
    depends_on: [D1]
    ask: "Which event is being modelled, and are competing events treated as censoring?"
    default: null

  - id: D3
    param: covariates
    kind: required
    source: data
    ask: "Which variables enter the model, and are any measured after baseline?"
    default: null

  - id: D4
    param: modelFamily
    kind: required
    source: user
    ask: "A proportional-hazards model whose coefficients are interpretable, or an ensemble that predicts better but does not yield hazard ratios?"
    default: "Cox proportional hazards"

  - id: D5
    param: regularization
    kind: required
    source: user
    depends_on: [D3, D4]
    ask: "With more covariates than events, should coefficients be penalized - and toward selection or toward shrinkage?"
    default: "none"

  - id: D6
    param: validationScheme
    kind: required
    source: user
    ask: "How should performance be estimated - cross-validation, a held-out split, or an external cohort?"
    default: "cross-validation"

  - id: D7
    param: ensembleHyperparameters
    kind: optional_conditional
    source: user
    depends_on: [D4]
    ask: "How many trees, how deep, and at what learning rate?"
    default: "100 estimators, depth 3, learning rate 0.1"
    skip_if: "a proportional-hazards model was chosen"

D1 and D2 are asked before anything else because a flipped event indicator fits a model of the time to not having the event, and every hazard ratio comes back inverted without any warning. D3 carries the other silent error: a covariate measured after baseline leaks the outcome into the predictors and produces an impressive concordance index that will not replicate.

Quick Start

from sksurv.datasets import load_breast_cancer
from sksurv.ensemble import RandomSurvivalForest
from sksurv.metrics import concordance_index_ipcw
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

rsf = RandomSurvivalForest(n_estimators=100, random_state=42)
rsf.fit(X_train, y_train)

risk_scores = rsf.predict(X_test)
c_index = concordance_index_ipcw(y_train, y_test, risk_scores)[0]
print(f"C-index: {c_index:.3f}")  # e.g., 0.68

# Individual survival curves
surv_fns = rsf.predict_survival_function(X_test[:2])
for fn in surv_fns:
    print(f"5-year survival: {fn(365 * 5):.3f}")

Core API

Module 1: Data Preparation

Create structured survival arrays and preprocess features.

import numpy as np
import pandas as pd
from sksurv.util import Surv
from sksurv.preprocessing import OneHotEncoder, encode_categorical
from sksurv.datasets import load_gbsg2, load_breast_cancer
from sklearn.preprocessing import StandardScaler

# Create survival outcome from arrays
event = np.array([True, False, True, True, False])
time = np.array([120.0, 365.0, 200.0, 90.0, 400.0])
y = Surv.from_arrays(event=event, time=time)
print(y.dtype)  # [('event', '?'), ('time', '<f8')]

# From DataFrame columns
# y = Surv.from_dataframe("event_col", "time_col", df)

# Load built-in datasets
# Available: load_gbsg2, load_breast_cancer, load_veterans_lung_cancer,
#            load_whas500, load_aids, load_flchain
X, y = load_gbsg2()
print(f"Shape: {X.shape}, Events: {y['event'].sum()}, "
      f"Censoring rate: {1 - y['event'].mean():.1%}")

# Encode categoricals (survival-aware one-hot)
X_encoded = encode_categorical(X)  # auto-detect and encode all categorical cols

# Standardize (critical for Cox and SVM models)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X_encoded)
from sksurv.io import loadarff

# Load ARFF format (Weka format)
data = loadarff("survival_data.arff")
X_arff, y_arff = data[0], data[1]  # DataFrame, structured array

Module 2: Cox Proportional Hazards

Semi-parametric model: h(t|x) = h_0(t) * exp(beta^T x). Interpretable coefficients as log hazard ratios.

from sksurv.linear_model import CoxPHSurvivalAnalysis, CoxnetSurvivalAnalysis, IPCRidge

# Standard Cox PH model
cox = CoxPHSurvivalAnalysis(alpha=0.0, ties="breslow")
cox.fit(X_train, y_train)
print(f"Coefficients: {cox.coef_}")  # log hazard ratios
# Hazard ratio interpretation: exp(coef) = HR for 1-unit increase
risk_scores = cox.predict(X_test)  # Higher = higher risk

# Survival function for individual patients
surv_funcs = cox.predict_survival_function(X_test[:3])
for fn in surv_funcs:
    print(f"5-year survival: {fn(365 * 5):.3f}")
# Penalized Cox (elastic net) -- for high-dimensional data (p > n)
coxnet = CoxnetSurvivalAnalysis(
    l1_ratio=0.9,           # 0=Ridge, 1=Lasso, between=Elastic Net
    alpha_min_ratio=0.01,   # smallest alpha / largest alpha ratio
    n_alphas=100,           # steps in regularization path
)
coxnet.fit(X_train, y_train)

# Feature selection: non-zero coefficients
selected = np.where(coxnet.coef_ != 0)[0]
print(f"Selected {len(selected)} / {X_train.shape[1]} features")

# IPCRidge: accelerated failure time model (predicts log survival time)
ipcridge = IPCRidge(alpha=1.0)
ipcridge.fit(X_train, y_train)
log_survival_time = ipcridge.predict(X_test)

Module 3: Ensemble Methods

Non-parametric tree-based models for complex non-linear relationships.

from sksurv.ensemble import (
    RandomSurvivalForest,
    GradientBoostingSurvivalAnalysis,
    ComponentwiseGradientBoostingSurvivalAnalysis,
    ExtraSurvivalTrees,
)

# Random Survival Forest -- robust, minimal tuning
rsf = RandomSurvivalForest(
    n_estimators=200, min_samples_split=10, min_samples_leaf=15,
    max_features="sqrt", random_state=42, n_jobs=-1,
)
rsf.fit(X_train, y_train)
risk = rsf.predict(X_test)

# Gradient Boosting -- best performance, needs tuning
gbs = GradientBoostingSurvivalAnalysis(
    loss="coxph",           # "coxph" or "ipcwls" (AFT)
    n_estimators=300, learning_rate=0.05, max_depth=3,
    subsample=0.8, dropout_rate=0.1, random_state=42,
)
gbs.fit(X_train, y_train)

# ComponentwiseGB -- linear model with automatic feature selection
cgbs = ComponentwiseGradientBoostingSurvivalAnalysis(
    n_estimators=100, learning_rate=0.1,
)
cgbs.fit(X_train, y_train)
print(f"Non-zero coefficients: {np.sum(cgbs.coef_ != 0)}")

# ExtraSurvivalTrees -- more regularized than RSF, faster training
est = ExtraSurvivalTrees(n_estimators=100, random_state=42)
est.fit(X_train, y_train)

# Survival curves from any ensemble model
surv_funcs = rsf.predict_survival_function(X_test[:1])
chf_funcs = rsf.predict_cumulative_hazard_function(X_test[:1])

Module 4: Survival SVMs

Margin-based learning for survival ranking. Always standardize features.

from sksurv.svm import FastSurvivalSVM, FastKernelSurvivalSVM, HingeLossSurvivalSVM

# Linear SVM -- fast, for linear relationships
lsvm = FastSurvivalSVM(alpha=1.0, rank_ratio=1.0, max_iter=100, random_state=42)
lsvm.fit(X_train_scaled, y_train)
risk = lsvm.predict(X_test_scaled)

# Kernel SVM -- for non-linear relationships (rbf, poly, sigmoid)
ksvm = FastKernelSurvivalSVM(
    alpha=1.0, kernel="rbf", gamma="scale",
    max_iter=50, random_state=42,
)
ksvm.fit(X_train_scaled, y_train)

# Hinge loss variant
hsvm = HingeLossSurvivalSVM(alpha=1.0, random_state=42)
hsvm.fit(X_train_scaled, y_train)
# NaiveSurvivalSVM also available but slower (O(n^3))
from sksurv.kernels import ClinicalKernelTransform

# Clinical kernel: combines clinical + molecular features
# Weighs clinical variables separately from high-dimensional molecular data
transform = ClinicalKernelTransform(fit_once=True)
transform.prepare(X_train)  # auto-detect clinical features
X_kern = transform.fit_transform(X_train)

Module 5: Non-Parametric Estimation

Estimate survival and hazard curves without model assumptions.

from sksurv.nonparametric import kaplan_meier_estimator, nelson_aalen_estimator
import matplotlib.pyplot as plt

# Kaplan-Meier survival curve
time_km, surv_prob = kaplan_meier_estimator(y["event"], y["time"])
plt.step(time_km, surv_prob, where="post")
plt.xlabel("Time (days)")
plt.ylabel("Survival probability")
plt.title("Kaplan-Meier Estimate")
plt.savefig("km_curve.png", dpi=150)

# With confidence intervals
time_km, surv_prob, conf_int = kaplan_meier_estimator(
    y["event"], y["time"], conf_type="log-log",
)

# Nelson-Aalen cumulative hazard
time_na, cum_hazard = nelson_aalen_estimator(y["event"], y["time"])

# Stratified KM by group
for group_name, mask in [("Treated", treated_mask), ("Control", control_mask)]:
    t, s = kaplan_meier_estimator(y[mask]["event"], y[mask]["time"])
    plt.step(t, s, where="post", label=group_name)
plt.legend()

Module 6: Evaluation Metrics

Censoring-aware metrics for discrimination and calibration.

from sksurv.metrics import (
    concordance_index_censored,
    concordance_index_ipcw,
    cumulative_dynamic_auc,
    integrated_brier_score,
    brier_score,
    as_concordance_index_ipcw_scorer,
    as_integrated_brier_score_scorer,
)
import numpy as np

risk_scores = model.predict(X_test)

# Harrell's C-index (simple, biased with high censoring)
c_harrell = concordance_index_censored(
    y_test["event"], y_test["time"], risk_scores
)[0]

# Uno's C-index (recommended -- robust to censoring)
c_uno = concordance_index_ipcw(y_train, y_test, risk_scores)[0]
print(f"C-index (Harrell): {c_harrell:.3f}, (Uno): {c_uno:.3f}")

# Time-dependent AUC at clinically relevant timepoints
times = np.array([365, 730, 1095])  # 1, 2, 3 years
auc, mean_a

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars367
CategoryAutomation
Updated1mo ago
Forks36

Languages

Python

Trust signals

88/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 medium