kermt-finetune
Finetune a pretrained KERMT encoder on a labeled CSV. Validate the checkpoint and data, prepare features, and run containerized training. Use a local checkpoint or optionally download a pinned Hugging Face model bundle using HF_TOKEN if configured.
Install / Use
npx skills add NVIDIA/skills --skill bionemo-kermt-finetuneInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of kermt-finetune
kermt-finetune scores 92/100 on our quality scale, 712th of 2,125 Automation skills we index (top 34%).
Its SKILL.md is 16 KB long, well organised into 11 sections with 1 code example: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so kermt-finetune is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
kermt-finetune compared with similar skills
All 4 of these similar skills score higher than kermt-finetune; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| kermt-finetune (this skill)by NVIDIA | 92 | 3.4k | 5d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.0k | 13d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 84.4k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
Frequently asked questions
- How do I install kermt-finetune?
- Run
npx skills add NVIDIA/skills --skill kermt-finetune. The install tabs above show the steps for each supported agent. - Which AI agents does kermt-finetune work with?
- It is written for Claude Code, Zed and OpenAI Codex, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is kermt-finetune safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is kermt-finetune still maintained?
- The repository was last updated 5 days ago, so kermt-finetune is actively maintained.
Skill content
View source on GitHubname: kermt-finetune description: Finetune a pretrained KERMT encoder on a labeled CSV. Validate the checkpoint and data, prepare features, and run containerized training. Use a local checkpoint or optionally download a pinned Hugging Face model bundle using HF_TOKEN if configured. Write model bundles, prepared data, logs, and trained models to user-selected host directories. license: Apache-2.0 compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron. metadata: owner: evax@nvidia.com classification: workflow-skill risk_tier: skill
Line/token budget: targets ~250 lines / ~3000 tokens — within the
500-line / 5000-token cap for skill files.
kermt-finetune
Finetune a pretrained KERMT encoder on a user-supplied labeled CSV. The skill is the workflow orchestrator: validate ckpt, validate data, prepare data, launch the runner detached, return a run directory + container name.
Skill and runtime paths
Set SKILL_DIR to the absolute path of this installed skill directory. Export
KERMT_REPO as the absolute path to the KERMT checkout used for model
execution. The bundled container helper mounts that checkout at
/workspace and this skill at /skill (read-only). Commands inside
the container use /skill/scripts/; defaults are bundled in config/.
See Released models for checkpoint bundle requirements.
Downloads and local outputs
The optional released-model branch reads config/released_model.json for the
Hugging Face repository, pinned revision, and filenames. The bundled
scripts/fetch_released_model.py downloads the model bundle over HTTPS into
the host directory the user selects. Public models work without credentials;
if HF_TOKEN is set, the container helper forwards it for Hugging Face
authentication. Prepared data, logs, and workflow results go into the chosen
run directory.
Hardware requirements
- GPUs: 1 by default (single-GPU); pass
--gpus 0(or whichever id) to select one. For faster training on a multi-GPU host, pass--num-gpus N(N>1) to run data-parallel DDP across N GPUs —--batch-sizeis then per-GPU (effective global batch = batch_size × N). - VRAM: ≥ 8 GB for the default
batch_size 32configuration. Lower VRAM works at smaller batch sizes — pass--batch-size Nto override. - Disk: a few GB per run (checkpoint + features + logs).
- Driver / CUDA: any host supporting CUDA 12.6 (the kermt image base).
kermt-setupvalidates this up-front.
Inputs
Required:
--csv <path>— labeled CSV. First column issmiles; every other column is a target.
Checkpoint (optional — defaults to the released model if omitted):
--ckpt <path>— input pretrain checkpoint (grover_base / cmim / hybrid). The validator refuses already-finetuned ckpts with a redirect tokermt-infer. If omitted, the skill offers to download the released pretrained hybrid model nvidia/NV-KERMT-70M-v2 and finetune from it — see "Resolve & validate the checkpoint" (workflow step 3).--pretrained-release— explicit opt-in to use the released model without the interactive prompt (for non-interactive / agent runs). Mutually exclusive with--ckpt.--model-dir <dir>— where to save the downloaded bundle (default$KERMT_REPO/models/NV-KERMT-70M-v2/). An already-complete bundle there is reused, not re-downloaded.
Optional:
-
--dataset-type {regression | classification | multiclass}— defaultregression(fromdefaults_finetune.json). Drives loss, metric defaults, and head initialization. For classification tasks pass--dataset-type classification. -
--targets COL [COL ...]— explicit target column names. If omitted, the validator auto-detects numeric non-smiles columns and the skill confirms with the user before proceeding. -
--val-csv <path>and--test-csv <path>— user-provided val + test splits. Either pass both or pass neither (the skill auto-splits using the configured--split-type). -
--split-type {random | scaffold_balanced | index_predetermined}— defaultscaffold_balancedfromdefaults_finetune.json.randomandscaffold_balanced: build the val/test split internally from the train CSV. No--val-csv/--test-csvneeded.index_predetermined: requires pre-split CSVs passed via--val-csv+--test-csv(and, separately, per-fold index files — seekermt/util/utils.split_data). Use this when the dataset ships its own canonical split (e.g.tests/data/Biogen_for_grover/scaffold/ balance/<endpoint>/{train,val,test}.csv).
-
--metric NAME—mae(regression default),auc(classification default), or any namekermt.util.metrics.get_metric_funcaccepts. -
--epochs N/--batch-size N/--init-lr F/--max-lr F/--final-lr F/--warmup-epochs F/--weight-decay F/--dropout F/--bond-drop-rate F/--dist-coff F/--early-stop-epoch N/--seed N— training-hyperparameter overrides. Anything not given is filled fromconfig/defaults_finetune.json. -
--ffn-hidden-size N/--ffn-num-layers N— shared FFN trunk dims. -
--ffn-num-task-specific-layers N/--ffn-task-specific-hidden-size H— per-target FFN heads (default 0 = off; useful for heterogeneous multi-target finetunes). Both must be set together when N > 0. -
--ensemble-size N/--num-folds N— multi-model / k-fold CV. Default 1 each. -
--gpus 0— single GPU id for single-process finetune (default 0). Ignored when--num-gpus > 1. -
--num-gpus N— number of GPUs for data-parallel DDP finetune. Default 1 (single-process, unchanged). N>1 runsmain.py finetunewithWORLD_SIZE=N(one process per GPU);--batch-sizeis per-GPU. -
--from-prepare <dir>— skip the prepare step and reuse an existingprepare_data.jsonin<dir>. Useful when iterating on hyperparameters.
Workflow
Let $KERMT_REPO be the path to your kermt repo checkout, and assume
kermt-setup has built kermt:latest. All paths below are on the host; the
helper bind-mounts them at known container paths.
-
Pre-flight: ensure container + system probe.
"$SKILL_DIR/scripts/kermt_container.sh" check_system | python -c " import json, sys; d = json.load(sys.stdin) if not d['ok']: print('System check failed:', d['gaps']); sys.exit(1) print(f'OK: {len(d[\"gpus\"])} GPU(s); CUDA via container toolkit') "Refuse to proceed if
ok: false. -
Compute run directory.
RUN_DIR=$KERMT_REPO/runs/finetune_$(date -u +%Y-%m-%dT%H-%M-%SZ) -
Resolve & validate the checkpoint.
Resolve — only if
--ckptwas omitted. Default to the released pretrained hybrid model nvidia/NV-KERMT-70M-v2:- Consent gate. Unless
--pretrained-releasewas passed, ask the user: "No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2 (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2) and finetune from it? [y/N]". Never download without an explicit yes (or--pretrained-release). If both--ckptand--pretrained-releaseare given, abort — they conflict. - Save location. Default
$KERMT_REPO/models/NV-KERMT-70M-v2/; honor--model-dir <dir>if given. An already-complete bundle is reused. - Download (foreground; ~282 MB on first fetch):
Parse the JSON; abort on"$SKILL_DIR/scripts/kermt_container.sh" run --model-dir <save-dir> -- \ "python /skill/scripts/fetch_released_model.py --out /model"ok: false(surfaceerrors). On success set<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt.
Validate the resolved (or user-provided) ckpt:
"$SKILL_DIR/scripts/kermt_container.sh" run --ckpt <user-ckpt> -- \ "python /skill/scripts/check_checkpoint.py --mode finetune_init --ckpt /ckpt"Parse the JSON. Abort on
ok: false. The validator rejects already- finetuned ckpts (has_task_ffn: true) with a redirect tokermt-infer. - Consent gate. Unless
-
Validate the data.
"$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> -- \ "python /skill/scripts/check_data.py --mode finetune --csv /data/<basename> [--targets COL1 COL2 ...]"If
--targetswas not given by the user, surfaceauto_detected_targetsfrom the JSON and ask the user to confirm before continuing. Abort onok: false. -
Prepare the data (skip if
--from-preparegiven).Pre-flight: check for sibling val.csv / test.csv. Before invoking prepare_data, inspect the parent directory of
<user-csv>. If a canonical-looking siblingval.csv(orval_*.csv— common variants includeval_T.csv,val_clean.csv) AND a matchingtest.csv/test_*.csvexist next to the train CSV, the dataset ships its own pre-defined split. In that case set--split-type index_predeterminedAND pass--val-csv/--test-csv— otherwise the configuredsplit_type(defaultscaffold_balanced) will re-split the train CSV from scratch and silently discard the user's val/test files. When in doubt — or when the sibling files use non-canonical suffixes (_T,_v2, etc.) — surface the situation to the user and ask which they want.Quoting target names. If any of the
--targetscolumn names contain shell metacharacters (>,&,|,(,),$, etc.), single-quote each one when passing on the CLI to keep the shell from eating part of the name. Example:--targets 'Log_Caco2_Papp_A>B' 'logD'. The CSV header itself is read directly by the downstream trainer and is unaffected, but the prepare_data.json manifest'stargets[]field captures whatever the shell delivers — unquoted metacharacters get truncated there.Mount note:
kermt_container.sh --data <host-csv>mounts the parent directory of<host-csv>at/data.--val-csvand--test-csvmust therefore reference files in that same parent directory. If val/test live in a separate directory (e.g. a siblingsplits/folder), mount the parent of all three using--data <dir>on a directory rather than a file."$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> --run-dir $RUN_DIR -- \ "python /skill/scripts/prepare_data.py --mode finetune \\ --csv /data/<basename> --out /runs/data \\ --split-type <split_type> \\ [--val-csv /data/<val-basename> --test-csv /data/<test-basename>] \\ [--val-frac 0.1 --test-frac 0.1 --seed 0] \\ --targets <COL1> [COL2 ...]"Outputs land at
$RUN_DIR/data/prepare_data.json. Forscaffold_balancedandindex_predetermined, prep emits a singleclean_full_csv+.npz; the runner passes them through tomain.py finetunewhich callssplit_datainternally with the user-supplied seed. -
Estimate runtime + echo applied defaults.
- Finetune wall time is typically minutes-to-hours on 1 GPU.
- Surface a summary of every flag that was filled from the defaults
vs user-supplied, so the user knows what was assumed. The runner
records this in
args_applied. - Sample message:
"Filling from defaults_finetune.json: epochs=30, batch_size=32, split_type=scaffold_balanced. Override any of these with --<flag>."
-
Targets confirmation gate (hard requirement). Before launching the runner, regardless of how the targets list was determined (CLI
--targets, auto-detection in step 4, or a user natural-language request like "finetune on Caco2 and HLM"), echo the final targets list to the user with an explicit count:"Will finetune on N target(s): COL1, COL2, ...". If the user's request specified a subset that doesn't match this list (e.g., they asked for 2 tasks via natural language bu
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.4k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
84.4k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
