launch-nemo-rl
Playbook for launching, monitoring, stopping, and debugging NeMo-RL recipes on a Kubernetes cluster via the nrl-k8s CLI. Covers ephemeral vs long-lived RayCluster modes, iterating on runs, and debugging hung or failed training jobs.
Install / Use
npx skills add NVIDIA/skills --skill launch-nemo-rlInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Tags
Our assessment of launch-nemo-rl
launch-nemo-rl scores 95/100 on our quality scale, 89th of 487 Operations skills we index (top 19%).
Its SKILL.md is 17 KB long, well organised into 33 sections with 14 code examples: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so launch-nemo-rl is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-29. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
launch-nemo-rl compared with similar skills
All 4 of these similar skills score higher than launch-nemo-rl; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| launch-nemo-rl (this skill)by NVIDIA | 95 | 3.4k | 5d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 7d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 7d ago | SKILL.md |
Frequently asked questions
- How do I install launch-nemo-rl?
- Run
npx skills add NVIDIA/skills --skill launch-nemo-rl. The install tabs above show the steps for each supported agent. - Which AI agents does launch-nemo-rl work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is launch-nemo-rl safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is launch-nemo-rl still maintained?
- The repository was last updated 5 days ago, so launch-nemo-rl is actively maintained.
Skill content
View source on GitHubname: launch-nemo-rl license: Apache-2.0 description: Playbook for launching, monitoring, stopping, and debugging NeMo-RL recipes on a Kubernetes cluster via the nrl-k8s CLI. Covers ephemeral vs long-lived RayCluster modes, iterating on runs, and debugging hung or failed training jobs. when_to_use:
- "run this recipe on k8s"
- "launch on the cluster"
- "submit a training job"
- "tear down the cluster"
- "resubmit as rayjob"
- "why is the run stuck"
- "how do I get logs for job X"
- "bring the cluster back up" allowed-tools: Bash Read Grep Glob Edit Write
launch-nemo-rl — running NeMo-RL recipes on Kubernetes via nrl-k8s
This is the playbook for the nrl-k8s CLI at infra/nrl_k8s/. Follow it when the user asks to launch / iterate / debug a NeMo-RL recipe on a Kubernetes cluster. Verify current state (kubectl, git log, the recipe + infra files) before acting — the cluster is shared and the cost of a wrong action is high.
1. One command, two modes
There is a single top-level submission command: nrl-k8s run. It has two lifecycle modes.
| Mode | Invocation | When to use | Cluster after? |
| :----------------- | :---------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------------- |
| Ephemeral (default) | nrl-k8s run | One-shot. KubeRay applies a RayJob, runs, tears the cluster down. Best for most runs. | No (auto) |
| Long-lived | nrl-k8s run --raycluster | Dev loop. Reuses a matching live cluster, applies if absent, warns + reuses on drift (pass --recreate to replace). Then submits daemons and training. First-choice for iteration. | Yes |
Ask: Do I need this cluster after the run? If yes, use --raycluster. Otherwise use the default (ephemeral).
The rest of the CLI is observability / stage-by-stage control:
| Command | Purpose |
| :---------------------- | :---------------------------------------------------------------------------------------------- |
| nrl-k8s check | Validate a recipe + infra pair; optionally write the fully-resolved manifests (-o). |
| nrl-k8s status | Per-role RayCluster state, head pod phase, worker pod phases, daemon job status. |
| nrl-k8s cluster up/down/list/dashboard | Manage RayClusters independently of a run (e.g. render a manifest with --dry-run). |
| nrl-k8s job list/logs/stop | Observability over Ray Jobs already submitted to a role's cluster. |
| nrl-k8s logs | Tail a role's pod / daemon logs without needing a submission id. |
2. Recipe + infra pair
Every launch takes two files. Pass the infra with --infra, not merged inline:
nrl-k8s run infra/nrl_k8s/examples/<recipe>.yaml \
--infra infra/nrl_k8s/examples/<recipe>.<profile>.infra.yaml
- Recipe (e.g.
qwen3_30b_math_8n_4gpu.yaml) — NeMo-RL config: model, GRPO/SFT knobs,cluster.{gpus_per_node,num_nodes}. Usesdefaults:to inherit fromexamples/configs/recipes/llm/.... - Infra (e.g.
*.<profile>.infra.yaml) — K8s/Ray shape: namespace, image, service account, RayCluster spec underkuberay:, optional Deployments underdeployments:,submit.submitter,launch.{mode,codeSource,codePath,entrypoint}. Pair names follow<recipe>.<profile>[.prod].infra.yamlwhere<profile>names the hardware target (e.g.gb300).
Example pairs in infra/nrl_k8s/examples/ — read the neighbouring files to see the current conventions for the target profile.
3. Long-lived mode flags
Three independent dimensions. --mode is a macro that picks defaults; individual flags override it.
--mode interactive → --submitter portForward --code-source upload (tails logs)
--mode batch → --submitter exec --code-source image (returns after nohup)
- Submitter:
portForwarduseskubectl port-forward+ Ray Job SDK (gets asubmission_idthe dashboard tracks).execuseskubectl exec+nohupon the head pod (no submission_id; driver appears astype=DRIVERin the dashboard). - Code source:
uploadstages a working_dir from the laptop (Ray 100 MiB cap).image/lustreexpect code on the pod's filesystem — paired with--code-path(typically/opt/nemo-rl), which is a subPath of the shared-filesystem PVC mount in the standard infra examples. - Wait:
--waittails logs until terminal;--no-waitreturns as soon as the driver is running.
Other long-lived-only flags:
--replace— stop any running training / daemon job before submitting new ones (suffixes daemon submissionIds with a timestamp so Ray accepts the resubmit).--recreate— delete + re-apply a RayCluster whose live spec has drifted from the rendered manifest (default is warn + reuse).--skip-daemons— bring up all declared clusters but only submit training. Use on disagg recipes where gym/generation are already healthy.
Gotcha: on infra where the entrypoint does cd /opt/nemo-rl (or another in-image / Lustre path) and loads the recipe from there, --code-source upload does NOT override the recipe on the pod — the uploaded working_dir sits in /tmp/ray/... but the entrypoint cds away from it. To actually test a local recipe change, either sync your edits to the shared filesystem mounted into the pods or flip the Hydra overrides in the entrypoint.
4. Ephemeral mode flags (--rayjob)
When --rayjob is set, run branches into the RayJob code path. Relevant flags:
--rayjob-name NAME— RayJob metadata name (defaults to the training cluster name).--shutdown / --no-shutdown— defaulttrue: KubeRay deletes the RayCluster once the Ray Job reaches a terminal state.--ttl SECONDS— default 3600s: keep the RayJob object around after the run finishes for post-mortem log access.--wait / --no-wait— defaultwait: polljobDeploymentStatusuntil Complete/Failed.--no-waitreturns as soon as the RayJob is applied.--timeout SECONDS— default 86400s (24h): bound the--waitpoll.--dry-run— render the RayJob manifest and print it; do not apply.
--replace / --recreate / --skip-daemons are silently ignored in --rayjob mode (KubeRay owns lifecycle).
5. Iterating on a config without touching the shared filesystem
When the recipe on the pod filesystem has the wrong value for your experiment, use Hydra overrides on the entrypoint instead of forking the recipe. Pattern:
entrypoint: |
set -eu
cd /opt/nemo-rl
RUN_ID="\${RAY_JOB_SUBMISSION_ID:-\${NRL_K8S_RUN_ID:-$(date -u +%Y%m%d-%H%M%S)}}"
python -u examples/run_grpo.py \
--config infra/nrl_k8s/examples/<recipe>.yaml \
logger.wandb_enabled=true \
logger.wandb.project=<project> \
"logger.wandb.name=<run-name>-\${RUN_ID}"
Escape ${…} with a backslash. OmegaConf otherwise interprets it as interpolation and errors on shell-style ${VAR:-default}. RUN_ID resolves to RAY_JOB_SUBMISSION_ID (injected by KubeRay in rayjob mode) → NRL_K8S_RUN_ID (injected by the CLI in long-lived mode) → local timestamp — so the name is unique across either path.
6. Per-profile concerns (hardware + scheduler + DRA)
Every infra YAML encodes a hardware/scheduler profile. The concrete examples in infra/nrl_k8s/examples/ are authoritative for the profiles they target — read the neighbouring infra file before writing a new one. Things that commonly vary:
- Per-node GPUs (e.g. 4 vs 8) — must match
cluster.gpus_per_nodein the recipe, otherwise workers stayPending. - Node selectors — head pods usually land on a CPU-only node pool; GPU workers match on
nvidia.com/gpu.productor a node-group label. - Scheduler — KAI (
schedulerName: kai-scheduler+kai.scheduler/queuelabel) with topology annotations (kai.scheduler/topology,kai.scheduler/topology-required-placement) gang-schedules workers into one clique. Without it, pods may land on different racks and NVLink/RoCE won't span them. - DRA claims — ComputeDomain + RoCE are attached via
resourceClaimsreferencingResourceClaimTemplates. The CLI auto-creates/deletes these when the worker pod spec contains DRA claim references — no manual setup needed. - Secrets — always via
secretKeyRef(wandb-api-key, image pull secret). Never embed. - Shared filesystem mounts — typically a Lustre PVC mounted twice: once at the code path (e.g.
/opt/nemo-rlwith a user-scopedsubPath) and once at a workspace root (e.g./mnt/rl-workspace) for datasets, HF cache, and checkpoints.
Before applying an infra, verify prereqs exist in the target namespace:
kubectl get pvc <workspace-pvc>
kubectl get secret <wandb-secret> <image-pull-secret>
kubectl get sa <service-account>
7. End-to-end workflows
7a. Fresh one-shot run (rayjob)
# From the NeMo-RL repo root:
nrl-k8s check <recipe> --infra <infra> # validate first
nrl-k8s run <recipe> --infra <infra> --rayjob --dry-run # render RayJob manifest
nrl-k8s run <recipe> --infra <infra> --rayjob --no-wait # apply, returns fast
Watch status + teardown (works even after your laptop disconnects because KubeRay owns the lifecycle):
kubectl get rayjob -n default <name> -w
kubectl get raycluster -n default # empty = teardown succeeded
7b. Dev loop (long-lived)
nrl-k8s run <recipe> --infra <infra> --run-id $(date +%Y%m%d-%H%M%S)
# Edits in the recipe? Just re-run — reuses the live cluster.
# Pod spec changed? Add --recreate to delete + re-apply.
# Disagg recipe with gym/gen already healthy? --skip-daemons.
7c. First-time disaggregated bring-up
nrl-k8s run <recipe> --infra <disagg-infra> --mode batch --code-source image
7d. Cluster-only lifecycle
nrl-k8s cluster up <recipe> --infra <infra> --target kuberay.training --wait
nrl-k8s cluster up <recipe> --infra <infra> --target kuberay.training --dry-run # render manifest
nrl-k8s cluster down <recipe> --infra <infra> --target kuberay.training --wait
nrl-k8s cluster down <recipe> --infra <infra> # tear down all
nrl-k8s cluster list -n default
nrl-k8s cluster dashboard <cluster-name> # port-forward + browser
7e. Deployments (e.g. nemo-skills sandbox)
# Bring up just the deployment
nrl-k8s cluster up <recipe> --infra <infra> --target deployments.nemo_skills
# Tear down just the deployment
nrl-k8s cluster down <recipe> --infra <infra> --target deployments.nemo_skills
# Tear down everything (RayClusters + Deployments)
nrl-k8s cluster down <recipe> --infra <infra>
The deployments: section in infra YAML declares Kubernetes Deployments managed alongside RayClusters. The CLI patches image, imagePullSecrets, and serviceAccountName from the top-level infra keys (same as RayClusters). Deployments start in parallel with cluster bring-up — no ordering dependency.
8. Monitoring a run
# Status
nrl-k8s status <recipe> --infra <infra>
kubectl get rayjob,raycluster -n default
# Follow the driver
nrl-k8s job list <recipe> --infra <infra> --role training
nrl-k8s job logs <run-id> <recipe> --infra <infra> --role training -f
When the nrl-k8s job logs -f subprocess dies (kubectl port-forward i/o timeout after ~15 min idle), just re-run it. The training job keeps going.
To fetch driver logs fo
Truncated for display — read the full file on GitHub.
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
