nemo-automodel-model-onboarding
Guide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation.
Install / Use
npx skills add NVIDIA/skills --skill nemo-automodel-model-onboardingInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of nemo-automodel-model-onboarding
nemo-automodel-model-onboarding scores 95/100 on our quality scale, 112th of 821 AI & Machine Learning skills we index (top 14%).
Its SKILL.md is 22 KB long, well organised into 33 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so nemo-automodel-model-onboarding is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-29. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
nemo-automodel-model-onboarding compared with similar skills
All 4 of these similar skills score higher than nemo-automodel-model-onboarding; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| nemo-automodel-model-onboarding (this skill)by NVIDIA | 95 | 3.4k | 5d ago | SKILL.md |
| claude-memby thedotmack | 100 | 94.9k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 84.5k | today | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install nemo-automodel-model-onboarding?
- Run
npx skills add NVIDIA/skills --skill nemo-automodel-model-onboarding. The install tabs above show the steps for each supported agent. - Which AI agents does nemo-automodel-model-onboarding work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is nemo-automodel-model-onboarding safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is nemo-automodel-model-onboarding still maintained?
- The repository was last updated 5 days ago, so nemo-automodel-model-onboarding is actively maintained.
Skill content
View source on GitHubname: nemo-automodel-model-onboarding description: Guide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation. when_to_use: Adding or modifying model architecture support in NeMo AutoModel, such as LLM/VLM/MoE model files, custom layers, state-dict adapters, registry entries, Hugging Face config mapping, or capability flags. license: Apache-2.0 metadata: author: NVIDIA tags: - nemo-automodel - model-onboarding
Adding Model Support to NeMo AutoModel
Purpose
This skill guides implementation of new model architectures in NeMo AutoModel. Follow the five phases in order.
<!-- NVSkills signature refresh requested after PR #2998 (2026-07-31). -->Instructions
When answering an onboarding question, keep the response in this order:
- Classify the architecture from
config.json. - Name the exact implementation files under
components/models/<name>/. - Identify registry and optional custom-config updates.
- State the validation tests that must be added before full checkpoint use.
For conceptual onboarding questions, answer from this skill without opening the pattern files unless the user asks you to edit code. Mention pattern filenames as references, then give the direct checklist.
Use direct action verbs: classify the model, name the files, map the weights, register the class, and add tests. Do not discuss distributed strategy, launcher configuration, or general recipe authoring unless the user explicitly connects it to onboarding a new architecture.
Examples
Use these compact answer patterns for common questions:
- Dense causal LM: classify as dense only when
architecturescontains aForCausalLMclass and expert fields such asnum_local_experts,n_routed_experts, ornum_experts_per_tokare absent. Createcomponents/models/<name>/model.pyand__init__.py; addstate_dict_adapter.pyonly for checkpoint weight conversion andconfig.pyonly if needed. RegisterMODEL_ARCH_MAPPINGin_transformers/registry.py, add example YAML, and add tiny-config unit tests plus layer-equivalence tests for rewritten layers. - MoE state dict: identify expert fields in
config.json, referencemoe-patterns.md, map router tensors separately, preserve routed-expert index order, map routed experts, shared experts, and gate/up/down projections, add adapter key-map tests and tiny-config numerical equivalence tests, and do not rely only onfrom_pretrained()or silent tensor reshapes. - VLM onboarding: classify as VLM only when
vision_config,text_config, and aForConditionalGenerationarchitecture are present. Referencevlm-patterns.mdand existing VLM implementations such asmistral4,kimivl, orkimi_k25_vl; check text backbone, vision tower, projector, processor assumptions, text and vision checkpoint compatibility (adapter mappings when needed), registry registration, and tiny image-text tests before full checkpoints. Do not treat VLM onboarding as a pure causal-LM path or skip processor/image tests.
For MoE state-dict and VLM questions, apply the checklists in Sections 2.4 and 2.5.
Routing Boundary
Use this skill only when the user is adding or modifying model architecture support: model files, custom layers, state-dict adapters, Hugging Face config mapping, registry entries, or model capability flags.
Do not use this skill for standalone training recipe YAML questions about optimizers, datasets, schedulers, validation datasets, or trainer wiring unless they are explicitly part of onboarding a new model architecture. Those recipe questions belong to the nemo-automodel-recipe-development skill.
In-scope examples:
- "Add support for a new Hugging Face causal LM architecture."
- "Map MoE router and expert weights from a Hugging Face checkpoint."
- "Register a new model class in NeMo AutoModel."
Out-of-scope examples:
- "Write a finetuning recipe YAML with optimizer and dataset sections."
- "Choose FSDP2, DDP, tensor parallel, or context parallel settings."
- "Configure Slurm, SkyPilot, containers, mounts, or launch dispatch."
Phase 1: Discovery
Before writing code, gather information about the target model.
1.1 Fetch HuggingFace config.json
Download the model's config.json from the HuggingFace Hub (or use AutoConfig.from_pretrained). Key fields to extract:
architectures-- determines the class name and registration key (e.g.,"LlamaForCausalLM","Qwen3MoeForCausalLM","Mistral3ForConditionalGeneration")model_type-- used for custom config registration in_CUSTOM_CONFIG_REGISTRATIONSif HF does not have a built-in config classhidden_size,intermediate_size,num_hidden_layers,num_attention_heads,num_key_value_heads-- sizingvocab_size-- needed for tiny test configstie_word_embeddings-- the saved setting in each supported checkpoint; do not infer it from a bare config constructorhidden_act-- activation function (e.g.,"silu"for SwiGLU)
1.2 Determine model type
| Type | Indicators | Pattern file |
|------|-----------|-------------|
| Dense LLM | ForCausalLM in architectures, no expert fields | llm-patterns.md |
| MoE LLM | n_routed_experts, num_local_experts, num_experts_per_tok in config | moe-patterns.md |
| VLM | ForConditionalGeneration in architectures, has vision_config + text_config | vlm-patterns.md |
1.3 Check for existing similar architectures
Look in components/models/ for architectures with similar attention or MLP patterns:
components/models/
llama/ # Standard GQA + SwiGLU with separate HF-compatible projections
qwen2/ # Same as Llama but with attention bias + QKV bias
baichuan/ # ALiBi attention variant
deepseek_v3/ # MLA attention + MoE (DeepSeek-style grouped experts)
mistral4/ # MLA + MoE + VLM (Pixtral vision)
kimivl/ # DeepSeek-V3 backbone + MoonVit vision
kimi_k25_vl/ # Updated KimiVL with different projector
qwen3_moe/ # Qwen3 with MoE layers
nemotron_v3/ # Hybrid mamba-attention
1.4 Identify custom components
Check whether the model needs:
- Custom attention: GQA (standard), MLA (DeepSeek/Mistral4), sliding window, bidirectional
- Custom RoPE: Standard (Llama), YaRN scaling, NTK-aware, complex-number (DeepSeek)
- Custom normalization: RMSNorm (standard), LayerNorm, different eps values
- Custom MLP: SwiGLU (standard), GeGLU, ReLU-squared, MoE routing
- Custom config class: Needed only if HF
AutoConfigcannot parse the model'sconfig.json(checkauto_mapfield)
1.5 Note dimensions for test config
For unit tests, create a tiny config. Target: ~1M parameters or less.
# Example tiny config for a Llama-like model:
tiny_config = LlamaConfig(
hidden_size=64,
intermediate_size=128,
num_hidden_layers=2,
num_attention_heads=4,
num_key_value_heads=2,
vocab_size=256,
max_position_embeddings=128,
)
Phase 2: Implementation
2.1 Create directory structure
components/models/<name>/
__init__.py
model.py
state_dict_adapter.py # Only if HF weight names or tensor layouts need conversion
config.py # Only if HF config is insufficient
layers.py # Only for MoE / MLA / other non-standard layers
rope_utils.py # Only for custom RoPE
2.2 Implementation order
Implement files in dependency order:
- config.py (if needed) -- Custom
PretrainedConfigsubclass - rope_utils.py (if needed) -- RoPE implementation
- layers.py (if needed) -- Attention, MLP, decoder block classes
- model.py -- The main
ForCausalLM(orForConditionalGeneration) class - state_dict_adapter.py (if needed) -- HF weight conversion. Treat checkpoint I/O performance as part of the implementation, and evaluate the low-memory DCP capability as described in Section 2.6.
- init.py -- Re-export the main model class
See the pattern files for detailed implementation guidance:
- Dense LLM: llm-patterns.md
- MoE: moe-patterns.md
- VLM: vlm-patterns.md
- Capabilities and fp32 precision: capabilities-and-precision.md
Most custom models need state_dict_adapter.py for HF weight conversion.
Omit the file and attribute only when HF names and tensor layouts already match
across supported backend/config variants, as in Llama, Qwen2, and Qwen3.
Weight tying remains the model's responsibility (Section 2.3).
2.3 Causal LM weight tying
Every registered model class with a causal lm_head must:
- Declare
tie_word_embeddings_support: TieSupportasBOTH,TIED_ONLY, orUNTIED_ONLY. - Call
reject_unsupported_tie_word_embeddings(type(self), config)at the top of__init__, using the original config before unwrappingtext_configorthinker_config.
Only classes with no causal LM head may be explicitly exempted from the registry test.
Choose the policy from the implementation and the actual supported checkpoint configs, not from a bare config constructor:
BOTH: tied and untied configurations are both supported.TIED_ONLY: only a tied configuration is supported.UNTIED_ONLY: only an untied configuration is supported.
Runtime helpers must treat TIED_ONLY and UNTIED_ONLY as authoritative and
only resolve a per-checkpoint config flag for BOTH. All current BOTH VLMs
honor the outer tie_word_embeddings flag, so do not add a model-specific
resolver until a supported BOTH model actually requires another config path.
For BOTH and TIED_ONLY, always declare _tied_weights_keys and implement
tie_weights() with the actual lm_head and input-embedding FQNs. Do not rely
on inherited Hugging Face tying, and re-tie after any language-model swap.
Add policy-specific tests:
BOTH: tied aliases; untied does not alias.TIED_ONLY: tied aliases; untied is rejected.UNTIED_ONLY: weights stay separate; tied is rejected.
Do not tie architectures with intentionally separate heads, asymmetric vocab sizes, or stages that do not own both tensors.
For from_pretrained, the checkpoint's saved tie_word_embeddings value is
authoritative, even for BOTH. The NeMoAuto* bridge rejects flips in either
direction. A model-owned from_pretrained that bypasses that bridge must call
reject_tie_word_embeddings_flip(checkpoint_config, requested_config, model_class_name).
2.4 MoE state-dict adapter checklist
For MoE models, verify all weights below. When their HF and native layouts differ, the adapter must explicitly map:
- Router weights, including gate bias or correction-bias tensors when the Hugging Face model has them.
- Expert weights, preserving expert index order across local and routed experts.
- Gate/up/down projections, including combined or split projection layouts.
- Shared experts separately from routed experts when the architecture has both.
Add tests that assert expected key mappings and run numerical equivalence with tiny configs before trying full checkpoints.
Do not use these shortcuts:
- Do not validate the adapter only by calling
from_pretrained(). - Do not accept missing or extra expert keys without an explicit mapping reason.
- Do not change dtype, transpose dimensions, or reshape tensors unless the HF and NeMo layouts require it and a test proves the conversion is reversible.
- Do not skip router or shared-expert tests because dense-layer tests pass.
2.5 VLM onboarding checklist
For VLMs, confirm the Hugging Face config has vision_config and text_config
and that architectures points to a conditional-generation class. Start from
the closest VLM pattern file, usually vlm-patterns.md, and
compare existi
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
94.9kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Understand-Anything
84.5kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.2kOpen-source super AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
