SkillAgentSearch skills...

nemo-automodel-model-onboarding

Guide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation.

Install / Use

npx skills add NVIDIA/skills --skill nemo-automodel-model-onboarding

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

95/100

Supported Platforms

Universal

Tags

Our assessment of nemo-automodel-model-onboarding

nemo-automodel-model-onboarding scores 95/100 on our quality scale, 112th of 821 AI & Machine Learning skills we index (top 14%).

Its SKILL.md is 22 KB long, well organised into 33 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.

With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
15/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 5 days ago, so nemo-automodel-model-onboarding is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-09-29. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

nemo-automodel-model-onboarding compared with similar skills

All 4 of these similar skills score higher than nemo-automodel-model-onboarding; compare them before choosing.

SkillScoreStarsUpdatedFormat
nemo-automodel-model-onboarding (this skill)by NVIDIA953.4k5d agoSKILL.md
claude-memby thedotmack10094.9ktodayCLAUDE.md
Understand-Anythingby Egonex-AI10084.5ktodayCLAUDE.md
headroomby headroomlabs-ai10074.0ktodayCLAUDE.md
CowAgentby zhayujie10047.2ktodayCLAUDE.md

Frequently asked questions

How do I install nemo-automodel-model-onboarding?
Run npx skills add NVIDIA/skills --skill nemo-automodel-model-onboarding. The install tabs above show the steps for each supported agent.
Which AI agents does nemo-automodel-model-onboarding work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is nemo-automodel-model-onboarding safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is nemo-automodel-model-onboarding still maintained?
The repository was last updated 5 days ago, so nemo-automodel-model-onboarding is actively maintained.

name: nemo-automodel-model-onboarding description: Guide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation. when_to_use: Adding or modifying model architecture support in NeMo AutoModel, such as LLM/VLM/MoE model files, custom layers, state-dict adapters, registry entries, Hugging Face config mapping, or capability flags. license: Apache-2.0 metadata: author: NVIDIA tags: - nemo-automodel - model-onboarding

Adding Model Support to NeMo AutoModel

Purpose

This skill guides implementation of new model architectures in NeMo AutoModel. Follow the five phases in order.

<!-- NVSkills signature refresh requested after PR #2998 (2026-07-31). -->

Instructions

When answering an onboarding question, keep the response in this order:

  1. Classify the architecture from config.json.
  2. Name the exact implementation files under components/models/<name>/.
  3. Identify registry and optional custom-config updates.
  4. State the validation tests that must be added before full checkpoint use.

For conceptual onboarding questions, answer from this skill without opening the pattern files unless the user asks you to edit code. Mention pattern filenames as references, then give the direct checklist.

Use direct action verbs: classify the model, name the files, map the weights, register the class, and add tests. Do not discuss distributed strategy, launcher configuration, or general recipe authoring unless the user explicitly connects it to onboarding a new architecture.

Examples

Use these compact answer patterns for common questions:

  • Dense causal LM: classify as dense only when architectures contains a ForCausalLM class and expert fields such as num_local_experts, n_routed_experts, or num_experts_per_tok are absent. Create components/models/<name>/model.py and __init__.py; add state_dict_adapter.py only for checkpoint weight conversion and config.py only if needed. Register MODEL_ARCH_MAPPING in _transformers/registry.py, add example YAML, and add tiny-config unit tests plus layer-equivalence tests for rewritten layers.
  • MoE state dict: identify expert fields in config.json, reference moe-patterns.md, map router tensors separately, preserve routed-expert index order, map routed experts, shared experts, and gate/up/down projections, add adapter key-map tests and tiny-config numerical equivalence tests, and do not rely only on from_pretrained() or silent tensor reshapes.
  • VLM onboarding: classify as VLM only when vision_config, text_config, and a ForConditionalGeneration architecture are present. Reference vlm-patterns.md and existing VLM implementations such as mistral4, kimivl, or kimi_k25_vl; check text backbone, vision tower, projector, processor assumptions, text and vision checkpoint compatibility (adapter mappings when needed), registry registration, and tiny image-text tests before full checkpoints. Do not treat VLM onboarding as a pure causal-LM path or skip processor/image tests.

For MoE state-dict and VLM questions, apply the checklists in Sections 2.4 and 2.5.

Routing Boundary

Use this skill only when the user is adding or modifying model architecture support: model files, custom layers, state-dict adapters, Hugging Face config mapping, registry entries, or model capability flags.

Do not use this skill for standalone training recipe YAML questions about optimizers, datasets, schedulers, validation datasets, or trainer wiring unless they are explicitly part of onboarding a new model architecture. Those recipe questions belong to the nemo-automodel-recipe-development skill.

In-scope examples:

  • "Add support for a new Hugging Face causal LM architecture."
  • "Map MoE router and expert weights from a Hugging Face checkpoint."
  • "Register a new model class in NeMo AutoModel."

Out-of-scope examples:

  • "Write a finetuning recipe YAML with optimizer and dataset sections."
  • "Choose FSDP2, DDP, tensor parallel, or context parallel settings."
  • "Configure Slurm, SkyPilot, containers, mounts, or launch dispatch."

Phase 1: Discovery

Before writing code, gather information about the target model.

1.1 Fetch HuggingFace config.json

Download the model's config.json from the HuggingFace Hub (or use AutoConfig.from_pretrained). Key fields to extract:

  • architectures -- determines the class name and registration key (e.g., "LlamaForCausalLM", "Qwen3MoeForCausalLM", "Mistral3ForConditionalGeneration")
  • model_type -- used for custom config registration in _CUSTOM_CONFIG_REGISTRATIONS if HF does not have a built-in config class
  • hidden_size, intermediate_size, num_hidden_layers, num_attention_heads, num_key_value_heads -- sizing
  • vocab_size -- needed for tiny test configs
  • tie_word_embeddings -- the saved setting in each supported checkpoint; do not infer it from a bare config constructor
  • hidden_act -- activation function (e.g., "silu" for SwiGLU)

1.2 Determine model type

| Type | Indicators | Pattern file | |------|-----------|-------------| | Dense LLM | ForCausalLM in architectures, no expert fields | llm-patterns.md | | MoE LLM | n_routed_experts, num_local_experts, num_experts_per_tok in config | moe-patterns.md | | VLM | ForConditionalGeneration in architectures, has vision_config + text_config | vlm-patterns.md |

1.3 Check for existing similar architectures

Look in components/models/ for architectures with similar attention or MLP patterns:

components/models/
  llama/           # Standard GQA + SwiGLU with separate HF-compatible projections
  qwen2/           # Same as Llama but with attention bias + QKV bias
  baichuan/        # ALiBi attention variant
  deepseek_v3/     # MLA attention + MoE (DeepSeek-style grouped experts)
  mistral4/        # MLA + MoE + VLM (Pixtral vision)
  kimivl/          # DeepSeek-V3 backbone + MoonVit vision
  kimi_k25_vl/     # Updated KimiVL with different projector
  qwen3_moe/       # Qwen3 with MoE layers
  nemotron_v3/     # Hybrid mamba-attention

1.4 Identify custom components

Check whether the model needs:

  • Custom attention: GQA (standard), MLA (DeepSeek/Mistral4), sliding window, bidirectional
  • Custom RoPE: Standard (Llama), YaRN scaling, NTK-aware, complex-number (DeepSeek)
  • Custom normalization: RMSNorm (standard), LayerNorm, different eps values
  • Custom MLP: SwiGLU (standard), GeGLU, ReLU-squared, MoE routing
  • Custom config class: Needed only if HF AutoConfig cannot parse the model's config.json (check auto_map field)

1.5 Note dimensions for test config

For unit tests, create a tiny config. Target: ~1M parameters or less.

# Example tiny config for a Llama-like model:
tiny_config = LlamaConfig(
    hidden_size=64,
    intermediate_size=128,
    num_hidden_layers=2,
    num_attention_heads=4,
    num_key_value_heads=2,
    vocab_size=256,
    max_position_embeddings=128,
)

Phase 2: Implementation

2.1 Create directory structure

components/models/<name>/
  __init__.py
  model.py
  state_dict_adapter.py # Only if HF weight names or tensor layouts need conversion
  config.py            # Only if HF config is insufficient
  layers.py            # Only for MoE / MLA / other non-standard layers
  rope_utils.py        # Only for custom RoPE

2.2 Implementation order

Implement files in dependency order:

  1. config.py (if needed) -- Custom PretrainedConfig subclass
  2. rope_utils.py (if needed) -- RoPE implementation
  3. layers.py (if needed) -- Attention, MLP, decoder block classes
  4. model.py -- The main ForCausalLM (or ForConditionalGeneration) class
  5. state_dict_adapter.py (if needed) -- HF weight conversion. Treat checkpoint I/O performance as part of the implementation, and evaluate the low-memory DCP capability as described in Section 2.6.
  6. init.py -- Re-export the main model class

See the pattern files for detailed implementation guidance:

Most custom models need state_dict_adapter.py for HF weight conversion. Omit the file and attribute only when HF names and tensor layouts already match across supported backend/config variants, as in Llama, Qwen2, and Qwen3. Weight tying remains the model's responsibility (Section 2.3).

2.3 Causal LM weight tying

Every registered model class with a causal lm_head must:

  • Declare tie_word_embeddings_support: TieSupport as BOTH, TIED_ONLY, or UNTIED_ONLY.
  • Call reject_unsupported_tie_word_embeddings(type(self), config) at the top of __init__, using the original config before unwrapping text_config or thinker_config.

Only classes with no causal LM head may be explicitly exempted from the registry test.

Choose the policy from the implementation and the actual supported checkpoint configs, not from a bare config constructor:

  • BOTH: tied and untied configurations are both supported.
  • TIED_ONLY: only a tied configuration is supported.
  • UNTIED_ONLY: only an untied configuration is supported.

Runtime helpers must treat TIED_ONLY and UNTIED_ONLY as authoritative and only resolve a per-checkpoint config flag for BOTH. All current BOTH VLMs honor the outer tie_word_embeddings flag, so do not add a model-specific resolver until a supported BOTH model actually requires another config path.

For BOTH and TIED_ONLY, always declare _tied_weights_keys and implement tie_weights() with the actual lm_head and input-embedding FQNs. Do not rely on inherited Hugging Face tying, and re-tie after any language-model swap.

Add policy-specific tests:

  • BOTH: tied aliases; untied does not alias.
  • TIED_ONLY: tied aliases; untied is rejected.
  • UNTIED_ONLY: weights stay separate; tied is rejected.

Do not tie architectures with intentionally separate heads, asymmetric vocab sizes, or stages that do not own both tensors.

For from_pretrained, the checkpoint's saved tie_word_embeddings value is authoritative, even for BOTH. The NeMoAuto* bridge rejects flips in either direction. A model-owned from_pretrained that bypasses that bridge must call reject_tie_word_embeddings_flip(checkpoint_config, requested_config, model_class_name).

2.4 MoE state-dict adapter checklist

For MoE models, verify all weights below. When their HF and native layouts differ, the adapter must explicitly map:

  • Router weights, including gate bias or correction-bias tensors when the Hugging Face model has them.
  • Expert weights, preserving expert index order across local and routed experts.
  • Gate/up/down projections, including combined or split projection layouts.
  • Shared experts separately from routed experts when the architecture has both.

Add tests that assert expected key mappings and run numerical equivalence with tiny configs before trying full checkpoints.

Do not use these shortcuts:

  • Do not validate the adapter only by calling from_pretrained().
  • Do not accept missing or extra expert keys without an explicit mapping reason.
  • Do not change dtype, transpose dimensions, or reshape tensors unless the HF and NeMo layouts require it and a test proves the conversion is reversible.
  • Do not skip router or shared-expert tests because dense-layer tests pass.

2.5 VLM onboarding checklist

For VLMs, confirm the Hugging Face config has vision_config and text_config and that architectures points to a conditional-generation class. Start from the closest VLM pattern file, usually vlm-patterns.md, and compare existi

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3.4k
CategoryAI
Updated5d ago
Forks412

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions