nanogpt-training
Train GPT-2 scale models (~124M parameters) efficiently on a single GPU. Covers GPT-124M architecture, tokenized dataset loading (e.g., HuggingFace Hub shards), modern optimizers (Muon, AdamW), mixed precision training, and training loop implementation.
Install / Use
npx skills add benchflow-ai/skillsbench --skill nanogpt-trainingInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of nanogpt-training
nanogpt-training scores 88/100 on our quality scale, 392nd of 970 AI & Machine Learning skills we index (top 41%).
Its SKILL.md is 3.4 KB long, well organised into 9 sections with 3 code examples: a solid amount of guidance for an agent.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so nanogpt-training is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
nanogpt-training compared with similar skills
All 4 of these similar skills score higher than nanogpt-training; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| nanogpt-training (this skill)by benchflow-ai | 88 | 1.8k | 2mo ago | SKILL.md |
| claude-memby thedotmack | 100 | 95.2k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.0k | today | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install nanogpt-training?
- Run
npx skills add benchflow-ai/skillsbench --skill nanogpt-training. The install tabs above show the steps for each supported agent. - Which AI agents does nanogpt-training work with?
- It is written for Zed, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is nanogpt-training safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is nanogpt-training still maintained?
- The repository was last updated about 2 months ago, so nanogpt-training is actively maintained.
Skill content
View source on GitHubname: nanogpt-training description: Train GPT-2 scale models (~124M parameters) efficiently on a single GPU. Covers GPT-124M architecture, tokenized dataset loading (e.g., HuggingFace Hub shards), modern optimizers (Muon, AdamW), mixed precision training, and training loop implementation.
NanoGPT Training
Overview
Training GPT-2 scale models (~124M parameters) efficiently on a single GPU. It provides:
- GPT-124M Architecture: Standard transformer with RoPE and modern optimizations
- Tokenized Datasets: Loading pre-tokenized shards from HuggingFace Hub or local files
- Modern Optimizers: Muon optimizer with Newton-Schulz orthogonalization
- Mixed Precision: bfloat16 training on A100 for 2x speedup
Training options:
- Baseline GPT: Standard residual connections
- Experimental residual variants: Optional alternative residual schemes for stability/efficiency
Quick Reference
| Topic | Reference | |-------|-----------| | Model Architecture | GPT Architecture | | Data Loading | Tokenized Data | | Optimizers | Optimizers | | Training Loop | Training Loop | | Hyperparameters | Hyperparameters |
Installation
pip install torch einops numpy huggingface_hub
Minimal Example
import modal
app = modal.App("gpt-training")
image = modal.Image.debian_slim(python_version="3.11").pip_install(
"torch", "einops", "numpy", "huggingface_hub"
)
@app.function(gpu="A100", image=image, timeout=3600)
def train():
import torch
from dataclasses import dataclass
@dataclass
class GPTConfig:
block_size: int = 1024
vocab_size: int = 50257
n_layer: int = 12
n_head: int = 12
n_embd: int = 768
dropout: float = 0.0
bias: bool = False
# Download data, build model, train
# ... (see references for full implementation)
return {"final_loss": final_loss}
@app.local_entrypoint()
def main():
results = train.remote()
print(results)
Common Imports
import torch
import torch.nn as nn
import torch.nn.functional as F
from torch.cuda.amp import autocast, GradScaler
from dataclasses import dataclass
from einops import rearrange, repeat, reduce
import numpy as np
import math
When to Use What
| Scenario | Approach | |----------|----------| | Standard GPT training | Use baseline model with standard residuals | | Stability experiments | Try alternative residual variants or extra streams | | Small experiments | Use T4/A10G GPU | | Full training | Use A100 with bfloat16 | | Custom data | Modify the dataset loader class | | Different model size | Adjust GPTConfig parameters |
Metrics to Monitor
| Metric | Typical Signal | Notes | |--------|----------------|-------| | Validation loss | Steady decrease | Absolute value depends on dataset/tokenizer | | Grad norm | Moderate, stable range | Large spikes indicate instability | | Training stability | Smooth curves | Frequent spikes suggest LR/batch issues | | Throughput | Consistent tokens/sec | Use for comparing configs |
External Resources
- nanoGPT: https://github.com/karpathy/nanoGPT
- build-nanogpt: https://github.com/karpathy/build-nanogpt
- modded-nanogpt: https://github.com/KellerJordan/modded-nanogpt
- FineWeb-Edu token shards: https://huggingface.co/datasets/karpathy/fineweb-edu-100B-gpt2-token-shards
Related Skills
claude-mem
95.2kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Understand-Anything
85.0kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.2kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
