finetuning
Fine-tune models on Microsoft Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation.
Install / Use
npx skills add microsoft/skills --skill finetuningInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Tags
Our assessment of finetuning
finetuning scores 86/100 on our quality scale, 291st of 508 Operations skills we index.
Its SKILL.md is 5.4 KB long, well organised into 8 sections and no code examples: a solid amount of guidance for an agent.
With 3,051 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so finetuning is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
finetuning compared with similar skills
All 4 of these similar skills score higher than finetuning; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| finetuning (this skill)by microsoft | 86 | 3.1k | 5d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 7d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 7d ago | SKILL.md |
Frequently asked questions
- How do I install finetuning?
- Run
npx skills add microsoft/skills --skill finetuning. The install tabs above show the steps for each supported agent. - Which AI agents does finetuning work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is finetuning safe to use?
- It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is finetuning still maintained?
- The repository was last updated 5 days ago, so finetuning is actively maintained.
Skill content
View source on GitHubname: finetuning description: "Fine-tune models on Microsoft Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation. USE FOR: fine-tune, SFT, DPO, RFT, training data, grader, distillation, fine-tuned model, training job, large file upload, calibrate grader, deploy fine-tuned model, evaluate fine-tuned model. DO NOT USE FOR: general model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer)." license: MIT metadata: author: Microsoft version: "0.0.0-placeholder"
Fine-Tuning on Microsoft Foundry
Fine-tune models using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset prep, training, deployment, and evaluation.
When to Use
Use this sub-skill when the user asks about:
- Fine-tuning a model (SFT, DPO, or RFT)
- Preparing, validating, or formatting training data
- Submitting, monitoring, or diagnosing training jobs
- Calibrating graders or pass thresholds for RFT
- Deploying or evaluating a fine-tuned model
- Choosing between training types (SFT vs DPO vs RFT)
- Distillation, synthetic data generation, or dataset quality scoring
- Large file uploads for training data
- Cleaning up fine-tuning resources (files, deployments)
Do NOT use for: General model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).
Workflows
| Stage | Guide | |-------|-------| | Quick start | workflows/quickstart.md | | Full pipeline | workflows/full-pipeline.md | | Create data | workflows/dataset-creation.md | | Iterate | workflows/iterative-training.md | | Diagnose | workflows/diagnose-poor-results.md |
References
| Topic | File | |-------|------| | SFT vs DPO vs RFT | references/training-types.md | | Hyperparameters | references/hyperparameters.md | | Data formats | references/dataset-formats.md | | Grader design (RFT) | references/grader-design.md | | Reward hacking | references/reward-hacking.md | | Agentic RFT (tools) | references/agentic-rft.md | | Deployment | references/deployment.md | | Training curves | references/training-curves.md | | Evaluation | references/evaluation.md | | Vision fine-tuning | references/vision-fine-tuning.md | | Large file uploads | references/large-file-uploads.md | | Platform gotchas | references/platform-gotchas.md |
Scripts
| Script | Purpose |
|--------|---------|
| scripts/submit_training.py | Submit SFT/DPO/RFT jobs |
| scripts/monitor_training.py | Poll job until completion |
| scripts/calibrate_grader.py | Find optimal RFT pass_threshold |
| scripts/check_training.py | Analyze curves, list checkpoints |
| scripts/deploy_model.py | Deploy via ARM REST API |
| scripts/evaluate_model.py | LLM judge evaluation |
| scripts/convert_dataset.py | Convert between SFT/DPO/RFT formats |
| scripts/generate_distillation_data.py | Generate synthetic training data |
| scripts/score_dataset.py | Quality scoring on training data |
| scripts/cleanup.py | Delete old files and deployments |
| scripts/validate/ | Data validators (SFT, DPO, RFT) + stats |
Rules
- Always baseline first — evaluate the base model before fine-tuning
- Validate data before submitting — run
scripts/validate/validate_sft.py - Calibrate RFT graders — target 25-50% failure rate on the base model
- Evaluate checkpoints — don't blindly deploy the final one
- Measure token cost alongside accuracy when comparing models
Quick Reference
| Task | Command |
|------|---------|
| Validate SFT data | python scripts/validate/validate_sft.py data.jsonl |
| Submit SFT job | python scripts/submit_training.py --model gpt-4.1-mini --training-file train.jsonl --validation-file val.jsonl --type sft |
| Monitor job | python scripts/monitor_training.py --job-id ftjob-xxx |
| Analyze curves | python scripts/check_training.py --job-id ftjob-xxx |
| Deploy model | python scripts/deploy_model.py --model-id ft:gpt-4.1-mini:... --name my-eval |
| Evaluate model | python scripts/evaluate_model.py --deployment-name my-eval --test-file test.jsonl |
Error Handling
| Error | Cause | Fix |
|-------|-------|-----|
| "API version not supported" | Older openai SDK on /v1/ endpoint | Upgrade to openai>=1.0 |
| "does not support fine-tuning with Standard TrainingType" | OSS model needs globalStandard | Use --use-rest flag or script auto-falls back |
| Job stuck in post-training eval | Under-provisioned tool endpoint (RFT) | Scale to S2+, enable Always On |
| "DeploymentNotReady" after ARM succeeds | ARM/data-plane race condition | Delete and recreate deployment, wait 5 min |
| Content safety block at deployment | PII-dense training data | Remove problematic document types |
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
