doca-hardware-safety
Use this skill whenever the agent is about to recommend or apply a change that touches DPU / NIC hardware state on a live system — mlxconfig firmware-parameter write, NIC firmware burn, BFB reflash, NIC ↔ DPU mode flip, SR-IOV or device-emulation slot enable, kernel boot-parameter change (IOMMU, hug…
Install / Use
npx skills add NVIDIA/skills --skill doca-hardware-safetyInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Tags
Our assessment of doca-hardware-safety
doca-hardware-safety scores 88/100 on our quality scale, 970th of 3,356 Development & Engineering skills we index (top 29%).
Its SKILL.md is 17 KB long, well organised into 8 sections and no code examples: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so doca-hardware-safety is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
doca-hardware-safety compared with similar skills
All 4 of these similar skills score higher than doca-hardware-safety; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| doca-hardware-safety (this skill)by NVIDIA | 88 | 3.4k | 5d ago | SKILL.md |
| ai-job-searchby MadsLorentzen | 100 | 44.4k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 2d ago | CLAUDE.md |
| algorithmic-artby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
Frequently asked questions
- How do I install doca-hardware-safety?
- Run
npx skills add NVIDIA/skills --skill doca-hardware-safety. The install tabs above show the steps for each supported agent. - Which AI agents does doca-hardware-safety work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is doca-hardware-safety safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is doca-hardware-safety still maintained?
- The repository was last updated 5 days ago, so doca-hardware-safety is actively maintained.
Skill content
View source on GitHublicense: Apache-2.0 name: doca-hardware-safety description: > Use this skill whenever the agent is about to recommend or apply a change that touches DPU / NIC hardware state on a live system — mlxconfig firmware-parameter write, NIC firmware burn, BFB reflash, NIC ↔ DPU mode flip, SR-IOV or device-emulation slot enable, kernel boot-parameter change (IOMMU, hugepages, VFIO), PCIe rebind / rescan / link-state flip, or BlueField cold reboot. Wraps the change in pre-flight inventory, OOB reachability, a maintenance window, the mlxconfig cold-power-cycle rule, replica rehearsal, and rollback. Trigger even when the user does not say "hardware safety" — implicit phrasings: "flip BlueField mode over SSH", "enable SR-IOV and reboot", "burned firmware but mlxconfig shows old value", "reflashed BFB and lost representors", "reflash during business hours", "vendor says this is one-way". Refuse for general DOCA orientation (doca-public-knowledge-map), install or env debug (doca-setup), and program-side debug (doca-debug, doca-programming-guide) — those belong to other skills. metadata: kind: library compatibility: > No DOCA install required to read this skill (it is an overlay loaded against any DOCA artifact skill); the validation steps within DO require a live DOCA install at /opt/mellanox/doca with a BlueField DPU or ConnectX NIC, plus out-of-band console reachability (BMC, RShim, or operator-managed console) for any link-breaking change.
DOCA hardware safety
Where to start: This skill is the bundle's single source of truth
for the discipline that wraps every change touching DPU / NIC hardware
state on a live system. Open
TASKS.md when the operator is about to apply a
hardware-touching change and needs the change-application discipline
(pre-flight inventory → out-of-band path → window → apply → verify
→ rollback). Open CAPABILITIES.md when the
question is what does hardware-safety even cover (the class of
changes in scope, the failure modes the policy prevents, the
observability surface that gates a change, and the meta-policy that
every per-artifact ## Safety policy overlays).
Every per-artifact skill (services, libraries, tools) in the bundle
that recommends a hardware-touching action overlays this meta-policy
with artifact-specific safety. The per-artifact ## Safety policy
anchors do NOT redefine the cross-cutting discipline — they layer the
artifact's own concerns on top of it. This skill is the layer they all
build on.
Example questions this skill answers well
The CLASSES of hardware-safety questions this skill is built to answer, each with one worked example. The agent should treat the class as load-bearing — the worked example is a single instance.
- "I'm about to apply a hardware-touching change. What do I have to
capture before I touch anything?" — worked example: "the
per-artifact skill told me to flip a firmware-level emulation slot;
what do I capture first?". Answered by the pre-flight inventory
in
TASKS.md ## configureplus the inventory taxonomy inCAPABILITIES.md ## Capabilities and modes. - "This change might drop the link I'm using to manage the
BlueField. Is that safe?" — worked example: "I'm about to flip
the BlueField between NIC and DPU mode over the same management
link". Answered by the out-of-band access rule in
CAPABILITIES.md ## Safety policyplus the OOB-precondition gate inTASKS.md ## configure. - "The per-artifact skill said to write an
mlxconfigparameter, then reboot. Is that the right sequence?" — worked example: "the storage-emulation skill told me to enable a firmware slot viamlxconfigand then warm-reboot to apply it". Answered by themlxconfig-class rule inCAPABILITIES.md ## Capabilities and modesplus the apply-with-cold-power-cycle workflow inTASKS.md ## modify. - "My deployment plan reflashes the BlueField BFB during business
hours. Is that OK?" — worked example: "I have a one-hour
window during the day; can I reflash now?". Answered by the
maintenance-window discipline in
CAPABILITIES.md ## Safety policyplus the firmware-burn workflow inTASKS.md ## modify. - "How do I prove the change works before I touch production?"
— worked example: "the change is small; can I skip the lab
replica". Answered by the replica-first rule in
TASKS.md ## testplus the pre-hardware-validation pattern inCAPABILITIES.md ## Capabilities and modes. - "How do I roll back if this change goes wrong?" — worked
example: "I just reflashed the BFB and the host can't see the
representors anymore". Answered by the rollback ladder in
TASKS.md ## debugplus the rollback-must-be-documented rule inCAPABILITIES.md ## Safety policy. - "This change doesn't have a documented rollback. Should I still
apply it?" — worked example: "the vendor says this firmware
rev is one-way". Answered by the refuse-and-escalate rule in
CAPABILITIES.md ## Safety policyplus the escalation path inTASKS.md ## debug.
When to load this skill
Load this skill whenever the agent is about to recommend, or is
helping the operator apply, a change that touches DPU / NIC hardware
state on a live system. The decision must be made before the
agent composes its first sentence — the activation checklist below
is the same one referenced from
AGENTS.md ## Cross-cutting overlay activation triggers,
mirrored here so a per-artifact skill that already loaded this skill
has the activation rule at hand.
Agent activation checklist — load this skill at the START of the answer when any cell below is true
| Trigger class | Concrete prompt-side signals (any one fires the overlay) |
| --- | --- |
| mlxconfig-class change | the prompt or the agent's next recommended action mentions mlxconfig directly; OR toggles BlueField between NIC / DPU / Separated-Host mode; OR enables SR-IOV; OR enables device-emulation slot (virtio-net, NVMe-emu, snap, virtio-blk-emu); OR changes BAR window / window size; OR sets any firmware-committed parameter that requires a reset to take effect |
| Firmware / BFB | NIC firmware burn (flint, mft, mlxfwmanager, mlxconfig with -y reset); BlueField BFB reflash (bfb-install, rshim); BlueField mode change that requires a BFB swap |
| Host kernel state | host kernel boot parameter change (IOMMU mode iommu=pt / intel_iommu=on, default_hugepagesz, hugepagesz, nr_hugepages, pci=resource_alignment, vfio-pci.ids); requires host reboot |
| Hugepages | hugepage reservation change (/sys/kernel/mm/hugepages/*/nr_hugepages, sysctl vm.nr_hugepages); hugepage mount change (mount -t hugetlbfs); this is global state shared with every DOCA / DPDK process on the host |
| PCIe state | PCIe rebind / echo > /sys/bus/pci/.../{bind,unbind,remove,rescan}; representor enable/disable; eswitch mode change (devlink dev eswitch set ... mode {switchdev,legacy}); link down/up on a port currently carrying traffic |
| BlueField reboot class | BlueField cold reboot, BlueField warm reboot to apply mlxconfig; any change whose blast radius is "every hosted service on this DPU restarts" |
| Per-artifact cross-link | any per-artifact skill's ## Safety policy cross-links here for the cross-cutting rule body |
When any cell above fires, the agent MUST load this skill before
composing the first sentence of the answer. For production it MUST
walk the change-application discipline in this order:
TASKS.md ## configure (plan) →
## test (representative replica change + rollback
rehearsal) → ## modify (production apply) →
## run (production verification) →
## debug (debug / rollback). It MUST cite the
activation explicitly in the answer (e.g. "because this touches
mlxconfig, the answer follows the doca-hardware-safety
discipline …") so the user can audit the reasoning.
The activation is mandatory, not advisory. The most common failure mode this overlay prevents is "the agent recommended a mlxconfig change with no maintenance window, no out-of-band path, and no rollback statement, the user applied it, the management link dropped, and the box was unrecoverable without a physical console." The cost of one unjustified activation (a few extra paragraphs in the answer) is trivial compared to the cost of one missed activation.
Refuse-and-escalate is a hard rule
If any of the following is true, the agent MUST stop and refuse to recommend the change — not soften the warning, not proceed with a "this is risky but here's how" answer, not defer the rollback question to "you should think about that":
- The change has no documented rollback path AND the user cannot provide one. (Per
CAPABILITIES.md ## Safety policyrollback-must-be-documented rule.) - The change is link-breaking AND the host has no out-of-band access path. (Per
CAPABILITIES.md ## Safety policyout-of-band-precondition rule.) - The change touches hardware state AND the user has not confirmed an
explicit, time-boxed maintenance window. (Per
CAPABILITIES.md ## Safety policymaintenance-window rule.) - Production application is contemplated before the change and its
rollback have passed on a representative non-prod replica. This
refusal is intent-based: it applies to a plan, recommendation, or
next action that would reach production early, not only when the
user explicitly asks for "direct application." A replica
mismatched on the required hardware, firmware, kernel, module, or
function-topology axes does not satisfy the gate; obtain a
representative replica or refuse and escalate. (Per
TASKS.md ## testreplica-first rule.)
In each of these cases the correct answer shape is "this change requires X (here is why); the bundle refuses to recommend it without X; here is the route to obtain X" — not silence and not improvisation. The refuse-and-escalate rule is what makes the bundle's hardware-safety guidance trustworthy to production operators.
Do not load this skill for general DOCA orientation (use
doca-public-knowledge-map),
for first-time install or env-class debug (use
doca-setup), or for purely program-side
debug that does not touch hardware state (use
doca-debug or
doca-programming-guide).
What this skill provides
This is a thin loader. The body keeps only the orientation needed to pick the right next file. The substantive content lives in two companion files:
CAPABILITIES.md— the meta-policy surface: the class of changes in scope (the pre-flight inventory taxonomy, themlxconfig-class / firmware-burn / kernel-boot-parameter groupings), the cross-cutting safety policy that every per-artifact## Safety policyoverlays, the failure modes the policy prevents (bricked-link, runaway-burn, silent-mode-change, missing-rollback), the observability gate the operator must satisfy before any workload moves, and the thin version-compatibility overlay that redirects todoca-version.TASKS.md— the change-application workflow
Truncated for display — read the full file on GitHub.
Related Skills
ai-job-search
44.4kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
