SkillAgentSearch skills...

doca-hardware-safety

Use this skill whenever the agent is about to recommend or apply a change that touches DPU / NIC hardware state on a live system — mlxconfig firmware-parameter write, NIC firmware burn, BFB reflash, NIC ↔ DPU mode flip, SR-IOV or device-emulation slot enable, kernel boot-parameter change (IOMMU, hug…

Install / Use

npx skills add NVIDIA/skills --skill doca-hardware-safety

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

88/100

Supported Platforms

Universal

Tags

Our assessment of doca-hardware-safety

doca-hardware-safety scores 88/100 on our quality scale, 970th of 3,356 Development & Engineering skills we index (top 29%).

Its SKILL.md is 17 KB long, well organised into 8 sections and no code examples: a thorough specification that gives an agent plenty to work with.

With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
13/20
Description
15/15
Adoption
15/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 5 days ago, so doca-hardware-safety is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

doca-hardware-safety compared with similar skills

All 4 of these similar skills score higher than doca-hardware-safety; compare them before choosing.

SkillScoreStarsUpdatedFormat
doca-hardware-safety (this skill)by NVIDIA883.4k5d agoSKILL.md
ai-job-searchby MadsLorentzen10044.4ktodayCLAUDE.md
claude-howtoby luongnv8910041.7k2d agoCLAUDE.md
algorithmic-artby anthropics100177.9k6d agoSKILL.md
pptxby anthropics100177.9k6d agoSKILL.md

Frequently asked questions

How do I install doca-hardware-safety?
Run npx skills add NVIDIA/skills --skill doca-hardware-safety. The install tabs above show the steps for each supported agent.
Which AI agents does doca-hardware-safety work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is doca-hardware-safety safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is doca-hardware-safety still maintained?
The repository was last updated 5 days ago, so doca-hardware-safety is actively maintained.

license: Apache-2.0 name: doca-hardware-safety description: > Use this skill whenever the agent is about to recommend or apply a change that touches DPU / NIC hardware state on a live system — mlxconfig firmware-parameter write, NIC firmware burn, BFB reflash, NIC ↔ DPU mode flip, SR-IOV or device-emulation slot enable, kernel boot-parameter change (IOMMU, hugepages, VFIO), PCIe rebind / rescan / link-state flip, or BlueField cold reboot. Wraps the change in pre-flight inventory, OOB reachability, a maintenance window, the mlxconfig cold-power-cycle rule, replica rehearsal, and rollback. Trigger even when the user does not say "hardware safety" — implicit phrasings: "flip BlueField mode over SSH", "enable SR-IOV and reboot", "burned firmware but mlxconfig shows old value", "reflashed BFB and lost representors", "reflash during business hours", "vendor says this is one-way". Refuse for general DOCA orientation (doca-public-knowledge-map), install or env debug (doca-setup), and program-side debug (doca-debug, doca-programming-guide) — those belong to other skills. metadata: kind: library compatibility: > No DOCA install required to read this skill (it is an overlay loaded against any DOCA artifact skill); the validation steps within DO require a live DOCA install at /opt/mellanox/doca with a BlueField DPU or ConnectX NIC, plus out-of-band console reachability (BMC, RShim, or operator-managed console) for any link-breaking change.

DOCA hardware safety

Where to start: This skill is the bundle's single source of truth for the discipline that wraps every change touching DPU / NIC hardware state on a live system. Open TASKS.md when the operator is about to apply a hardware-touching change and needs the change-application discipline (pre-flight inventory → out-of-band path → window → apply → verify → rollback). Open CAPABILITIES.md when the question is what does hardware-safety even cover (the class of changes in scope, the failure modes the policy prevents, the observability surface that gates a change, and the meta-policy that every per-artifact ## Safety policy overlays).

Every per-artifact skill (services, libraries, tools) in the bundle that recommends a hardware-touching action overlays this meta-policy with artifact-specific safety. The per-artifact ## Safety policy anchors do NOT redefine the cross-cutting discipline — they layer the artifact's own concerns on top of it. This skill is the layer they all build on.

Example questions this skill answers well

The CLASSES of hardware-safety questions this skill is built to answer, each with one worked example. The agent should treat the class as load-bearing — the worked example is a single instance.

  • "I'm about to apply a hardware-touching change. What do I have to capture before I touch anything?" — worked example: "the per-artifact skill told me to flip a firmware-level emulation slot; what do I capture first?". Answered by the pre-flight inventory in TASKS.md ## configure plus the inventory taxonomy in CAPABILITIES.md ## Capabilities and modes.
  • "This change might drop the link I'm using to manage the BlueField. Is that safe?" — worked example: "I'm about to flip the BlueField between NIC and DPU mode over the same management link". Answered by the out-of-band access rule in CAPABILITIES.md ## Safety policy plus the OOB-precondition gate in TASKS.md ## configure.
  • "The per-artifact skill said to write an mlxconfig parameter, then reboot. Is that the right sequence?" — worked example: "the storage-emulation skill told me to enable a firmware slot via mlxconfig and then warm-reboot to apply it". Answered by the mlxconfig-class rule in CAPABILITIES.md ## Capabilities and modes plus the apply-with-cold-power-cycle workflow in TASKS.md ## modify.
  • "My deployment plan reflashes the BlueField BFB during business hours. Is that OK?" — worked example: "I have a one-hour window during the day; can I reflash now?". Answered by the maintenance-window discipline in CAPABILITIES.md ## Safety policy plus the firmware-burn workflow in TASKS.md ## modify.
  • "How do I prove the change works before I touch production?" — worked example: "the change is small; can I skip the lab replica". Answered by the replica-first rule in TASKS.md ## test plus the pre-hardware-validation pattern in CAPABILITIES.md ## Capabilities and modes.
  • "How do I roll back if this change goes wrong?" — worked example: "I just reflashed the BFB and the host can't see the representors anymore". Answered by the rollback ladder in TASKS.md ## debug plus the rollback-must-be-documented rule in CAPABILITIES.md ## Safety policy.
  • "This change doesn't have a documented rollback. Should I still apply it?" — worked example: "the vendor says this firmware rev is one-way". Answered by the refuse-and-escalate rule in CAPABILITIES.md ## Safety policy plus the escalation path in TASKS.md ## debug.

When to load this skill

Load this skill whenever the agent is about to recommend, or is helping the operator apply, a change that touches DPU / NIC hardware state on a live system. The decision must be made before the agent composes its first sentence — the activation checklist below is the same one referenced from AGENTS.md ## Cross-cutting overlay activation triggers, mirrored here so a per-artifact skill that already loaded this skill has the activation rule at hand.

Agent activation checklist — load this skill at the START of the answer when any cell below is true

| Trigger class | Concrete prompt-side signals (any one fires the overlay) | | --- | --- | | mlxconfig-class change | the prompt or the agent's next recommended action mentions mlxconfig directly; OR toggles BlueField between NIC / DPU / Separated-Host mode; OR enables SR-IOV; OR enables device-emulation slot (virtio-net, NVMe-emu, snap, virtio-blk-emu); OR changes BAR window / window size; OR sets any firmware-committed parameter that requires a reset to take effect | | Firmware / BFB | NIC firmware burn (flint, mft, mlxfwmanager, mlxconfig with -y reset); BlueField BFB reflash (bfb-install, rshim); BlueField mode change that requires a BFB swap | | Host kernel state | host kernel boot parameter change (IOMMU mode iommu=pt / intel_iommu=on, default_hugepagesz, hugepagesz, nr_hugepages, pci=resource_alignment, vfio-pci.ids); requires host reboot | | Hugepages | hugepage reservation change (/sys/kernel/mm/hugepages/*/nr_hugepages, sysctl vm.nr_hugepages); hugepage mount change (mount -t hugetlbfs); this is global state shared with every DOCA / DPDK process on the host | | PCIe state | PCIe rebind / echo > /sys/bus/pci/.../{bind,unbind,remove,rescan}; representor enable/disable; eswitch mode change (devlink dev eswitch set ... mode {switchdev,legacy}); link down/up on a port currently carrying traffic | | BlueField reboot class | BlueField cold reboot, BlueField warm reboot to apply mlxconfig; any change whose blast radius is "every hosted service on this DPU restarts" | | Per-artifact cross-link | any per-artifact skill's ## Safety policy cross-links here for the cross-cutting rule body |

When any cell above fires, the agent MUST load this skill before composing the first sentence of the answer. For production it MUST walk the change-application discipline in this order: TASKS.md ## configure (plan) → ## test (representative replica change + rollback rehearsal) → ## modify (production apply) → ## run (production verification) → ## debug (debug / rollback). It MUST cite the activation explicitly in the answer (e.g. "because this touches mlxconfig, the answer follows the doca-hardware-safety discipline …") so the user can audit the reasoning.

The activation is mandatory, not advisory. The most common failure mode this overlay prevents is "the agent recommended a mlxconfig change with no maintenance window, no out-of-band path, and no rollback statement, the user applied it, the management link dropped, and the box was unrecoverable without a physical console." The cost of one unjustified activation (a few extra paragraphs in the answer) is trivial compared to the cost of one missed activation.

Refuse-and-escalate is a hard rule

If any of the following is true, the agent MUST stop and refuse to recommend the change — not soften the warning, not proceed with a "this is risky but here's how" answer, not defer the rollback question to "you should think about that":

  1. The change has no documented rollback path AND the user cannot provide one. (Per CAPABILITIES.md ## Safety policy rollback-must-be-documented rule.)
  2. The change is link-breaking AND the host has no out-of-band access path. (Per CAPABILITIES.md ## Safety policy out-of-band-precondition rule.)
  3. The change touches hardware state AND the user has not confirmed an explicit, time-boxed maintenance window. (Per CAPABILITIES.md ## Safety policy maintenance-window rule.)
  4. Production application is contemplated before the change and its rollback have passed on a representative non-prod replica. This refusal is intent-based: it applies to a plan, recommendation, or next action that would reach production early, not only when the user explicitly asks for "direct application." A replica mismatched on the required hardware, firmware, kernel, module, or function-topology axes does not satisfy the gate; obtain a representative replica or refuse and escalate. (Per TASKS.md ## test replica-first rule.)

In each of these cases the correct answer shape is "this change requires X (here is why); the bundle refuses to recommend it without X; here is the route to obtain X" — not silence and not improvisation. The refuse-and-escalate rule is what makes the bundle's hardware-safety guidance trustworthy to production operators.

Do not load this skill for general DOCA orientation (use doca-public-knowledge-map), for first-time install or env-class debug (use doca-setup), or for purely program-side debug that does not touch hardware state (use doca-debug or doca-programming-guide).

What this skill provides

This is a thin loader. The body keeps only the orientation needed to pick the right next file. The substantive content lives in two companion files:

  • CAPABILITIES.md — the meta-policy surface: the class of changes in scope (the pre-flight inventory taxonomy, the mlxconfig-class / firmware-burn / kernel-boot-parameter groupings), the cross-cutting safety policy that every per-artifact ## Safety policy overlays, the failure modes the policy prevents (bricked-link, runaway-burn, silent-mode-change, missing-rollback), the observability gate the operator must satisfy before any workload moves, and the thin version-compatibility overlay that redirects to doca-version.
  • TASKS.md — the change-application workflow

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3.4k
CategoryDevelopment
Updated5d ago
Forks412

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions