doca-gpunetio
Use this skill when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via doca_gpu_eth_rxq / doca_gpu_eth_txq, standing up the per-CUDA-device doca_gpu context, designing the persistent CUDA kernel that drains the GPU-visible queue, runn…
Install / Use
npx skills add NVIDIA/skills --skill doca-gpunetioInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of doca-gpunetio
doca-gpunetio scores 88/100 on our quality scale, 969th of 3,356 Development & Engineering skills we index (top 29%).
Its SKILL.md is 15 KB long, well organised into 8 sections and no code examples: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so doca-gpunetio is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
doca-gpunetio compared with similar skills
All 4 of these similar skills score higher than doca-gpunetio; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| doca-gpunetio (this skill)by NVIDIA | 88 | 3.4k | 5d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.0k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | today | CLAUDE.md |
| ai-job-searchby MadsLorentzen | 100 | 44.4k | today | CLAUDE.md |
| claude-howtoby luongnv89 | 100 | 41.7k | 2d ago | CLAUDE.md |
Frequently asked questions
- How do I install doca-gpunetio?
- Run
npx skills add NVIDIA/skills --skill doca-gpunetio. The install tabs above show the steps for each supported agent. - Which AI agents does doca-gpunetio work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is doca-gpunetio safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is doca-gpunetio still maintained?
- The repository was last updated 5 days ago, so doca-gpunetio is actively maintained.
Skill content
View source on GitHublicense: Apache-2.0
name: doca-gpunetio
description: >
Use this skill when the user is doing hands-on DOCA GPUNetIO
programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth
queue via doca_gpu_eth_rxq / doca_gpu_eth_txq, standing up the
per-CUDA-device doca_gpu context, designing the persistent CUDA
kernel that drains the GPU-visible queue, running the dual
capability check (DOCA cap-query plus cudaGetDeviceProperties),
registering cudaMalloc pools via doca_buf_arr_create_, or
debugging DOCA_ERROR_ returns from the GPUNetIO API. Trigger
even when the user does not explicitly mention "DOCA GPUNetIO"
or "persistent kernel" — typical implicit phrasings include
"CUDA kernel reading packets directly from the NIC",
"GPU-initiated networking on BlueField", "DOCA_ERROR_DRIVER on
doca_gpu_create", "nvidia_peermem not loaded",
"kernel-per-packet is too slow", or "which GPU supports GPU-side
packet I/O". Refuse and route elsewhere for general CUDA
programming, DOCA Ethernet queue bring-up, DOCA DPA, or
DOCA install — those belong to other skills.
metadata:
kind: library
compatibility: >
Requires DOCA SDK at /opt/mellanox/doca on Linux (Ubuntu 22.04/24.04 or
RHEL/SLES) with a BlueField DPU or ConnectX NIC. Reads the local install
via pkg-config doca-gpunetio. Requires an NVIDIA GPU with CUDA toolkit
(matched to DOCA per the DOCA Compatibility Policy) and the
nvidia_peermem kernel module loaded for GPUDirect RDMA; some samples
need an InfiniBand-capable RNIC.
DOCA GPUNetIO
Where to start: This skill assumes DOCA is already installed,
the CUDA toolkit is installed and matched to the DOCA install, and
the user is doing hands-on GPUNetIO work — i.e. wiring a DOCA
network queue into a CUDA kernel on an NVIDIA GPU. Open
TASKS.md if the user wants to do something
(configure / build / modify / run / test / debug); open
CAPABILITIES.md when the question is what
can GPUNetIO express on this version + this GPU. If the user has
not installed DOCA yet, route to
doca-setup first; if the user has
not set up the underlying Ethernet RX/TX queues yet, that is a
DOCA Ethernet question — route to
doca-eth.
Example questions this skill answers well
The CLASSES of GPUNetIO questions this skill is built to answer, each with one worked example. The agent should treat the class as the load-bearing piece — the worked example is a single instance.
- "How do I get a CUDA kernel to receive packets directly from
the NIC?" — worked example: "persistent kernel on one GPU
reads packets from a
doca_gpu_eth_rxqbuilt on top of a representordoca_eth_rxqand counts them per-flow". Answered by the persistent-kernel pattern inCAPABILITIES.md ## Capabilities and modes- the GPU-side bring-up workflow in
TASKS.md ## configure.
- the GPU-side bring-up workflow in
- "Can I run GPUNetIO on this GPU?" — worked example: "my
host has one Ampere card and one Turing card; which one
supports GPU-initiated networking?". Answered by the dual
capability-discovery rule (DOCA cap-query AND
cudaGetDevicePropertiesagainst the CUDA device ordinal) inCAPABILITIES.md ## Capabilities and modes- the device-enumeration step in
TASKS.md ## configure.
- the device-enumeration step in
- "Why does my GPUNetIO setup fail with
DOCA_ERROR_NOT_SUPPORTEDeven though doca-eth came up fine?" — worked example: "nvidia_peermemis not loaded so GPUDirect RDMA is unavailable". Answered by the env preconditions inCAPABILITIES.md ## Safety policy- the env checklist in
TASKS.md ## configurestep 1.
- the env checklist in
- "How do I move data between CUDA-allocated buffers and a DOCA
queue?" — worked example: "use
cudaMallocfor the receive buffer pool and register it with DOCA viadoca_buf_arr_create_*before starting the context". Answered by the CUDA-allocator- DOCA-registration overlay in
CAPABILITIES.md ## Safety policy - the buffer-prep step in
TASKS.md ## configurestep 4.
- DOCA-registration overlay in
- "Is the GPUNetIO API I'm reading about on my installed DOCA +
CUDA combination?" — worked example: "is the persistent-kernel
helper available with the CUDA toolkit version I have?".
Answered by the version-compatibility overlay in
CAPABILITIES.md ## Version compatibilitywhich cross-links the canonical detection chain indoca-versionand adds the GPUNetIO-specific DOCA must match CUDA overlay. - "What does this
DOCA_ERROR_*from a GPUNetIO call mean and which layer caused it?" — worked example: "DOCA_ERROR_DRIVERondoca_gpu_*_create— is it DOCA, CUDA, or the underlying doca-eth queue?". Answered by the GPUNetIO overlay on the cross-library taxonomy inCAPABILITIES.md ## Error taxonomy- the layered ladder in
TASKS.md ## debugthat escalates todoca-debug.
- the layered ladder in
Audience
This skill serves external developers building applications
that consume the DOCA GPUNetIO library — i.e., users whose code
calls doca_gpu_* from host C/C++ to stand up the per-GPU
context and the GPU-visible queue handles, and whose CUDA kernel
(.cu translation unit) uses those handles from device code to
submit / receive packets. The canonical target shape is the GPU
Packet Processing reference application: a CUDA persistent
kernel on an NVIDIA GPU that polls a GPU-visible RX queue and
processes packets in-place on the GPU. It is not for NVIDIA
developers contributing to DOCA GPUNetIO itself.
Language scope. DOCA GPUNetIO ships as a C / CUDA library
with pkg-config module name doca-gpunetio. The host-side API
is C; the device-side API is CUDA C++ used inside a .cu
kernel. The shipped samples and the GPU Packet Processing
reference application are written in C + CUDA C++ (NVIDIA's
choice). Other-language consumers are limited in practice — the
device-side API has no FFI escape hatch because the kernel must
be a CUDA translation unit — but a Rust / Go / Python host-side
wrapper that drives the host-side doca_gpu_* setup and
launches a CUDA kernel built separately is still useful, and the
skill keeps the lifecycle, capability-discovery, env-precondition,
and error-taxonomy guidance language-neutral.
When to load this skill
Load this skill when the user is doing hands-on DOCA GPUNetIO work, in any host language plus CUDA. Concretely:
- Initializing a
doca_gpuagainst a specific CUDA device ordinal on a host with one or more NVIDIA GPUs. - Creating a GPU-visible queue handle (
doca_gpu_eth_rxq,doca_gpu_eth_txq) on top of an existingdoca_eth_rxq/doca_eth_txqfrom DOCA Ethernet, and passing the handle into a CUDA kernel for device-side use. - Writing or modifying the persistent CUDA kernel that drains the GPU-visible RX queue in a long-running loop (the canonical GPU Packet Processing shape).
- Allocating GPU buffers via
cudaMallocand registering them with DOCA via thedoca_buf_arr_create_*family beforedoca_ctx_start(). - Checking which GPUNetIO features are supported on the active
doca_devinfo(DOCA cap-query family) AND on the candidate CUDA device (cudaGetDevicePropertiesand CUDA-driver-version checks). - Debugging a
DOCA_ERROR_*returned from a GPUNetIO call — in particular disambiguating DOCA capability missing from CUDA device too old fromnvidia_peermemnot loaded from CUDA driver + DOCA version skew. - Designing host-side bindings for non-C languages that drive a CUDA kernel they built separately — the env-precondition and capability-discovery rules in this skill still apply.
Do not load this skill for general DOCA orientation, install
of DOCA or the CUDA toolkit, the underlying DOCA Ethernet queue
setup, or non-GPUNetIO library questions. For those, route
through doca-public-knowledge-map
to the matching upstream guide.
What this skill provides
This is a thin loader. The body keeps only the orientation needed to pick the right next file. The substantive GPUNetIO-specific material lives in two companion files:
CAPABILITIES.md— what GPUNetIO can express on this version- this GPU: the
doca_gpuper-device context, the GPU-visible RX / TX queue handles layered on doca-eth, the persistent CUDA-kernel pattern as the default usage shape, the capability-query surface (the doca-ethdoca_eth_rxq_cap_is_type_supported/doca_eth_rxq_cap_get_*family indoca_eth_rxq.h, plus the matchingdoca_eth_txq_cap_*family, on the DOCA side, pluscudaGetDevicePropertieson the CUDA side), the GPUNetIO error taxonomy mapped onto the cross-libraryDOCA_ERROR_*set, the observability surface (CUDA-side counters + DOCA-side per-task completion), and the safety policy that gates env preconditions (CUDA + DOCA version match,nvidia_peermem, CUDA buffer registration).
- this GPU: the
TASKS.md— step-by-step workflows for the six in-scope GPUNetIO verbs:configure,build,modify,run,test,debug. Plus a## rollbackoverlay (GPUNetIO-specific five-step teardown that signals the persistent kernel to drain, unregisters GPU buffers in reverse-register order, and leaves the parent doca-eth queue intact) and the 5-phase universal debug-loop instantiation appended to## debug. Plus aDeferred task verbsblock that points out-of-scope questions at the right next skill.
The skill assumes a host where DOCA is already installed at the
standard location, an NVIDIA GPU is physically present, the CUDA
toolkit is installed and its version is matched to the DOCA
install per the DOCA Compatibility Policy, and the underlying
DOCA Ethernet RX/TX queues are already configured (this skill
sits on top of doca-eth, not below it). It does not cover
installing DOCA or the CUDA toolkit — that path goes through
doca-setup.
What this skill deliberately does not ship
This skill is agent guidance, not a samples or templates bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:
- Pre-written DOCA GPUNetIO application source code or CUDA
kernel source, in any language. The verified GPUNetIO
source is the shipped C + CUDA samples at
/opt/mellanox/doca/samples/doca_gpunetio/and the GPU Packet Processing reference application. The agent's job is to route the user to those files and prescribe a minimum-diff modification on them via the universal modify-a-sample workflow indoca-programming-guide, layered with the GPUNetIO-specific overrides inTASKS.md ## modify. - Standalone build manifests (
meson.build,CMakeLists.txt, …) parked inside the skill. The agent constructs the build manifest in the user's project directory against the user's installed DOCA + CUDA toolkit, wherepkg-config --modversion doca-gpunetio,pkg-config --modversion doca-common, andnvcc --versionform the version gate. - A
samples/,bindings/, orreference/subtree of any kind. A mock or incomplete artifact in this skill's tree, even one labeled "reference", is misleading: users will read it as buildable.
Loading order
- Read this
SKILL.mdfirst to confirm the user's question is in scope. - **For the GPUNetIO capability matrix, the
doca_gpuper-device context, the persistent-kernel pattern, the dual capability query, the env-precondition policy, the error taxonomy, the observability surfa
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ai-job-search
44.4kThe job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
claude-howto
41.7kA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
