foundationpose-setup
Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. Use for SAM3/TAO dependency conflicts, CUDA library failures, and depth-engine shape or precision decisions.
Install / Use
npx skills add NVIDIA/skills --skill foundationpose-setupInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of foundationpose-setup
foundationpose-setup scores 85/100 on our quality scale, 1337th of 2,125 Automation skills we index.
Its SKILL.md is 5.9 KB long, split into 7 sections and no code examples: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so foundationpose-setup is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
foundationpose-setup compared with similar skills
All 4 of these similar skills score higher than foundationpose-setup; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| foundationpose-setup (this skill)by NVIDIA | 85 | 3.4k | 5d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.0k | 13d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 84.4k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 6d ago | SKILL.md |
Frequently asked questions
- How do I install foundationpose-setup?
- Run
npx skills add NVIDIA/skills --skill foundationpose-setup. The install tabs above show the steps for each supported agent. - Which AI agents does foundationpose-setup work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is foundationpose-setup safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is foundationpose-setup still maintained?
- The repository was last updated 5 days ago, so foundationpose-setup is actively maintained.
Skill content
View source on GitHubname: foundationpose-setup description: Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. Use for SAM3/TAO dependency conflicts, CUDA library failures, and depth-engine shape or precision decisions. license: Apache-2.0 metadata: author: "zwdoescode zhengwang@nvidia.com" version: "0.1.0"
FoundationPose perception pipeline setup
Purpose
Prepare the FoundationPose perception pipeline
for depth, segmentation, and pose inference. Environment installation and engine construction
belong here; dataset adaptation, inference, and pose evaluation belong to
foundationpose-pipeline when that skill is installed.
Requirements
Work from the product checkout, not the installed skill directory. Find the user's checkout by
checking for pyproject.toml (project foundationpose-perception-pipeline),
tools/build_tao_engine.py, and config/defaults.yaml. If absent and setup was requested, clone
the product URL above into the user's workspace and enter it. For advice-only requests, use the
supplied diagnostics without cloning or installing anything. Commands below use paths relative
to the product root; references/ links are relative to this skill.
Read the checkout's README.md Requirements and Install sections for the matching revision.
The supported stack requires Linux x86_64, glibc >= 2.38, GLIBCXX_3.4.31, NVIDIA driver >= 580,
a CUDA toolkit >= 12.8 with nvcc, Python 3.12, uv, Git, Docker with GPU access, wget, and unzip.
Start GPU sizing at 24 GB and measure the densest scene; 32 GB was tested. Budget depth-cache
disk as roughly width * height * 4 * 3 bytes per scene, plus predictions and models.
Keep sibling directories for sam3/, foundation-pose-inference-library/, and models/ beside
the product checkout. models/ contains the deployable ONNX and engine, not FoundationStereo source.
Instructions
-
Preflight before installing. Check glibc, GLIBCXX, driver,
nvcc, tools, and Docker GPU access. An Ubuntu 22.04 host with glibc 2.35 cannot load the shipped FoundationPose library; report the unsupported runtime and stop setup there. Do not replace system libc or try to solve this withLD_LIBRARY_PATH. See installation. -
Install into the product's Python 3.12 venv. Follow installation for uv, SAM3, the FoundationPose build, and TAO Deploy. On a fresh venv use
uv sync --extra foundationpose; on an existing venv useuv sync --inexact --extra foundationposeto preserve out-of-band packages. -
Verify checkpoint access. SAM3 is gated at Hugging Face; an existing authorized token or usable cached checkpoint is sufficient. Request user action only if access is missing. FoundationPose and the documented FoundationStereo export are public Hugging Face downloads; credential hunting is not the first response to a network failure.
-
Prepare the depth engine. Read engine construction. Use the user's ONNX location or the sibling
models/directory. Adapt the dataset before measuring the engine shape; usetools/bop_adapt/adapt.py --config <profile> --src <source>as described in the checkout's README Dataset adaptation section. Only registered adapters are supported. Build with--shape-from-sceneon an adapted scene and FP32 unless the user requests a precision experiment. Setoverrides.depth.engineinconfig/<profile>.yaml. -
Set runtime paths and verify. From the product root:
source .venv/bin/activate export FOUNDATIONPOSE_ROOT="$(realpath ../foundation-pose-inference-library)" PIPELINE_SITE="$(realpath .venv/lib/python3.12/site-packages)" export LD_LIBRARY_PATH="${PIPELINE_SITE}/tensorrt_libs:${PIPELINE_SITE}/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}" python tools/verify_sam3.py python tools/verify_foundationpose.py python tools/verify_foundationstereo.py --config <profile> --engine <engine-path> python test/check_engine_depth_smoke.py --config <profile> --engine <engine-path>The first three verify components; the last also needs an adapted dataset. Expect
backend=tao,normalization=imagenet, the intended fixed shape, and nocropping N rowswarning. An unloaded or unavailable engine is an incomplete verification, not a pass.
Troubleshooting
| Symptom | Action |
|---|---|
| GLIBC_2.38 not found | Use a supported OS/runtime; a venv or library search path cannot upgrade host libc. |
| libcudart.so.13 missing | Check the product venv runtime wheels and absolute library paths before retrying pose. |
| SAM3 breaks after sync | Use --inexact; confirm numpy 1.26.x and reinstall the sibling SAM3 package if pruned. |
| pycuda build cannot find cuda.h | Check the CUDA toolkit, nvcc on PATH, or CUDA_ROOT. |
| TAO import or dependency conflict | Use TAO Deploy 7.1.0 with --no-deps; sync declared dependencies with --inexact. |
| Engine sidecar mismatch or cropping | Rebuild for this GPU, TensorRT version, precision, and adapted scene shape. |
Examples
- "Install the FoundationPose perception pipeline on this Ubuntu 24.04 GPU machine."
- "The pipeline cannot load libcudart.so.13 after I moved the checkout."
- "Build the TAO depth engine for my adapted T-LESS scenes."
Limitations and completion
Engines are machine-specific and must not be committed. With no dataset, download the ONNX and report shape-dependent engine construction and scene validation as pending; do not invent a rig shape. The pipeline's Apache license does not cover separately downloaded model weights; retain their upstream terms and SAM3's access requirements.
Report which preflight, install, checkpoint, engine, and verification steps actually passed, the checkout and engine paths, versions used, and remaining blockers. Do not equate installation or a smoke check with measured pose accuracy.
Related Skills
Agent-Reach
86.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.4k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
84.4k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
