nvmolkit-usage
Use when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches.
Install / Use
npx skills add NVIDIA/skills --skill bionemo-nvmolkit-usageInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Content & MediaSupported Platforms
Our assessment of nvmolkit-usage
nvmolkit-usage scores 95/100 on our quality scale, 115th of 770 Content & Media skills we index (top 15%).
Its SKILL.md is 19 KB long, well organised into 29 sections with 8 code examples: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so nvmolkit-usage is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-09-29. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
nvmolkit-usage compared with similar skills
All 4 of these similar skills score higher than nvmolkit-usage; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| nvmolkit-usage (this skill)by NVIDIA | 95 | 3.4k | 5d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.0k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | today | CLAUDE.md |
| crawl4aiby unclecode | 100 | 84.4k | 3d ago | MCP Server |
| Scraplingby D4Vinci | 100 | 84.4k | today | MCP Server |
Frequently asked questions
- How do I install nvmolkit-usage?
- Run
npx skills add NVIDIA/skills --skill nvmolkit-usage. The install tabs above show the steps for each supported agent. - Which AI agents does nvmolkit-usage work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is nvmolkit-usage safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is nvmolkit-usage still maintained?
- The repository was last updated 5 days ago, so nvmolkit-usage is actively maintained.
Skill content
View source on GitHubname: nvmolkit-usage description: >- Use when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches. license: Apache-2.0 metadata: author: Kevin Boyd (@scal444) owner: Kevin Boyd (@scal444) risk-tier: skill tags: [cheminformatics, rdkit, cuda]
nvMolKit usage
Purpose
GPU-accelerated, batched implementations of common RDKit operations. APIs mirror RDKit where possible but are batch-oriented: they take lists of rdkit.Chem.Mol (or lists of fingerprints) and process them in parallel on one or more GPUs. nvMolKit links against RDKit at build time; inputs and outputs are real RDKit Mol objects.
This skill covers the installed Python API. Building nvMolKit from source is out of scope.
Where nvMolKit does well
Reach for nvMolKit when:
- The workload is a large batch of molecules processed together (typically thousands or more).
- The metric is throughput / total wall time across the batch, not per-molecule latency.
- The same operation is repeated identically across the batch (fingerprinting a library, embedding/minimizing many conformers, bulk pairwise similarity), so the GPU stays saturated.
Requirements
- An NVIDIA GPU with compute capability 7.0 (V100) or higher
- A CUDA driver compatible with CUDA 12.6+.
- A working
torchinstall with CUDA support (nvMolKit returns GPU tensors viatorch's CUDA array interface).
When helping with installation, make the user choose a PyTorch CUDA backend that the host driver supports before installing nvMolKit. nvMolKit's PyPI wheels are built with CUDA Toolkit 12.9 and depend on CUDA 12 runtime packages, but pip/uv can still select a CUDA 13 PyTorch wheel unless the install command says otherwise.
- Conda: prefer conda-forge
pytorch-gpu; pincuda-version=12.6or another CUDA version supported by the driver. - pip: send the user to the PyTorch install selector or previous-versions page to install
torchfor a CUDA 12.x backend before installing nvMolKit. - uv: install nvMolKit with an explicit backend, e.g.
uv pip install --torch-backend=cu128 nvmolkit.
Inputs
- Required: choose an operation and supply molecules or fingerprints from the user's code or molecular dataset. Parse SMILES with RDKit and reject failed parses (
None). - Molecular operations use RDKit
Molobjects. Add hydrogens for ETKDG; minimization and conformer comparisons need existing conformers. - Fingerprint similarity takes packed
AsyncGpuResult, torch tensors, or NumPy arrays: one molecule per row, withint32oruint32words. - Optional: take conformer counts, fingerprint settings, cutoffs, output modes, and hardware options from the user's requested workflow; otherwise use the documented API defaults.
Limitations
- CUDA is required; there is no CPU fallback. Use RDKit directly when CPU execution is needed.
- Plain RDKit is usually preferable for single-molecule work or operations that cannot be batched.
- ETKDG does not support custom bounds matrices, custom CPCI, coordinate maps, or separate-fragment embedding.
- Substructure search does not support chirality-aware matching, enhanced stereochemistry, or other advanced RDKit
SubstructMatchParametersoptions.
Instructions
- Run the smoke test below before writing nvMolKit code.
- Choose an API from the entry-point table and apply its input requirements.
- Handle its result as described below; synchronize asynchronous GPU results before host reads.
Verify the install before writing real code
import nvmolkit
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
print("nvmolkit:", nvmolkit.__version__)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())
mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
result = fpgen.GetFingerprints(mols)
torch.cuda.synchronize()
fps = result.torch()
print("fps shape:", tuple(fps.shape), "dtype:", fps.dtype)
# Expected: shape (3, 32), dtype torch.int32 (1024 bits packed into 32 int32s per row)
If this fails, point the user at the installation guide rather than guessing.
Entry points
| Task | Module | Primary entry point |
|---|---|---|
| Morgan fingerprints | nvmolkit.fingerprints | MorganFingerprintGenerator(radius, fpSize).GetFingerprints(mols) |
| Bulk Tanimoto / cosine similarity | nvmolkit.similarity | crossTanimotoSimilarity(...), crossCosineSimilarity(...), plus *MemoryConstrained variants for results too large to fit in GPU memory |
| ETKDG conformer embedding | nvmolkit.embedMolecules | EmbedMolecules(molecules, params, confsPerMolecule, ...) |
| MMFF94 optimization (one-shot) | nvmolkit.mmffOptimization | MMFFOptimizeMoleculesConfs(molecules, ..., minimizerKind=..., fireOptions=...) |
| UFF optimization (one-shot) | nvmolkit.uffOptimization | UFFOptimizeMoleculesConfs(molecules, ..., minimizerKind=..., fireOptions=...) |
| Forcefield with custom options + constraints | nvmolkit.batchedForcefield | MMFFBatchedForcefield(mols, properties=..., nonBondedThreshold=..., ignoreInterfragInteractions=..., hardwareOptions=...), UFFBatchedForcefield(mols, vdwThreshold=..., ...). Per-molecule view ff[i] exposes add_distance_constraint, add_position_constraint, add_angle_constraint, add_torsion_constraint. Methods: .compute_energy(), .compute_gradients(), .minimize(maxIters, forceTol, minimizerKind=..., fireOptions=...) |
| Pairwise conformer RMSD | nvmolkit.conformerRmsd | GetConformerRMSMatrix(mol), GetConformerRMSMatrixBatch(mols) |
| Torsion Fingerprint Deviation (TFD) | nvmolkit.tfd | GetTFDMatrix(mol), GetTFDMatrices(mols) |
| Butina clustering | nvmolkit.clustering | butina(distance_matrix, cutoff) (precomputed matrix), fused_butina(fingerprints, cutoff) (memory-efficient, on-the-fly); both support explicit RDKit and device output modes |
| Substructure search | nvmolkit.substructure | hasSubstructMatch, countSubstructMatches, getSubstructMatches |
| Maximum common substructure | nvmolkit.mcs | findMCS(mols, ...) for all pairs, explicit pairs, or two paired molecule lists |
| Hardware tuning (batch size, GPU IDs) | nvmolkit.types | HardwareOptions(...) passed to ETKDG / MMFF / UFF |
| Optional autotuning | nvmolkit.autotune | tune_embed_molecules, tune_mmff_optimize, tune_uff_optimize, tune_batched_forcefield, tune_substructure, tune_mcs. Requires the optuna package |
Result types and execution model
Two return shapes carry GPU-resident output, depending on what the operation produces.
AsyncGpuResult
Used by operations that return a single flat tensor (fingerprints, similarity matrices, RMSD/TFD vectors, Butina inputs). Key behaviors:
- Asynchronous. The kernel may not have completed when the call returns.
result.torch()returns a zero-copytorch.Tensoron the GPU. Caller is responsible for synchronizing before reading values on the host.result.numpy()synchronizes and returns a CPU numpy array.- Exposes
__cuda_array_interface__, so it can be passed directly into other nvMolKit functions (e.g. fingerprints → similarity) with no host round-trip.
CUDA stream control
A subset of the AsyncGpuResult-returning APIs accept an optional stream: torch.cuda.Stream | None = None argument so callers can submit nvMolKit work to a non-default stream and overlap it with their own kernels. When omitted, the call uses the current torch stream.
APIs that take a stream argument:
MorganFingerprintGenerator.GetFingerprintscrossTanimotoSimilarity,crossCosineSimilarity, and their*MemoryConstrainedvariantsbutina,fused_butinaGetConformerRMSMatrix,GetConformerRMSMatrixBatch
Other APIs (ETKDG, MMFF/UFF optimization, TFD, substructure search, MCS) are synchronous to the caller — no stream plumbing needed.
Typical pattern:
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity
stream = torch.cuda.Stream()
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]
with torch.cuda.stream(stream):
fps = fpgen.GetFingerprints(mols, stream=stream)
sim = crossTanimotoSimilarity(fps, stream=stream)
stream.synchronize()
print(sim.torch())
Device3DResult
Used by ETKDG embedding and MMFF/UFF optimization (one-shot and BatchedForcefield) when called with output=CoordinateOutput.DEVICE. The GPU-resident equivalent of writing conformers back to Mol objects. Fields:
values:AsyncGpuResultof shape(total_atoms, 3)float64. Concatenated conformer coordinates in CSR-style layout.atom_starts,mol_indices,conf_indices:AsyncGpuResultint32 buffers describing the layout (values[atom_starts[i]:atom_starts[i+1]]is conformeri's atoms).energies,converged:AsyncGpuResultbuffers populated only for MMFF/UFF minimization (not for plain ETKDG).gpu_id: device the buffers live on. ThetargetGpuargument on each API picks this;targetGpu=-1uses the default consolidation device..per_molecule()returns nestedlist[list[torch.Tensor]]of per-conformer views;.dense(pad_value=nan)materializes a padded(n_mols, max_confs, max_atoms, 3)tensor.
The default mode (CoordinateOutput.RDKIT_CONFORMERS) still writes optimized coordinates back into each Mol and returns Python lists of energies/convergence flags. Reach for CoordinateOutput.DEVICE when chaining downstream GPU work (e.g. ETKDG → MMFF → similarity scoring) without host round-trips.
MCSBatchResult
findMCS is synchronous and returns an MCSBatchResult backed by CPU NumPy
arrays. Results are always flat: result[k] (or result.get_result(k))
materializes the result at pair position k, not generally the result for
molecule k. Use result.pairs[k] to identify that pair. In all_pairs mode
these are the generated pairs over mols; in pairs mode they exactly preserve
the supplied pair sequence; in paired_lists mode item k compares mols[k]
with mols_b[k], while result.pairs[k] uses the combined-table indices
(k, len(mols) + k). Each MCSResult has pair, num_atoms, num_bonds,
canceled, atom_mapping, and bond_mapping; the two columns of each mapping
index the first and second molecule of that result pair, respectively.
Configuration
For ETKDG, forcefield, substructure, or MCS tuning, read the advanced configuration reference. It lists configuration fields, defaults, GPU selection, and autotuning APIs.
Examples
Morgan fingerprints + bulk Tanimoto similarity
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity
smiles = ["CCO", "CCN", "c1ccccc1", "CC(=O)O", "CCOCC"]
mols = [Chem.MolFromSmiles(smi) for smi in smiles]
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
fps = fpgen.GetFingerprints(mols)
sim = crossTanimotoSimilarity(fps)
torch.cuda.synchronize()
print(sim.torch())
Inputs are list[Mol]. Output of GetFingerprints is an AsyncGpuResult wrapping an (n_mols, fpSize / 32) int32 tensor of packed bits. Pass it straight into crossTanimotoSimilarity for an (n, n) similarity matrix; pass two fingerprint sets for an (n, m) cross-matrix. For sets too large to materialize on the GPU, use crossTanimotoSimilarityMemoryConstrained (chunked compute, returns numpy on CPU).
ETKDG conformer embe
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
crawl4ai
84.4kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Scrapling
84.4k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
