SkillAgentSearch skills...

nvmolkit-usage

Use when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches.

Install / Use

npx skills add NVIDIA/skills --skill bionemo-nvmolkit-usage

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

95/100

Supported Platforms

Universal

Our assessment of nvmolkit-usage

nvmolkit-usage scores 95/100 on our quality scale, 115th of 770 Content & Media skills we index (top 15%).

Its SKILL.md is 19 KB long, well organised into 29 sections with 8 code examples: a thorough specification that gives an agent plenty to work with.

With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
15/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 5 days ago, so nvmolkit-usage is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.

Automated pattern scan on 2026-09-29. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

nvmolkit-usage compared with similar skills

All 4 of these similar skills score higher than nvmolkit-usage; compare them before choosing.

SkillScoreStarsUpdatedFormat
nvmolkit-usage (this skill)by NVIDIA953.4k5d agoSKILL.md
Agent-Reachby Panniantong10086.0k13d agoCLAUDE.md
headroomby headroomlabs-ai10074.0ktodayCLAUDE.md
crawl4aiby unclecode10084.4k3d agoMCP Server
Scraplingby D4Vinci10084.4ktodayMCP Server

Frequently asked questions

How do I install nvmolkit-usage?
Run npx skills add NVIDIA/skills --skill nvmolkit-usage. The install tabs above show the steps for each supported agent.
Which AI agents does nvmolkit-usage work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is nvmolkit-usage safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is nvmolkit-usage still maintained?
The repository was last updated 5 days ago, so nvmolkit-usage is actively maintained.

name: nvmolkit-usage description: >- Use when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches. license: Apache-2.0 metadata: author: Kevin Boyd (@scal444) owner: Kevin Boyd (@scal444) risk-tier: skill tags: [cheminformatics, rdkit, cuda]

nvMolKit usage

Purpose

GPU-accelerated, batched implementations of common RDKit operations. APIs mirror RDKit where possible but are batch-oriented: they take lists of rdkit.Chem.Mol (or lists of fingerprints) and process them in parallel on one or more GPUs. nvMolKit links against RDKit at build time; inputs and outputs are real RDKit Mol objects.

This skill covers the installed Python API. Building nvMolKit from source is out of scope.

Where nvMolKit does well

Reach for nvMolKit when:

  • The workload is a large batch of molecules processed together (typically thousands or more).
  • The metric is throughput / total wall time across the batch, not per-molecule latency.
  • The same operation is repeated identically across the batch (fingerprinting a library, embedding/minimizing many conformers, bulk pairwise similarity), so the GPU stays saturated.

Requirements

  • An NVIDIA GPU with compute capability 7.0 (V100) or higher
  • A CUDA driver compatible with CUDA 12.6+.
  • A working torch install with CUDA support (nvMolKit returns GPU tensors via torch's CUDA array interface).

When helping with installation, make the user choose a PyTorch CUDA backend that the host driver supports before installing nvMolKit. nvMolKit's PyPI wheels are built with CUDA Toolkit 12.9 and depend on CUDA 12 runtime packages, but pip/uv can still select a CUDA 13 PyTorch wheel unless the install command says otherwise.

  • Conda: prefer conda-forge pytorch-gpu; pin cuda-version=12.6 or another CUDA version supported by the driver.
  • pip: send the user to the PyTorch install selector or previous-versions page to install torch for a CUDA 12.x backend before installing nvMolKit.
  • uv: install nvMolKit with an explicit backend, e.g. uv pip install --torch-backend=cu128 nvmolkit.

Inputs

  • Required: choose an operation and supply molecules or fingerprints from the user's code or molecular dataset. Parse SMILES with RDKit and reject failed parses (None).
  • Molecular operations use RDKit Mol objects. Add hydrogens for ETKDG; minimization and conformer comparisons need existing conformers.
  • Fingerprint similarity takes packed AsyncGpuResult, torch tensors, or NumPy arrays: one molecule per row, with int32 or uint32 words.
  • Optional: take conformer counts, fingerprint settings, cutoffs, output modes, and hardware options from the user's requested workflow; otherwise use the documented API defaults.

Limitations

  • CUDA is required; there is no CPU fallback. Use RDKit directly when CPU execution is needed.
  • Plain RDKit is usually preferable for single-molecule work or operations that cannot be batched.
  • ETKDG does not support custom bounds matrices, custom CPCI, coordinate maps, or separate-fragment embedding.
  • Substructure search does not support chirality-aware matching, enhanced stereochemistry, or other advanced RDKit SubstructMatchParameters options.

Instructions

  1. Run the smoke test below before writing nvMolKit code.
  2. Choose an API from the entry-point table and apply its input requirements.
  3. Handle its result as described below; synchronize asynchronous GPU results before host reads.

Verify the install before writing real code

import nvmolkit
import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator

print("nvmolkit:", nvmolkit.__version__)
print("cuda available:", torch.cuda.is_available())
print("device count:", torch.cuda.device_count())

mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
result = fpgen.GetFingerprints(mols)
torch.cuda.synchronize()
fps = result.torch()
print("fps shape:", tuple(fps.shape), "dtype:", fps.dtype)
# Expected: shape (3, 32), dtype torch.int32  (1024 bits packed into 32 int32s per row)

If this fails, point the user at the installation guide rather than guessing.

Entry points

| Task | Module | Primary entry point | |---|---|---| | Morgan fingerprints | nvmolkit.fingerprints | MorganFingerprintGenerator(radius, fpSize).GetFingerprints(mols) | | Bulk Tanimoto / cosine similarity | nvmolkit.similarity | crossTanimotoSimilarity(...), crossCosineSimilarity(...), plus *MemoryConstrained variants for results too large to fit in GPU memory | | ETKDG conformer embedding | nvmolkit.embedMolecules | EmbedMolecules(molecules, params, confsPerMolecule, ...) | | MMFF94 optimization (one-shot) | nvmolkit.mmffOptimization | MMFFOptimizeMoleculesConfs(molecules, ..., minimizerKind=..., fireOptions=...) | | UFF optimization (one-shot) | nvmolkit.uffOptimization | UFFOptimizeMoleculesConfs(molecules, ..., minimizerKind=..., fireOptions=...) | | Forcefield with custom options + constraints | nvmolkit.batchedForcefield | MMFFBatchedForcefield(mols, properties=..., nonBondedThreshold=..., ignoreInterfragInteractions=..., hardwareOptions=...), UFFBatchedForcefield(mols, vdwThreshold=..., ...). Per-molecule view ff[i] exposes add_distance_constraint, add_position_constraint, add_angle_constraint, add_torsion_constraint. Methods: .compute_energy(), .compute_gradients(), .minimize(maxIters, forceTol, minimizerKind=..., fireOptions=...) | | Pairwise conformer RMSD | nvmolkit.conformerRmsd | GetConformerRMSMatrix(mol), GetConformerRMSMatrixBatch(mols) | | Torsion Fingerprint Deviation (TFD) | nvmolkit.tfd | GetTFDMatrix(mol), GetTFDMatrices(mols) | | Butina clustering | nvmolkit.clustering | butina(distance_matrix, cutoff) (precomputed matrix), fused_butina(fingerprints, cutoff) (memory-efficient, on-the-fly); both support explicit RDKit and device output modes | | Substructure search | nvmolkit.substructure | hasSubstructMatch, countSubstructMatches, getSubstructMatches | | Maximum common substructure | nvmolkit.mcs | findMCS(mols, ...) for all pairs, explicit pairs, or two paired molecule lists | | Hardware tuning (batch size, GPU IDs) | nvmolkit.types | HardwareOptions(...) passed to ETKDG / MMFF / UFF | | Optional autotuning | nvmolkit.autotune | tune_embed_molecules, tune_mmff_optimize, tune_uff_optimize, tune_batched_forcefield, tune_substructure, tune_mcs. Requires the optuna package |

Result types and execution model

Two return shapes carry GPU-resident output, depending on what the operation produces.

AsyncGpuResult

Used by operations that return a single flat tensor (fingerprints, similarity matrices, RMSD/TFD vectors, Butina inputs). Key behaviors:

  • Asynchronous. The kernel may not have completed when the call returns.
  • result.torch() returns a zero-copy torch.Tensor on the GPU. Caller is responsible for synchronizing before reading values on the host.
  • result.numpy() synchronizes and returns a CPU numpy array.
  • Exposes __cuda_array_interface__, so it can be passed directly into other nvMolKit functions (e.g. fingerprints → similarity) with no host round-trip.

CUDA stream control

A subset of the AsyncGpuResult-returning APIs accept an optional stream: torch.cuda.Stream | None = None argument so callers can submit nvMolKit work to a non-default stream and overlap it with their own kernels. When omitted, the call uses the current torch stream.

APIs that take a stream argument:

  • MorganFingerprintGenerator.GetFingerprints
  • crossTanimotoSimilarity, crossCosineSimilarity, and their *MemoryConstrained variants
  • butina, fused_butina
  • GetConformerRMSMatrix, GetConformerRMSMatrixBatch

Other APIs (ETKDG, MMFF/UFF optimization, TFD, substructure search, MCS) are synchronous to the caller — no stream plumbing needed.

Typical pattern:

import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity

stream = torch.cuda.Stream()
fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
mols = [Chem.MolFromSmiles(smi) for smi in ["CCO", "c1ccccc1", "CC(=O)O"]]

with torch.cuda.stream(stream):
    fps = fpgen.GetFingerprints(mols, stream=stream)
    sim = crossTanimotoSimilarity(fps, stream=stream)
stream.synchronize()
print(sim.torch())

Device3DResult

Used by ETKDG embedding and MMFF/UFF optimization (one-shot and BatchedForcefield) when called with output=CoordinateOutput.DEVICE. The GPU-resident equivalent of writing conformers back to Mol objects. Fields:

  • values: AsyncGpuResult of shape (total_atoms, 3) float64. Concatenated conformer coordinates in CSR-style layout.
  • atom_starts, mol_indices, conf_indices: AsyncGpuResult int32 buffers describing the layout (values[atom_starts[i]:atom_starts[i+1]] is conformer i's atoms).
  • energies, converged: AsyncGpuResult buffers populated only for MMFF/UFF minimization (not for plain ETKDG).
  • gpu_id: device the buffers live on. The targetGpu argument on each API picks this; targetGpu=-1 uses the default consolidation device.
  • .per_molecule() returns nested list[list[torch.Tensor]] of per-conformer views; .dense(pad_value=nan) materializes a padded (n_mols, max_confs, max_atoms, 3) tensor.

The default mode (CoordinateOutput.RDKIT_CONFORMERS) still writes optimized coordinates back into each Mol and returns Python lists of energies/convergence flags. Reach for CoordinateOutput.DEVICE when chaining downstream GPU work (e.g. ETKDG → MMFF → similarity scoring) without host round-trips.

MCSBatchResult

findMCS is synchronous and returns an MCSBatchResult backed by CPU NumPy arrays. Results are always flat: result[k] (or result.get_result(k)) materializes the result at pair position k, not generally the result for molecule k. Use result.pairs[k] to identify that pair. In all_pairs mode these are the generated pairs over mols; in pairs mode they exactly preserve the supplied pair sequence; in paired_lists mode item k compares mols[k] with mols_b[k], while result.pairs[k] uses the combined-table indices (k, len(mols) + k). Each MCSResult has pair, num_atoms, num_bonds, canceled, atom_mapping, and bond_mapping; the two columns of each mapping index the first and second molecule of that result pair, respectively.

Configuration

For ETKDG, forcefield, substructure, or MCS tuning, read the advanced configuration reference. It lists configuration fields, defaults, GPU selection, and autotuning APIs.

Examples

Morgan fingerprints + bulk Tanimoto similarity

import torch
from rdkit import Chem
from nvmolkit.fingerprints import MorganFingerprintGenerator
from nvmolkit.similarity import crossTanimotoSimilarity

smiles = ["CCO", "CCN", "c1ccccc1", "CC(=O)O", "CCOCC"]
mols = [Chem.MolFromSmiles(smi) for smi in smiles]

fpgen = MorganFingerprintGenerator(radius=2, fpSize=1024)
fps = fpgen.GetFingerprints(mols)

sim = crossTanimotoSimilarity(fps)
torch.cuda.synchronize()
print(sim.torch())

Inputs are list[Mol]. Output of GetFingerprints is an AsyncGpuResult wrapping an (n_mols, fpSize / 32) int32 tensor of packed bits. Pass it straight into crossTanimotoSimilarity for an (n, n) similarity matrix; pass two fingerprint sets for an (n, m) cross-matrix. For sets too large to materialize on the GPU, use crossTanimotoSimilarityMemoryConstrained (chunked compute, returns numpy on CPU).

ETKDG conformer embe

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3.4k
CategoryContent
Updated5d ago
Forks412

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions