umap-learn
UMAP dimensionality reduction for visualization, clustering prep, and feature engineering. Fast nonlinear manifold learning preserving local and global structure.
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill umap-learnInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Project & Program ManagementSupported Platforms
Our assessment of umap-learn
umap-learn scores 91/100 on our quality scale, 12th of 81 Project & Program Management skills we index (top 15%).
Its SKILL.md is 19 KB long, well organised into 61 sections with 15 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so umap-learn is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
umap-learn compared with similar skills
All 4 of these similar skills score higher than umap-learn; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| umap-learn (this skill)by jaechang-hits | 91 | 367 | 37d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install umap-learn?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill umap-learn. The install tabs above show the steps for each supported agent. - Which AI agents does umap-learn work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is umap-learn safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is umap-learn still maintained?
- The repository was last updated 37 days ago, so umap-learn is actively maintained.
Skill content
View source on GitHubname: umap-learn description: >- UMAP dimensionality reduction for visualization, clustering prep, and feature engineering. Fast nonlinear manifold learning preserving local and global structure. Standard UMAP (fit/transform, sklearn-compatible), supervised/semi-supervised, Parametric UMAP (NN encoder/decoder, TensorFlow), DensMAP (density), AlignedUMAP (temporal/batch). 15+ distance metrics, custom Numba metrics, precomputed distances. For linear reduction use PCA; for neighborhood graphs use sklearn NearestNeighbors. license: BSD-3-Clause
UMAP-Learn
Overview
UMAP (Uniform Manifold Approximation and Projection) is a dimensionality reduction algorithm for visualization and general non-linear dimensionality reduction. It is faster than t-SNE, scales to larger datasets, preserves both local and global structure, and supports supervised learning and embedding of new data points.
When to Use
- Reducing high-dimensional data to 2D/3D for visualization
- Preprocessing for density-based clustering (HDBSCAN, DBSCAN)
- Feature engineering in ML pipelines (transform new data into learned embedding)
- Supervised/semi-supervised embedding with partial labels
- Tracking embeddings across time points or batches (AlignedUMAP)
- Density-preserving embeddings (DensMAP)
- Neural network-based embedding with custom architectures (Parametric UMAP)
- For linear dimensionality reduction use PCA (scikit-learn)
- For neighborhood-graph construction without embedding use scikit-learn NearestNeighbors
Prerequisites
pip install umap-learn
# For Parametric UMAP (neural network variant)
pip install umap-learn[parametric_umap] # requires TensorFlow 2.x
Critical: Always standardize features before applying UMAP to ensure equal weighting across dimensions.
Quick Start
import umap
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_digits
# Load and scale data
X, y = load_digits(return_X_y=True)
X_scaled = StandardScaler().fit_transform(X)
# Fit and transform
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled)
print(f"Input: {X_scaled.shape}, Output: {embedding.shape}")
# Input: (1797, 64), Output: (1797, 2)
Core API
1. Standard UMAP
Basic dimensionality reduction following scikit-learn conventions.
import umap
from sklearn.preprocessing import StandardScaler
X_scaled = StandardScaler().fit_transform(data)
# Method 1: fit_transform (single step)
embedding = umap.UMAP(
n_neighbors=15, # local neighborhood size (2-200)
min_dist=0.1, # min distance between embedded points (0.0-0.99)
n_components=2, # output dimensions
metric='euclidean', # distance metric
random_state=42, # reproducibility
).fit_transform(X_scaled)
print(f"Embedding shape: {embedding.shape}")
# Method 2: fit + access (for reuse)
reducer = umap.UMAP(random_state=42)
reducer.fit(X_scaled)
embedding = reducer.embedding_ # trained embedding
graph = reducer.graph_ # fuzzy simplicial set (sparse matrix)
# Visualization
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 6))
plt.scatter(embedding[:, 0], embedding[:, 1], c=labels, cmap='Spectral', s=5)
plt.colorbar()
plt.title('UMAP Embedding')
plt.tight_layout()
plt.savefig('umap_embedding.png', dpi=150)
2. Supervised & Semi-Supervised UMAP
Incorporate label information to guide embedding via the y parameter.
import umap
# Supervised — all labels known
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled, y=labels)
# Semi-supervised — partial labels (mark unlabeled as -1)
semi_labels = labels.copy()
semi_labels[unlabeled_indices] = -1
embedding = umap.UMAP(random_state=42).fit_transform(X_scaled, y=semi_labels)
# Control label influence with target_weight (0.0=unsupervised, 1.0=fully supervised)
reducer = umap.UMAP(
target_weight=0.7, # emphasize labels
target_metric='categorical', # for classification; use distance metric for regression
random_state=42
)
embedding = reducer.fit_transform(X_scaled, y=labels)
print(f"Supervised embedding: {embedding.shape}")
3. Transform New Data
Project unseen data into the trained embedding space.
import umap
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Fit on training data
reducer = umap.UMAP(n_components=10, random_state=42)
X_train_emb = reducer.fit_transform(X_train_scaled)
# Transform test data
X_test_emb = reducer.transform(X_test_scaled)
print(f"Train: {X_train_emb.shape}, Test: {X_test_emb.shape}")
# Works in sklearn Pipelines
from sklearn.pipeline import Pipeline
from sklearn.svm import SVC
pipeline = Pipeline([
('scaler', StandardScaler()),
('umap', umap.UMAP(n_components=10, random_state=42)),
('classifier', SVC())
])
pipeline.fit(X_train, y_train)
accuracy = pipeline.score(X_test, y_test)
print(f"Pipeline accuracy: {accuracy:.3f}")
4. Parametric UMAP
Neural network-based embedding via TensorFlow/Keras. Enables efficient transform, reconstruction, and custom architectures.
from umap.parametric_umap import ParametricUMAP
# Default architecture (3-layer, 100-neuron FC network)
embedder = ParametricUMAP(n_components=2, random_state=42)
embedding = embedder.fit_transform(X_scaled)
new_emb = embedder.transform(new_data) # fast neural network inference
print(f"Parametric embedding: {embedding.shape}")
import tensorflow as tf
from umap.parametric_umap import ParametricUMAP
# Custom encoder/decoder for autoencoder mode
input_dim = X_scaled.shape[1]
encoder = tf.keras.Sequential([
tf.keras.layers.InputLayer(input_shape=(input_dim,)),
tf.keras.layers.Dense(128, activation='relu'),
tf.keras.layers.Dense(64, activation='relu'),
tf.keras.layers.Dense(2),
])
decoder = tf.keras.Sequential([
tf.keras.layers.InputLayer(input_shape=(2,)),
tf.keras.layers.Dense(64, activation='relu'),
tf.keras.layers.Dense(128, activation='relu'),
tf.keras.layers.Dense(input_dim),
])
embedder = ParametricUMAP(
encoder=encoder, decoder=decoder, dims=(input_dim,),
parametric_reconstruction=True, autoencoder_loss=True,
n_training_epochs=10, batch_size=128,
n_neighbors=15, min_dist=0.1, random_state=42
)
embedding = embedder.fit_transform(X_scaled)
reconstructed = embedder.inverse_transform(embedding)
print(f"Reconstruction error: {np.mean((X_scaled - reconstructed)**2):.4f}")
5. DensMAP
Variant preserving local density information in the embedding.
import umap
reducer = umap.UMAP(
densmap=True, # enable DensMAP
dens_lambda=2.0, # density preservation weight
dens_frac=0.3, # fraction for density estimation
output_dens=True, # output density estimates
n_neighbors=15,
min_dist=0.1,
random_state=42
)
embedding = reducer.fit_transform(X_scaled)
# Access density estimates
original_density = reducer.rad_orig_ # density in original space
embedded_density = reducer.rad_emb_ # density in embedded space
print(f"DensMAP embedding: {embedding.shape}")
print(f"Density correlation: {np.corrcoef(original_density, embedded_density)[0,1]:.3f}")
6. AlignedUMAP
Align embeddings across multiple related datasets (time points, batches).
from umap import AlignedUMAP
# Multiple related datasets
datasets = [day1_data, day2_data, day3_data]
mapper = AlignedUMAP(
n_neighbors=15,
alignment_regularisation=1e-2, # alignment strength
alignment_window_size=2, # align with N adjacent datasets
n_components=2,
random_state=42
)
mapper.fit(datasets)
aligned_embeddings = mapper.embeddings_ # list of aligned embedding arrays
print(f"Aligned {len(aligned_embeddings)} datasets")
for i, emb in enumerate(aligned_embeddings):
print(f" Dataset {i}: {emb.shape}")
Key Concepts
Parameter Tuning Guide
| Parameter | Low | Medium (default) | High | Effect |
|-----------|-----|-------------------|------|--------|
| n_neighbors | 2-5 | 15 | 50-200 | Local detail vs global structure |
| min_dist | 0.0 | 0.1 | 0.5-0.99 | Tight clusters vs spread out |
| n_components | 2 | 2 | 5-50 | Visualization vs ML/clustering |
| spread | 0.5 | 1.0 | 2.0 | Embedding scale (with min_dist) |
Configuration by Use-Case
| Use-Case | n_neighbors | min_dist | n_components | metric | |----------|-------------|----------|-------------|--------| | Visualization | 15 | 0.1 | 2 | euclidean | | Clustering (HDBSCAN) | 30 | 0.0 | 5-10 | euclidean | | Text/document embedding | 15 | 0.1 | 2 | cosine | | Global structure | 100 | 0.5 | 2 | euclidean | | ML feature engineering | 15-30 | 0.1 | 10-50 | euclidean | | Binary/set data | 15 | 0.1 | 2 | hamming/jaccard |
Supported Metrics
Minkowski family: euclidean, manhattan, chebyshev, minkowski. Spatial: canberra, braycurtis, haversine. Correlation: cosine, correlation. Binary: hamming, jaccard, dice, russellrao, rogerstanimoto, sokalmichener, sokalsneath, yule. Special: precomputed (distance matrix), custom Numba-compiled callables.
Standard UMAP vs Parametric UMAP
| Feature | Standard | Parametric | |---------|----------|-----------| | Backend | Direct optimization | TensorFlow neural network | | Transform speed | Moderate | Fast (neural net inference) | | Inverse transform | Approximate, expensive | Decoder network, fast | | Custom architecture | No | Yes (CNNs, RNNs, etc.) | | Requirements | umap-learn | umap-learn + TensorFlow 2.x | | Best for | Quick exploration | Production pipelines, reconstruction |
Common Workflows
Workflow 1: UMAP + HDBSCAN Clustering Pipeline
import umap
import hdbscan
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import adjusted_rand_score
# Step 1: Preprocess
X_scaled = StandardScaler().fit_transform(data)
print(f"Input shape: {X_scaled.shape}")
# Step 2: UMAP for clustering (NOT visualization parameters)
reducer = umap.UMAP(
n_neighbors=30, # more global structure for clustering
min_dist=0.0, # allow tight packing
n_components=10, # higher dims preserve density better than 2D
metric='euclidean',
random_state=42
)
embedding = reducer.fit_transform(X_scaled)
# Step 3: HDBSCAN clustering
clusterer = hdbscan.HDBSCAN(min_cluster_size=15, min_samples=5)
cluster_labels = clusterer.fit_predict(embedding)
n_clusters = len(set(cluster_labels)) - (1 if -1 in cluster_labels else 0)
noise = sum(cluster_labels == -1)
print(f"Clusters: {n_clusters}, Noise: {noise}")
# Step 4: Separate 2D embedding for visualization
vis_emb = umap.UMAP(n_neighbors=15, min_dist=0.1, random_state=42).fit_transform(X_scaled)
plt.scatter(vis_emb[:, 0], vis_emb[:, 1], c=cluster_labels, cmap='Spectral', s=5)
plt.colorbar()
plt.title(f'HDBSCAN Clusters (n={n_clusters})')
plt.tight_layout()
plt.savefig('umap_clusters.png', dpi=150)
Workflow 2: Supervised Embedding for Classification
import umap
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
from sklearn.metrics import classification_report
# Split and scale
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)
# Supervised UMAP for feature engineering
reducer = umap.UMAP(n_components=10, random_state=42)
X_train_emb = reducer.fit_transform(X_train_s, y=y_train)
X_test_emb = reducer.transform(X_test_s)
# Downstream classifier
clf = SVC(kernel='rbf')
clf.fit(X_train_emb, y_train)
y_pred = clf.predict(X_test_emb)
print(classification_report(y_test,
Truncated for display — read the full file on GitHub.
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
