embedding-strategies
Select and optimize embedding models for semantic search and RAG applications
Install / Use
npx skills add wshobson/agents --skill embedding-strategiesInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of embedding-strategies
embedding-strategies scores 89/100 on our quality scale, 197th of 628 AI & Machine Learning skills we index (top 32%).
Its SKILL.md is 2.8 KB long, well organised into 9 sections with 1 code example: a solid amount of guidance for an agent.
With 39,920 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so embedding-strategies is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-09-26. Automated pattern scan on 2026-09-25. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
embedding-strategies compared with similar skills
All 4 of these similar skills score higher than embedding-strategies; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| embedding-strategies (this skill)by wshobson | 89 | 39.9k | 5d ago | SKILL.md |
| claude-memby thedotmack | 100 | 94.7k | 1d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 84.2k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 73.8k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.1k | today | CLAUDE.md |
Frequently asked questions
- How do I install embedding-strategies?
- Run
npx skills add wshobson/agents --skill embedding-strategies. The install tabs above show the steps for each supported agent. - Which AI agents does embedding-strategies work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is embedding-strategies safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is embedding-strategies still maintained?
- The repository was last updated 5 days ago, so embedding-strategies is actively maintained.
Skill content
View source on GitHubname: embedding-strategies description: Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.
Embedding Strategies
Guide to selecting and optimizing embedding models for vector search applications.
When to Use This Skill
- Choosing embedding models for RAG
- Optimizing chunking strategies
- Fine-tuning embeddings for domains
- Comparing embedding model performance
- Reducing embedding dimensions
- Handling multilingual content
Core Concepts
1. Embedding Model Comparison (2026)
| Model | Dimensions | Max Tokens | Best For | | -------------------------- | ---------- | ---------- | ----------------------------------- | | voyage-3-large | 1024 | 32000 | Claude apps (Anthropic recommended) | | voyage-3 | 1024 | 32000 | Claude apps, cost-effective | | voyage-code-3 | 1024 | 32000 | Code search | | voyage-finance-2 | 1024 | 32000 | Financial documents | | voyage-law-2 | 1024 | 32000 | Legal documents | | text-embedding-3-large | 3072 | 8191 | OpenAI apps, high accuracy | | text-embedding-3-small | 1536 | 8191 | OpenAI apps, cost-effective | | bge-large-en-v1.5 | 1024 | 512 | Open source, local deployment | | all-MiniLM-L6-v2 | 384 | 256 | Fast, lightweight | | multilingual-e5-large | 1024 | 512 | Multi-language |
2. Embedding Pipeline
Document → Chunking → Preprocessing → Embedding Model → Vector
↓
[Overlap, Size] [Clean, Normalize] [API/Local]
Templates and detailed worked examples
Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.
Best Practices
Do's
- Match model to use case: Code vs prose vs multilingual
- Chunk thoughtfully: Preserve semantic boundaries
- Normalize embeddings: For cosine similarity search
- Batch requests: More efficient than one-by-one
- Cache embeddings: Avoid recomputing for static content
- Use Voyage AI for Claude apps: Recommended by Anthropic
Don'ts
- Don't ignore token limits: Truncation loses information
- Don't mix embedding models: Incompatible vector spaces
- Don't skip preprocessing: Garbage in, garbage out
- Don't over-chunk: Lose important context
- Don't forget metadata: Essential for filtering and debugging
Related Skills
claude-mem
94.7kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Understand-Anything
84.2kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
73.8kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.1kOpen-source super AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
