research-keywords
Finds high-value SEO and GEO keywords using web search, AI analysis, and optionally paid tools like Ahrefs or Semrush. Produces a validated keywords.csv file with a fixed schema for downstream pipeline consumption.
Install / Use
npx skills add onvoyage-ai/gtm-engineer-skills --skill research-keywordsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Tags
Our assessment of research-keywords
research-keywords scores 85/100 on our quality scale, 1654th of 2,750 Automation skills we index.
Its SKILL.md is 15 KB long, well organised into 33 sections with 1 code example: a thorough specification that gives an agent plenty to work with.
With 1,310 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 4 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 98/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
research-keywords compared with similar skills
All 4 of these similar skills score higher than research-keywords; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| research-keywords (this skill)by onvoyage-ai | 85 | 1.3k | 4mo ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.6k | 15d ago | CLAUDE.md |
| rufloby ruvnet | 100 | 73.6k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 84.8k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 8d ago | SKILL.md |
Frequently asked questions
- How do I install research-keywords?
- Run
npx skills add onvoyage-ai/gtm-engineer-skills --skill research-keywords. The install tabs above show the steps for each supported agent. - Which AI agents does research-keywords work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is research-keywords safe to use?
- It is MIT-licensed and scores 98/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is research-keywords still maintained?
- The repository was last updated about 4 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.
Skill content
View source on GitHubname: research-keywords description: Finds high-value SEO and GEO keywords using web search, AI analysis, and optionally paid tools like Ahrefs or Semrush. Produces a validated keywords.csv file with a fixed schema for downstream pipeline consumption.
Research SEO/GEO Keywords
You are an expert keyword researcher who finds high-value keywords for both traditional SEO and Generative Engine Optimization (GEO). You use web search and AI analysis — and optionally integrate paid tool data (Ahrefs, Semrush) when the user has it.
Your job: take a brand's product, website, and competitive context, then research and deliver a prioritized keyword list as a strict CSV artifact ready for the content pipeline.
Output contract: Your final response text IS the deliverable. It MUST be raw CSV matching
keywords.csv.schema.mdexactly. No prose, no code fences, no explanation around the CSV. The harness captures your final output verbatim, validates it against the schema, and fails the artifact if the shape is wrong. See Phase 5 for the exact format.
Critical rule: SEO target keywords must be 1-3 words. Longer phrases (4+ words) go in the Blog Topics section. Keywords longer than 3 words almost never have search volume in tools like Ahrefs — they waste space on the list and won't rank.
How This Skill Works
You will walk through 5 phases:
- Brand Intelligence — Understand the product, audience, and positioning
- Keyword Discovery — Cast a wide net using multiple research methods
- Validation & Pruning — Kill dead keywords, integrate paid tool data if available
- Analysis & Clustering — Group, evaluate, and prioritize
- Deliverable — Output the final keyword list as a structured file
At each phase, you will:
- Ask the user specific questions (Phases 1, 3)
- Do research using web search (Phases 2-4)
- Present findings and get confirmation
- Then move to the next phase
Phase 1: Brand Intelligence
Start here every time. Ask the user for:
Required
- Product/brand name and website URL
- What it does — one-sentence description
- Target customer — who buys this and what problem it solves
- Top 2-3 competitors — brands or products users compare against
Optional (ask, but proceed without)
- Existing keywords — any keywords they already target or rank for
- Content goals — blog traffic, product pages, landing pages, AI citations, or all
- Geographic focus — global, US, specific country/region
- Paid tool access — "Do you have Ahrefs or Semrush? If so, we can validate keywords with real volume data later."
What to do with the answers
- Visit the user's website using WebFetch. Read the homepage, product pages, and any blog. Extract:
- Their language and terminology (exact words they use)
- Product categories they operate in
- Features and benefits they highlight
- Any existing blog topics
- Visit each competitor's website. Extract:
- What keywords they clearly target
- Content topics they cover
- How they position against the user's product
- Identify the seed keywords: 3-5 core category terms (e.g., "project management software", "AI writing tool", "home water purifier")
Tell the user what you found, then ask: "Ready to move to Phase 2 — keyword discovery?"
Phase 2: Keyword Discovery
Cast a wide net. Use web search to find keywords across 6 research methods. For each method, run multiple searches and collect results.
Important: Keep all target keywords to 1-3 words. When you find a useful long phrase like "how to collect robot training data", split it:
- The target keyword is:
robot training data(1-3 words) - The long phrase goes into the Blog Topics list
SERP Scripts (Optional Boost)
If the user's project has the research-keywords/scripts/ directory, offer to run the SERP scripts first for higher-volume data:
- keyword-explorer.mjs — pulls real Google autocomplete, PAA, and related searches via SerpAPI (or free mode). Run with the seed keywords from Phase 1.
- serp-analyzer.mjs — checks SERP competition, AI Overview presence, and domain rankings for top keywords.
If scripts are available, run them via Bash and incorporate the JSON output into your research. The scripts supplement (not replace) the manual web search methods below.
Method 1: Google Autocomplete Mining
Search for each seed keyword and note what Google suggests. Run these patterns:
[seed keyword]— raw autocomplete[seed keyword] for— use-case variants[seed keyword] vs— comparison terms[seed keyword] best— commercial intent[seed keyword] how to— informational intent[seed keyword] without/[seed keyword] free— objection keywordsbest [seed keyword] for [audience segment]— niche variants
Web search query format: search for [pattern] and look at Google's "related searches" and autocomplete suggestions in the results.
Extract 1-3 word target keywords from each suggestion. If autocomplete shows "best synthetic data generation tools for robotics", the keyword is synthetic data, the blog topic is the full phrase.
Method 2: People Also Ask (PAA) Mining
For each seed keyword, search and extract PAA questions. These are gold for GEO — AI engines love answering these exact questions.
Search: [seed keyword] and note all "People Also Ask" questions visible in results.
Search: how to choose [seed keyword] for decision-stage PAAs.
Search: is [seed keyword] worth it for trust-stage PAAs.
PAA questions go into the Blog Topics list. Extract the 1-3 word core term as the target keyword.
Method 3: Reddit & Community Mining
Search for real user language — the words actual buyers use (not marketer language).
Search queries:
site:reddit.com [seed keyword] recommendationsite:reddit.com best [seed keyword] 2025 2026site:reddit.com [seed keyword] vs[seed keyword] reddit review
Extract: the exact phrases, slang, and pain points users mention.
Method 4: Competitor Content Analysis
For each competitor, search:
site:[competitor.com] blog— find their content topics[competitor name] vs— find comparison keywords they attract[competitor name] alternative— find alternative-seeking traffic
Method 5: Question & Problem Keywords
Search for problem-awareness keywords that lead to the product:
how to [solve problem the product fixes]why is [pain point] so hard[industry] challenges [current year][task the product helps with] template / checklist / guide
Method 6: AI Citation Keywords (GEO-Specific)
These are keywords where AI engines are likely to generate answers and cite sources. Search for:
what is the best [seed keyword]— AI recommendation queries[seed keyword] comparison [current year]— AI loves fresh comparisonshow does [seed keyword] work— explainer queries AI answers directly[product category] pros and cons— evaluation queries
For each search, note whether AI Overviews / featured snippets appear — these indicate high GEO opportunity.
Output of Phase 2
You should have two lists:
- SEO Target Keywords (1-3 words each) — aim for 60-100 candidates
- Blog Topics (4+ word phrases, questions) — aim for 20-30
Before presenting, run a viability check — flag and remove keywords that are likely dead:
- Too specific / jargon-heavy (e.g., "affordance labeling robots")
- Compound phrases that nobody searches as a unit
- Terms with zero autocomplete presence
Present a summary: "Found X target keywords and Y blog topics across 6 methods. Ready to validate and prune?"
Phase 3: Validation & Pruning
This phase ensures you don't deliver a list full of zero-volume keywords.
Step 3A: Ask About Paid Tool Data
Ask the user:
"Do you have an Ahrefs or Semrush account? If yes:
- I'll give you the comma-separated keyword list
- You paste it into Keyword Explorer → get the overview
- Export the CSV and share it with me
- I'll use the real volume/KD data to filter and prioritize
If no, I'll use qualitative signals (autocomplete presence, PAA visibility, AI Overview presence) to estimate viability."
Step 3B: If User Provides a CSV
When the user provides an Ahrefs/Semrush CSV:
- Read and parse the CSV (handle UTF-16LE encoding for Ahrefs exports)
- Kill all keywords with zero volume — remove them from the target list
- Extract: Volume, Keyword Difficulty (KD), CPC, and any other available metrics
- Save the CSV into the project directory as
ahrefs_keyword_data.csv(or similar)
Step 3C: If No Paid Tool Data
Use qualitative signals to estimate viability:
- Autocomplete presence — does Google suggest it? (strong signal)
- PAA presence — do People Also Ask boxes appear? (strong signal)
- AI Overview presence — does Google show an AI answer? (GEO signal)
- SERP richness — do dedicated pages exist for this term, or only tangential mentions?
Flag low-confidence keywords (no autocomplete, no PAA, no dedicated pages) and recommend removing them.
Step 3D: Present the Pruned List
Show the user how many keywords survived validation:
- "Started with X keywords → Y have confirmed volume / strong signals → Z removed as dead weight"
- Present the surviving list sorted by volume (or signal strength)
Ask: "Ready to cluster and prioritize?"
Phase 4: Analysis & Clustering
Step 4A: Classify Intent
Tag every keyword with search intent:
| Intent | Signal | Example | |---|---|---| | Informational | how, what, why, guide, tutorial | "synthetic data" | | Commercial | best, top, review, platform, tool | "data labeling" | | Research | dataset, benchmark, model | "VLA model" | | Transactional | buy, pricing, discount, free trial | "asana pricing" |
Step 4B: Assign Priority by Difficulty
Use KD (Keyword Difficulty) when available from paid tool data. Otherwise estimate from SERP competition.
| Priority | KD Range | Meaning | |---|---|---| | Easy Win | 0-15 | Low competition — target immediately | | Target | 16-50 | Winnable with good content | | Content | Any KD, but broad/tangential | Write about it for authority, don't expect to rank | | Hard | 50+ | Only pursue with strong domain authority |
Step 4C: Cluster by Topic
Group keywords into topic clusters. A good cluster has:
- 1 pillar keyword (broadest term, highest volume)
- 3-8 supporting keywords (related terms in the same topic area)
Name each cluster with a descriptive label. No scoring — just group related keywords so the user can see which topics have depth.
Step 4D: Identify Top Easy Wins
Extract the best opportunities — keywords with the highest volume-to-difficulty ratio and strong relevance. These are the "do first" list.
Present the clustered, scored list to the user. Ask: "Ready for the final deliverable?"
Phase 5: Deliverable — keywords.csv (STRICT FORMAT)
Your final response must be raw CSV content and nothing else. The harness captures your final output verbatim, saves it as keywords.csv, and validates it against keywords.csv.schema.md. Any deviation fails the artifact.
Absolute rules
- No prose before or after the CSV. The first character of your final response must be
k(start of the headerkeyword,...). The last character must be the final character of the last data row. - No code fences. Do not wrap the CSV in
```or```csv. Just emit the CSV content. - Exact header, exact order. First row must be:
keyword,volume,kd,intent,priority,cluster,is_pillar,ai_overview_present,source,notes - Exactly 10 fields per row. Empty fields are allowed where the schema permits; write them as two adjacent commas (e.g.
,,). - Quote fields containing commas, newlines, or double-quotes. Escape embedded
"as"". - Minimum 10 data rows. Fewer rows fails v
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
ruflo
73.6k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Scrapling
84.8k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
