geo-scope
Open framework for empirical AI visibility benchmarks, multi-model provider observation, and reproducible GEO research.
Install / Use
claude mcp add tmolavi -- npx -y github:tmolavi/geo-scopeIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of geo-scope
geo-scope scores 72/100 on our quality scale, 872nd of 955 AI & Machine Learning skills we index.
Its MCP Server is 18 KB long, well organised into 32 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 2 days ago, so geo-scope is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
geo-scope compared with similar skills
All 4 of these similar skills score higher than geo-scope; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| geo-scope (this skill)by tmolavi | 72 | 3 | 2d ago | MCP Server |
| claude-memby thedotmack | 100 | 98.1k | 1d ago | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 93.9k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.6k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.7k | today | CLAUDE.md |
Frequently asked questions
- How do I install geo-scope?
- Run
claude mcp add tmolavi -- npx -y github:tmolavi/geo-scope. The install tabs above show the steps for each supported agent. - Which AI agents does geo-scope work with?
- It is written for Claude Code, Claude Desktop and Gemini CLI, as a MCP Server file. Other agents that read the same format can often use it too.
- Is geo-scope safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is geo-scope still maintained?
- The repository was last updated 2 days ago, so geo-scope is actively maintained.
Skill content
View source on GitHub⟠ GEO-Scope
Empirical AI Answer Visibility Measurement Framework
An open-source framework for empirical measurement of AI answer visibility, entity mentions, recommendations, and citations across generative AI systems.
Introduction • What It Measures • Scientific Foundation • Benchmarks • Reproducibility • Research • Quickstart • MCP
</div>1. Introduction
Generative AI systems and search-grounded answer engines are rapidly becoming the primary discovery layer for users seeking products, vendors, services, and factual insights.
GEO-Scope is an evidence-first, open-source measurement framework designed to empirically quantify and preserve auditable evidence of how generative AI systems surface entities. It records, normalizes, and analyzes observable AI completions under documented, neutral prompt sets without relying on speculative ranking algorithms or ungrounded claims.
Core Observable Outputs Measured:
- Entity Mentions: Observable presence of brands, products, technologies, and public figures in generated text.
- Recommendations: Explicit linguistic endorsements and ordered top-position recommendations.
- Citations: Grounding source URLs and referenced web domains returned by search-augmented models.
- Attribution: Textual credit linking specific facts, statistics, or claims to source entities.
- Provider Differences: Distributional shifts between live search-grounded answer engines and parametric foundation LLMs.
- Multilingual Behavior: Cross-lingual response variations across 26+ evaluated languages.
2. What GEO-Scope Measures
GEO-Scope enforces a strict taxonomic separation between four independent visibility dimensions:
┌─────────────────────────────────────────────────────────────────────────┐
│ AI RESPONSE VISIBILITY MATRIX │
├───────────────────┬─────────────────────────────────────────────────────┤
│ Mention │ Did the entity appear anywhere in the completion? │
│ Recommendation │ Was the entity explicitly endorsed or recommended? │
│ Citation │ Was a source URL or grounding domain link provided? │
│ Attribution │ Was specific data/claim textually credited to it? │
│ Rank │ Extracted ONLY when a valid ordered list exists. │
└───────────────────┴─────────────────────────────────────────────────────┘
- Mention (
mentioned: true/false):- Captures whether the target entity (or associated canonical aliases/founders) appeared in the generated completion.
- Evaluated via Unicode NFKC normalization, Arabic/Persian letter unification, Zero-Width Non-Joiner (ZWNJ) handling, and negative homonym collision filtering.
- Recommendation (
recommended: true/false):- Strict rule:
mentioned != recommended. - Evaluated based on explicit linguistic recommendation markers (e.g., "We recommend...", "Top pick", "گزینه پیشنهادی") or inclusion in an ordered list answering a recommendation query.
- Strict rule:
- Citation (
cited: true/false):- Identifies presence of target entity web domains in grounding references, markdown hyperlinks, or structured provider citation chunks.
- Attribution (
attributed: true/false):- Distinct from citation: detects explicit textual sourcing phrasing (e.g., "According to [Entity]...", "طبق گزارش [موجودیت]") even if an active URL link was omitted by the model.
- Rank (
rank: 1..N | null):- Extracted strictly from numbered lists or ordinal items. If an informational question yields an unranked mention, rank is set to
nullto prevent artificial ranking bias.
- Extracted strictly from numbered lists or ordinal items. If an informational question yields an unranked mention, rank is set to
3. What GEO-Scope Does NOT Measure
To maintain scientific integrity, GEO-Scope clearly outlines its epistemic boundaries:
- ❌ It does NOT reverse-engineer internal ranking algorithms: GEO-Scope observes external API completions; it cannot inspect internal model weights, attention matrices, or proprietary ranking formulas.
- ❌ It does NOT inspect hidden training data: Observed entity knowledge reflects generated outputs, not full visibility into private training corpora.
- ❌ It does NOT claim causal ranking factors: All reported metrics represent descriptive statistical associations under documented prompts, not causal guarantees.
- ❌ It does NOT guarantee SEO or AI visibility improvements: Measurements provide historical observation, not predictive visibility promises.
- ❌ It does NOT treat simulation as live empirical data: Simulation fixtures are strictly quarantined for testing and CI.
4. Architecture
GEO-Scope operates as a modular, six-stage evidence pipeline:
┌─────────────────────────────────────────────────────────────┐
│ Provider Layer │
│ ┌──────────────────────────┐ ┌────────────────────────┐ │
│ │ Search Answer Engines │ │ Parametric Base LLMs │ │
│ │ (Perplexity, Gemini...) │ │ (OpenAI, Claude...) │ │
│ └──────────────────────────┘ └────────────────────────┘ │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Measurement Engine │
│ (Prompt Provenance · Zero Silent Fallback · Runs) │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Raw Response Storage │
│ (Unparsed API Payloads · Latency · Token Usage) │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Observation Parser │
│ (Multi-Lingual Normalizer · Homonyms · Citations) │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Metrics Calculation │
│ (OMR · Rec Share · Citation Rate · Honest Denominators) │
└──────────────────────────────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Reports + Replayable Bundle │
│ (JSONL Bundles · SHA-256 Checksums · Markdown Summaries) │
└─────────────────────────────────────────────────────────────┘
5. Execution Modes
GEO-Scope provides three mutually exclusive execution modes:
demo (Simulation Fixture)
- Purpose: Rapid offline testing, development fixtures, and CI validation.
- Behavior: Uses local mock completions with prefixed IDs (
simulated_*) and a clear simulation banner. - Guarantee: Simulation data is strictly rejected by the release quality gate and can never enter published empirical benchmarks.
measure (Live Empirical Execution)
- Purpose: Real-world observation runs against live generative AI endpoints.
- Behavior: Dispatches neutral prompt bundles to configured API providers with zero silent fallback.
- Preservation: Saves full unmodified payloads to
raw_responses.jsonlwith exact model governance metadata (requested_provider,actual_provider,search_grounded).
replay (Deterministic Offline Replay)
- Purpose: Independent auditability and benchmark verification without API calls or cost.
- Behavior: Reruns the observation parser and metric calculations directly against preserved
raw_responses.jsonl. - Integrity: Verifies that recomputed metrics match published results bit-for-bit.
6. Measurement Contract v1 & Scientific Foundation
GEO-Scope does not claim universal AI visibility truth. It measures empirical observations under declared, reproducible measurement configurations.
[!IMPORTANT] Fundamental Measurement Axiom
"AI visibility is an observation under a declared measurement system, not a universal ground-truth ranking."
- 🏛️ Scientific Measurement Gate:
docs/SCIENTIFIC_MEASUREMENT_GATE.md - 📄 Methodology:
docs/METHODOLOGY.md - 📐 Machine-Readable Schema:
schemas/v0.3/manifest.schema.json - 🧪 Validation Example Fixture:
examples/public_demo/manifest.json - 📊 Documentation Index:
docs/index.md
Core Measurement Principles
- Mention Definition: A response-level binary observation indicating whether the target entity appears at least once in the completion. Multiple mentions in a single answer do not artificially inflate response-level mention counts.
- Citation Separation: Strict 4-way separation between
entity_mentionedin text,target_domain_cited(root domain),target_url_cited(deep link), andthird_party_source_cited(external authority/review links). Mention and citation are never treated as equivalent. - Recommendation Semantics: Evaluated as true only when the model semantically recommends, selects, or endorses the entity. Ambiguous detections are gated and marked
experimental. - Comparability Rules: Machine-readable comparability verification. Two studies are marked
comparable: trueonly when prompt universe, market/language, provider/model family, measurement definitions, and observation windows match. - Raw Evidence Traceability: Every public observation is linked to prompt ID, raw response or cryptographic SHA-256 hash (
response_hash_sha256), extracted entities, citations, and execution configuration hash.
7. Published Benchmark Releases
GEO-Scope maintains immutable, peer-review-ready benchmark releases under benchmark/releases/ (see the documentation index):
| Benchmark Release | Prompt Count | Observations | Providers | Cryptographic Status | Documentation | |:---|:---|:---|:---|:---|:---| | [`global-ai-an
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
98.1kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
93.9kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.6kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.7kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
