DataAgentConnector
Blueprint for spinning up read-only SQL data agents from any SQLAlchemy-compatible database. Indexes schema and column content into LanceDB and exposes MCP + REST tools so LLMs and UIs can safely explore your data.
Install / Use
claude mcp add MagnusS0 -- npx -y github:MagnusS0/DataAgentConnectorIf the server publishes to npm under a different name, use that package instead β check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of DataAgentConnector
DataAgentConnector scores 70/100 on our quality scale, 894th of 968 AI & Machine Learning skills we index.
Its MCP Server is 4.6 KB long, well organised into 8 sections with 1 code example: a solid amount of guidance for an agent.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated about 9 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 86/100, with 2 cautions from licensing, adoption, age or documentation. These come from repository metadata, not a code audit β read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-09. It catches known dangerous patterns, not every risk β read a skill before letting an agent act on it.
DataAgentConnector compared with similar skills
All 4 of these similar skills score higher than DataAgentConnector; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| DataAgentConnector (this skill)by MagnusS0 | 70 | 3 | 9mo ago | MCP Server |
| claude-memby thedotmack | 100 | 98.6k | today | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 94.3k | 1d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.7k | 3d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.8k | today | CLAUDE.md |
Frequently asked questions
- How do I install DataAgentConnector?
- Run
claude mcp add MagnusS0 -- npx -y github:MagnusS0/DataAgentConnector. The install tabs above show the steps for each supported agent. - Which AI agents does DataAgentConnector work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is DataAgentConnector safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is Apache-2.0-licensed and scores 86/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is DataAgentConnector still maintained?
- The repository was last updated about 9 months ago. That is recent enough to be usable, but agent tooling moves fast, so check the instructions against your agent's current version.
Skill content
View source on GitHubData Agent Connector
This provides a blueprint to go from database connection string to up and running SQL Data Agent in seconds. Connect to any database that can be used with SQLAlchemy (might need to install specific engine). Builds search and convenience tools on top of a DB, made for SQL agents to interact with through the MCP protocol. Plus mirrored REST endpoint that can be connected to UIs etc.
Key Features
- Read-only SQL gateway: SQLAlchemy engines are locked to safe commands defined in
[tool.dac.settings.allowed_sql_commands]. - Automatic metadata: LLM agents summarise tables; LanceDB stores summaries plus sentence-transformer embeddings (embeddings are currently unused).
- Column-content retrieval: Distinct textual values are sampled, filtered, and indexed with LanceDB BM25 for direct content search in columns.
- MCP + REST:
/mcpserves FastMCP, while/widgets/*exposes REST endpoints for UI integration with OpenBB (customize this to your preferred UI). - Config-driven:
databases.tomldeclares available data sources,pyproject.tomlunder[tool.dac.settings]for runtime settings and.env(or environment variables) configures the LLM provider.
MCP (Mounted at /mcp)
| Tool | Summary | | --- | --- | | get_databases | Lists registered databases and descriptions. | | show_tables / show_views | Enumerates tables/views with cached annotations where available. | | describe_table / describe_view | Returns DDL-like metadata or view SQL. | | get_distinct_values | Pulls sample categorical values (limit enforced). | | preview_table | Returns first rows of non-binary columns. | | find_relevant_columns_and_content | BM25 search over distinct textual values with score filtering. | | query_database | Executes read-only SQL with a configurable row cap (mcp_query_limit). | | join_path | Suggests shortest join sequences or Steiner-tree paths across tables. |
Getting Started
-
Clone the repository:
git clone https://github.com/MagnusS0/DataAgentConnector.git cd DataAgentConnector -
Install dependencies:
uv sync --group ai -
Configure your databases in
databases.toml:[databases.my_database] connection_string = "sqlite:///path/to/your/database.db" description = "My local SQLite database" [databases.another_database] connection_string = "postgresql://user:password@localhost:5432/another_database" description = "Another PostgreSQL database" -
Set up your LLM provider in
.env:LLM_API_KEY=your_api_key_here LLM_MODEL_NAME=default-model LLM_BASE_URL=https://api.your-llm-provider.com -
Run the application:
uv run uvicorn app.main:app --reload
Project Structure
DataAgentConnector/
βββ app/
β βββ agents/
β βββ core/
β βββ domain/
β βββ interfaces/
β βββ models/
β βββ schemas/
β βββ repositories/
β βββ services/
β βββ main.py
βββ databases.toml
βββ pyproject.toml
βββ .env
βββ README.md
Indexing & Metadata Pipeline
- Column extraction (
app/domain/extract_colum_content.py) samples distinct textual values while filtering binary, numeric, or overly long fields; tunable viatool.dac.settings.fts_extraction_options. - FTS indexing (
app/services/indexing_service.py) persists values into LanceDB tables namedcolumn_contents_<database>and builds BM25 indexes. - Annotation workflow (
app/services/annotation_service.py) runs LLM prompts with table metadata, previews, and sampled values (schema hashes used to skip already processed tables), embeddings are added via sentence-transformers.
FK Graph & Join Paths
Foreign key constraints are analyzed to build a cached CSR adjacency matrix (app/domain/fk_analyzer.py) where tables are nodes and FKs are edges. For two tables, BFS finds the shortest join sequence. For 3+ tables, an approximate Steiner tree (MST on all-pairs distances) computes the minimal spanning network, returning ordered JoinStep objects with FK column mappings.
This allows agents to request optimal join paths across multiple tables when formulating SQL queries. Even when there is no direct foreign key relationship defined in the database schema.
Stats for the interested user
Indexing and annotating all of BIRD-SQL training databases (69 databases) results in:
- Table annotations stored successfully in ~200 seconds
- Content FTS indices created successfully in ~5 seconds
Hardware: Intel i9-14900K, 64GB RAM, RTX 3090 running Menlo/Jan-nano (4B params) using vLLM
Related Skills
claude-mem
98.6kPersistent Context Across Sessions for Every Agent β Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
94.3kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu β one CLI, zero API fees.
Understand-Anything
85.7kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.8kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit β see the Safety scan above for what the skill file itself contains.
