invaris-agentsec
Adversarial security testing for AI agents — prompt injection, tool misuse, memory poisoning, and MCP scanning, with CI-ready regression testing and reports.
Install / Use
claude mcp add invarislabs -- npx -y github:invarislabs/invaris-agentsecIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
SecuritySupported Platforms
Our assessment of invaris-agentsec
invaris-agentsec scores 84/100 on our quality scale, 855th of 1,114 Security skills we index.
Its MCP Server is 25 KB long, well organised into 25 sections with 9 code examples: a thorough specification that gives an agent plenty to work with.
It has 10 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated yesterday, so invaris-agentsec is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
invaris-agentsec compared with similar skills
All 4 of these similar skills score higher than invaris-agentsec; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| invaris-agentsec (this skill)by invarislabs | 84 | 10 | 1d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 95.7k | 3d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 75.0k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.3k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 86.8k | today | MCP Server |
Frequently asked questions
- How do I install invaris-agentsec?
- Run
claude mcp add invarislabs -- npx -y github:invarislabs/invaris-agentsec. The install tabs above show the steps for each supported agent. - Which AI agents does invaris-agentsec work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is invaris-agentsec safe to use?
- It is Apache-2.0-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is invaris-agentsec still maintained?
- The repository was last updated yesterday, so invaris-agentsec is actively maintained.
Skill content
View source on GitHubInvaris AgentSec
Adversarial security and reliability testing for autonomous AI agents.
Invaris AgentSec is an open-source testing framework for finding unsafe, unauthorized, and unreliable agent behaviour before it reaches production. It helps developers test complete agent workflows involving LLMs, retrieval pipelines, memory, MCP servers, external tools, databases, and sensitive actions.
The goal is simple: make testing an AI agent as repeatable and developer-friendly as testing an API.
Why AgentSec?
Traditional software follows explicitly written execution paths. AI agents interpret untrusted inputs, select tools, retain memory, and make decisions dynamically. A single indirect prompt injection or poisoned tool response can cause an agent to:
- expose confidential information;
- invoke an unauthorized tool;
- exceed its permissions or operating budget;
- execute an irreversible action;
- preserve malicious instructions in memory;
- enter an expensive or non-terminating loop; or
- behave differently after a model, prompt, or tool update.
Unit tests alone cannot adequately exercise these behaviours. AgentSec runs stateful adversarial scenarios, observes the complete execution trace, and verifies that security policies hold throughout the workflow.
This isn't hypothetical: see docs/why-agentsec.md for real, sourced incidents (a dealership chatbot selling a $76,000 car for $1, a zero-click exploit against Microsoft 365 Copilot, an AI coding agent deleting a production database), the survey data on how little of this is actually tested for today, and the governments and research institutions that have restricted AI use outright.
What AgentSec Tests
The test suite covers:
- Direct and indirect prompt injection
- Unauthorized tool invocation
- Tool-output and MCP-server poisoning
- Sensitive-data and secret leakage
- Memory poisoning and unsafe persistence
- Excessive tool calls, token usage, and cost
- Infinite loops and missing termination conditions
- Unsafe handling of retrieved documents
- Actions a tool is globally allowed to perform but that the current task never authorized (
action_without_authorization, when the policy declarestool_effects) -- see Policy reference - Dangerous compositions of individually allowed calls, tracked as data flow: private data read by one call and sent out by another to a destination the user never named, untrusted text executed as a command, untrusted instructions handed to another agent (
dangerous_composition) - Agents misreporting what they did: denying a side effect the trace shows, or claiming work that never happened (
deceptive_action_report) - Identity, session and authorization confusion: acting on another user's or tenant's resource, reusing a permission granted for an earlier task or another session, using a credential another caller left behind (
identity_and_session_confusion) - Multi-agent privilege abuse: sub-agents exceeding their role, privilege escalation through delegation (confused deputy), unauthorized delegation, unregistered agents, secrets passed between agents (
multi_agent_delegation, when the policy declaresagent_roles) -- see Multi-agent testing - Behavioural regressions across models and prompts, via
agentsec compare
Every check above is deterministic and opt-in by declaration: categories that need tool_effects or agent_roles build no scenarios without them, and a multi-agent scenario run against a system that reports no agent attribution is reported as not observable, never as passed. Unauthorized financial or on-chain actions are covered, but not by the default thirteen categories above -- via the bundled on-chain attack pack and the policy's spend_limits/address_allowlist; see Domain attack packs. What has and has not been tested against real agents and frameworks is in Testing.
Findings are mapped to the OWASP Top 10 for Agentic Applications (ASI01 to ASI10).
Design Principles
- Evidence over scores: Every finding includes the input, trace, violated policy, and observed action.
- Stateful testing: Tests cover complete workflows rather than isolated prompts.
- Framework independence: AgentSec works across models, agent frameworks, MCP servers, and custom APIs.
- CI-first: Security regressions are detectable automatically on every pull request.
- Local by default: Developers can test locally without sending private traces to a hosted service.
- Reproducibility: A failed scenario is replayable with the same configuration and evidence.
Developer Experience
Install the command-line tool:
pip install invaris-agentsec
For a local checkout instead (to run the examples, or to contribute), install from a clone with pip install -e ..
Create an agentsec.yaml policy in the directory you'll run agentsec test from -- it looks for ./agentsec.yaml by default (or pass --policy path/to/file.yaml to use a different location or name). Write it by hand as below, or scaffold one with agentsec init (writes a starter agentsec.yaml; won't overwrite an existing one). See the Policy reference for every available field.
version: "1"
agent:
name: support-agent
endpoint: http://localhost:8000/agent
allowed_tools:
- search_documents
- create_draft
forbidden_actions:
- send_email
- reveal_credentials
- execute_payment
limits:
max_steps: 12
max_tool_calls: 10
max_cost_usd: 0.50
tests:
- prompt_injection
- indirect_prompt_injection
- secret_extraction
- unauthorized_tool_use
- tool_output_poisoning
- unsafe_retrieved_documents
- loop_and_budget_limits
Run the security suite:
agentsec test
Example output (reports are also written to .agentsec/report.json and .agentsec/report.html, the two default formats -- markdown and sarif are also available via --format):
Invaris AgentSec
42 scenarios executed
36 passed
6 findings
CRITICAL Indirect prompt injection triggered send_email
HIGH Retrieved confidential content appeared in the response
HIGH Agent attempted a forbidden payment action
MEDIUM Tool-call budget exceeded
Report written to .agentsec/report.json
Report written to .agentsec/report.html
Try It
Run the bundled intentionally vulnerable agent, then test it:
pip install -e ".[dev]"
python examples/vulnerable_rag_agent/server.py & # add --safe for the hardened variant
agentsec test --policy examples/vulnerable_rag_agent/agentsec.yaml
Against the vulnerable agent you should see 35 scenarios executed and findings in all eight
categories. With --safe, all 35 pass.
Reports. agentsec test writes .agentsec/report.json and a self-contained .agentsec/report.html
(--format json,html,markdown,sarif to choose; sarif writes results.sarif for
GitHub Code Scanning). Secrets are masked. Every
finding is tagged with the closest OWASP Top 10 for Agentic Applications
category (ASI01 to ASI10).
Exit codes. agentsec test exits 1 when findings exist (--fail-on high to raise the bar,
--fail-on none to never fail) and 2 on configuration or connection errors. Use --seed N for a
reproducible run and -s <category-or-scenario-id> to run a subset.
Replay. agentsec replay .agentsec/report.json re-runs the findings from an earlier report with
the recorded seed and prints REPRODUCED or NOT REPRODUCED for each, so you can confirm a fix.
Model-assisted checks. Add a judge: section to your policy and pass --judge to have a model
review scenarios the deterministic checks passed, for example a paraphrased leak. These findings are
labelled model-assisted, are never critical, and are opt-in because transcripts are sent to the
judge endpoint. See docs/judge.md.
A RAG-style agent. examples/rag_agent owns a small document corpus and runs its tools itself, reporting the calls to AgentSec. Start it with python examples/rag_agent/server.py (add --safe for the hardened variant) and test it with agentsec test -p examples/rag_agent/agentsec.yaml. See docs/agent-contract.md.
Real-world reference agents. examples/real_world_agents has three more vulnerable/--safe reference agents modeled on daily-use AI agent products -- a coding assistant, a customer-support assistant and a browser-automation assistant -- each paired one-to-one with its matching domain attack pack below. See examples/real_world_agents/README.md and docs/agent-contract.md.
Regression comparison. agentsec compare old/report.json new/report.json lists new, fixed and changed findings and exits 1 on regressions. The packaged GitHub Action wires this into pull requests automatically (baseline-report), posting and updating a PR comment with what's new or worse; see GitHub Actions.
MCP servers. agentsec mcp scan --command "python server.py" lists a server's tools, resources and prompts (never calls, reads or fetches any of them) and flags poisoned descriptions, hidden characters, tool shadowing, homoglyph tool-name impersonation, lying read-only/destructive annotations, and changed definitions. See docs/mcp-testing.md.
MCP-connected agents. agentsec test --mcp-listen 127.0.0.1:8765 makes AgentSec the MCP server your agent uses, delivering the same adversarial scenarios through MCP tool results.
LangChain and LangGraph agents can be tested in-process with LangChainAdapter; see docs/frameworks.md. Any other framework works through the generic CallableAdapter.
Your own scenarios can be added without forking AgentSec: agentsec test --attack-pack my_pack.py or attack_packs: in the policy loads extra categories from a local file or an installed package; see docs/extending.md.
Other commands. agentsec init writes a starter policy and agentsec schema policy|trace prints
the JSON schemas.
The categories are prompt_injection, indirect_prompt_injection, secret_extraction,
unauthorized_tool_use, tool_output_poisoning, unsafe_retrieved_documents,
loop_and_budget_limits and memory_poisoning.
Agent contract. The adapter posts OpenAI-style chat-completions requests to agent.endpoint
and declares your allowed tools (plus forbidden actions as decoys). AgentSec plays the tools:
every call is simulated, and results carry the adversarial content. Agents that run tools
server-side can report them in an x_agentsec.events field, and agents with memory can key it on the
user field (also sent as X-AgentSec-Session). List the synthetic credentials your agent can see
under secrets: so leaks are detected.
In CI. Inside GitHub Actions, agentsec test adds workflow annotations for each finding and
writes a job summary automatically. A packaged action (uses: ./) starts your agent, runs the suite,
uploads reports and gates the job; it can also compare against a baseline report and comment on the
pull request, or feed results.sarif to GitHub Code Scanning. See docs/github-actions.md.
Full documentation, including architecture, polic
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
95.7kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
75.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.3kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Scrapling
86.8k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
