agent-reliability-analyzer
MCP serverAI & modelsLets your agent find and fix reliability and safety gaps in agent code across nine SDKs.
Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.
Add to setup to save this item as a reference. ahel cannot run it, and signing in will not install it.
About this server
Find and fix reliability and safety gaps in agent code, across nine agent SDKs.
Getting started
- Save this item in Your setup as a reference.
- Read the source or reference documentation for its setup requirements. Saving it here does not connect it to your AI.
- Check this page for availability before trying to install it through ahel.
From the project's README
As published by trustabl/agent-reliability-analyzer in README.md.
Trustabl — find and fix AI agent reliability gaps
Find what will make your AI agent fail — then fix it with one command.
Deterministic static analysis for agent code, across nine SDKs and seven languages. It runs entirely on your machine: no cloud scanner, no account, no code upload, no LLM.
Trustabl scans an agent repository for the gaps that break agents in production: tool descriptions too vague for a model to know when to use them, missing retry and timeout handling, untyped parameters, absent guardrails, and tool grants that exceed what the agent claims to do. Then it applies the fix directly to your source.
Every other agent scanner hands you a report. Trustabl hands you a patch.
Reliability is engineered before deployment, not observed after it.
Scanning runs entirely on your machine. No cloud scanner, no account, no code upload, no LLM — deterministic static analysis.
docker run --rm -v "$PWD:/repo" ghcr.io/trustabl/trustabl:latest scan /repo # try it, nothing installed
brew install trustabl/tap/trustabl # macOS / Linux
scoop bucket add trustabl https://github.com/trustabl/scoop-bucket # Windows
scoop install trustabl
trustabl scan . # find issues (fully local)
trustabl scan . --format json > scan.json
trustabl enrich --input scan.json --repo . --diff --apply # preview, then fix
9 SDKs · 7 languages · human, JSON, or SARIF 2.1.0 output ·
CI-friendly exit codes · also runs as a local stdio MCP server (trustabl mcp).
Deepest coverage for Claude Agent SDK, OpenAI Agents SDK, Google ADK, and MCP servers. Also scans LangChain / LangGraph, CrewAI, AutoGen / AG2, Pydantic AI, and the Vercel AI SDK — see COVERAGE.md for the full matrix.
Scanning needs no key and no network. Applying fixes uses your own Anthropic, OpenAI, or Google key.
What a scan looks like
Real output, scanning an agent repo that ships a Claude Agent SDK client, an MCP server, a subagent, and two skills:
$ trustabl scan .
Scan summary
Languages: typescript, javascript
SDKs: claude_agent_sdk
Tool definitions: 2 (custom tools with function bodies)
Agent tool grants: 14 (tool names the agent may call)
MCP servers: 1 (createSdkMcpServer)
Subagents: 1 (inbox-searcher)
Skills: 2 (action-creator, listener-creator)
Findings: 14
Overall score: 66%
Surface readiness
skill:action-creator 45% (5 findings)
skill:listener-creator 54% (4 findings)
agent:AIClient.queryStream 64% (2 findings)
subagent:inbox-searcher 79% (1 finding)
tool:search_inbox 98% (1 finding)
tool:read_emails 100% (0 findings)
Findings
inbox-searcher
[CSDK-201] HIGH Subagent is granted Bash (agent/.claude/agents/inbox-searcher.md:1-5)
A subagent with shell access can run arbitrary commands if it is
compromised or misdirected.
fix: Remove `Bash` from the subagent's `tools:` list unless shell access
is essential. Prefer specific tools (Read, Grep, Glob) for read-only roles.
Three things worth noticing:
- It scores each surface separately. A repo-wide number hides the problem;
skill:action-creatorat 45% tells you where to look first. - Every finding carries a fix, not just a complaint — and
trustabl enrich --applywrites those fixes to source. - The exit code is the CI gate.
0when nothing reaches medium severity,1when something does, so a pipeline can stop a bad agent from shipping.
AI agents pass their demo and fail in production
The failures are rarely exotic. They are the same handful of gaps, over and over:
- An agent holding shell tools with no input guardrails
- A tool making an HTTP call with no timeout, hanging the whole run
- Untyped tool parameters, so the model guesses and guesses wrong
- No observability wired at all, so the first three have no trace to read
- A tool description so vague the model calls the wrong tool — 56% of MCP tool descriptions fail to state their purpose clearly, and 97.1% carry at least one description defect (Hasan et al., MCP Tool Descriptions Are Smelly!, arXiv 2602.14878 — 856 tools across 103 servers)
- A subagent granted
Bashdespite a read-only description - A skill that auto-approves unrestricted shell access
- No retry handling, so one transient 500 fails the task
- An agent loop with no iteration bound, burning tokens until it is killed
- A user-controlled URL flowing into a fetch — SSRF, from your agent
Every one of these is visible in the source before the agent ever runs. Trustabl finds them in seconds, offline, with no LLM — and fixes them.
The rest of this document explains what Trustabl reasons about and how the scan works, then covers building and running it. For the full implementation reference see ARCHITECTURE.md; for the at-a-glance SDK coverage matrix see COVERAGE.md.
What it analyzes — the five-scope model
Trustabl does not treat a repository as one undifferentiated blob. Every rule is classified into exactly one of five scopes, and each scope receives a different typed input:
tool— fires once per tool definition. Input: aToolDef(a@function_tool/@tool/@claude_toolfunction, a Claude TStool(name, description, schema, handler)factory call, aFunctionTool(fn)ADK wrapper, an@server.toolMCP registration, or a bare shell-invoking function) plus its parsed file. Catches a missing docstring, an HTTP call with no timeout, untyped parameters, or an unnormalized path flowing intoopen(). (Hosted tools likeWebSearchTool()are agent-scope edge data, captured asHostedToolDef, notToolDef.)agent— fires once per agent declaration. Input: anAgentDef— a PythonAgent(...)/SandboxAgent(...)/AgentDefinition(...)call, a Claude TS typed-constAgentDefinition, a Claude TS sub-agent inline inoptions.agents, or the Claude TSquery(...)main-thread agent (QueryMainAgent) — with every constructor kwarg captured and its edges to tools, handoffs, and guardrails resolved. Catches an agent with shell tools and noinput_guardrails,tool_use_behavior="stop_on_first_tool"paired with filesystem-touching tools, or a main-thread agent with unrestrictedallowedTools.subagent— fires once per Claude Code subagent markdown declaration. Discovery is hybrid: canonical.claude/agents/*.md(any path depth, monorepo-safe) PLUS a frontmatter-shape fallback over all markdown files (gated onname+tools/model) that catches flat-collection repos which ship subagents undercategories/*.md,plugins/<x>/agents/*.md, or similar layouts. Input: aSubagentDefparsed from frontmatter —name,description,tools[](verbatim) +ToolGrants[](parsed permission grammar),disallowedTools,model,permissionMode(incl.bypassPermissions),mcpServers,skills,isolation,hasHooks. Catches a subagent granted the built-inBashtool despite a read-only description (CSDK-110). Subagent presence alone contributesclaude_agent_sdktoSDKsDetected, so the Claude pack loads and CSDK-110 fires on pure-markdown subagent collections.skill— fires once per Claude Code skill (SKILL.md, any path depth). Input: aSkillDefparsed from frontmatter —name,description,allowed-tools→ToolGrants[],disable-model-invocation— plus body facts (dynamic-context exec commands, external URLs, prompt-injection markers) and a bundled-file inventory. Catches a skill that auto-approves unrestrictedBash(CSKILL-001), runs a dynamic-context command that performs network egress or reads secrets before the model sees it (CSKILL-003), or is model-invocable while granting side-effecting tools (CSKILL-050). Skills are markdown, so skill rules carry nolanguage:; theclaude_skillpack loads whenever aSKILL.mdis present.repo— fires once per scan against the whole inventory. Catches project-wide gaps such as the OpenAI Agents SDK being present with no custom trace processor configured.
The agent is the unit of analysis, not the repo
A repo can declare zero, one, or many agents, across one or more SDKs. Two agents in the same repo can be in completely different security postures — one wired with input/output guardrails, the other not. Agent-scoped findings therefore attribute to a specific agent at its constructor call site; flattening them to a single repo-level verdict would lose that attribution and be wrong. Discovery builds a small per-repo graph (tools, agents, subagents, and the edges between them) so agent-scope and subagent-scope rules can query it.
Rules are scoped to one SDK and one language
A Claude-SDK rule and an OpenAI-Agents-SDK rule that detect the same
conceptual problem (a missing timeout, say) are two separate rules with
SDK-specific explanation and fix text — there is no cross-SDK casting.
When a repo declares agents from multiple SDKs side by side, each agent is
checked only against the rules for the SDK that declared it. The same
holds across languages: a language: python rule will not fire on a
TypeScript agent.
How it reasons — the scanning pipeline
trustabl scans in four steps. Each step's output is the typed input to the next, with no shared state between runs — and the inventory the early steps build is what makes policy selection data-driven rather than statically configured.
The binary ships with no embedded rules. Before the pipeline runs,
Trustabl resolves its detection rules from a separate git repository
(agent-reliability-rules) —
fetching the latest, caching the clone locally, and falling back to the
cache when the network is unreachable. This decouples rule updates from
binary releases: rules can be added or changed without rebuilding the
scanner. The resolved rules commit is recorded in the result and folded
into the ScanID, so a scan is honest about which rules produced it.
If no rules can be fetched and none are cached, the scan exits 2 and
tells you to run trustabl rules pull — Trustabl never runs rule-less.
flowchart LR
target[("Agent repo<br/>(local path or GitHub URL)")]
recon["Recon<br/>files · SDK deps"]
inv["Inventory<br/>Python + TS AST:<br/>tools · agents ·<br/>subagents · MCP servers"]
pol["Policy selection<br/>load rules per<br/>detected SDK ·<br/>META findings"]
ana["Analysis<br/>tool · agent · subagent ·<br/>repo detectors"]
score["Scoring<br/>per-surface score ·<br/>overall readiness"]
out[("ScanResult<br/>findings · scores<br/>(human / JSON / SARIF)")]
target --> recon --> inv --> pol --> ana --> score --> out
- Recon — walk the repo and answer "what's in here" cheaply, without
parsing any source language: languages present (by extension), SDK
dependencies declared in manifests (
pyproject.toml/requirements.txt/Pipfile/poetry.lock/package.jsonfor theclaude-agent-sdk/@anthropic-ai/claude-agent-sdk/openai-agents/@openai/agents/google-adk/@google/adkneedles), the file inventory, and discovered agent components (MCP configs, hook scripts,CLAUDE.mdandAGENTS.mdguidance docs,.claude/agents/*.mdsubagents at any depth,SKILL.mdskills, slash commands at both.claude/commands/*.mdand<plugin-root>/commands/*.md,.claude-plugin/{plugin,marketplace}.jsonmanifests, sandbox policies). No tree-sitter parses happen here — this step decides whether the expensive AST work is even worth attempting. - Inventory — for each language Recon cleared, do the AST work and
extract a typed inventory:
ToolDefs with their config and body facts,AgentDefs with all kwargs captured,SubagentDefs /SkillDefs /SlashCommandDefs /PluginManifests parsed from markdown and JSON frontmatter,MCPServerDefs, guardrails, sessions, and the resolved edges between agents and the tools/guardrails they reference. Detectors read fields off these structs — they never re-parse raw source. - Policy selection — load only the rule packs for SDKs actually
observed in code. An SDK seen in code with no shipped pack emits a
META-001info finding ("Trustabl does not currently audit this SDK") — silence on an unknown SDK is wrong. A dep declared but never used in code emits a different info finding flagging the drift. - Analysis — run the selected scope-aware detectors against the inventory. Findings carry the scope they fired at and attribute to the right location: tool file/line, agent call site, subagent markdown file, or the manifest.
Three properties fall out of this staging, by design:
- Performance. A repo with no Python skips Python AST work; a repo with only Claude TS code skips Python AST work AND OpenAI policy loading.
- Honest coverage. An "unaudited SDK" info finding is louder than a
zero-findings clean bill of health on an SDK Trustabl doesn't know. A
META-004finding further distinguishes "audited and clean" from "could not audit — discovery extracted nothing a rule targets." - Determinism is a contract. Same inputs → same
ScanID, and the report is byte-stable across runs (findings sorted by(RuleID, FilePath, Line), inventory slices sorted deterministically). CI consumers can diff scans without spurious churn.
See ARCHITECTURE.md § 2 for the full diagram with typed inputs at each step.
What's wired today
Tool/agent AST discovery is wired for:
- Python — Claude Agent SDK (decorators), OpenAI Agents SDK, Google
ADK, LangChain / LangGraph, CrewAI, AutoGen / AG2, and Pydantic AI.
Discovery extracts tool definitions, agent constructors, hosted
tools, MCP servers, guardrails, sessions. The bare
Agent(...)constructor shared by OpenAI / ADK / CrewAI / Pydantic AI is import-gated per SDK so the classes never cross-match, and the shared@tooldecorator is routed to the owning SDK by its import binding. - TypeScript — Claude Agent SDK (the
tool()factory, thequery()main-threadQueryMainAgent, inline-in-query()sub-agents, typed-constAgentDefinitions,createSdkMcpServerand the fouroptions.mcpServersconfig literals), OpenAI Agents SDK (thetool({...})factory,new Agent({...})andAgent.create({...}), 9 hosted-tool factories, MCP server classes across 3 transports plus theMCPServerswrapper, 4defineXguardrail factories, and theMemorySession/OpenAIConversationsSession/OpenAIResponsesCompactionSessionsession classes — gated on imports from@openai/agents,@openai/agents-core, or@openai/agents-openai), and Google ADK (thenew FunctionTool({...})constructor, 5 agent constructors —new LlmAgent({...})/SequentialAgent/ParallelAgent/LoopAgent/RoutedAgent— 13 hosted-tool classes, andsubAgentsedges — gated on imports from@google/adk), LangChain / LangGraph (thetool(fn, {...})factory,DynamicStructuredTool/DynamicTool, andcreateReactAgent/createAgent/new AgentExecutor— gated on the@langchain/*/langchain/langgraphecosystem), and the Vercel AI SDK (thetool({...})/dynamicTool({...})single-object factory, the call-basedgenerateText/streamText/generateObject/streamObjectagents and the classToolLoopAgent/Experimental_Agent, withtoolswalked as an object/record, plus the<provider>.tools.*()hosted tools — gated on the bareaiimport). Handles.ts/.tsx/.mts/.ctsplus JavaScript.js/.jsx/.mjs/.cjswith thetree-sitter-typescriptandtree-sitter-tsxgrammars (JavaScript routes to the tsx grammar — a JS superset — and is audited by the samelanguage: typescriptrule packs). TypeScript rule packs ship for the Claude Agent SDK (CSDK-010/011/012/013/014/016 tool rules; CSDK-120/130/131 agent rules), OpenAI Agents SDK (OAI-016/017/019/022/024 tool rules; OAI-105/116 agent rules), Google ADK (ADK-013/015/016 tool rules; ADK-109 agent rule), MCP (MCP-011/012/013/014 tool rules), LangChain (LC-010/011/012/013/014 tool rules; LC-111 agent rule), and the Vercel AI SDK (VAI-001..008 tool/agent rules; VAI-012 repo rule). A TS repo for any of these no longer produces a blanketMETA-004; seeCOVERAGE.mdfor the full matrix.
JavaScript (.js / .jsx / .mjs / .cjs) is AST-parsed through the shared
TypeScript-family pipeline: its tools and agents are discovered, tagged
javascript, and audited by the language: typescript rule packs (both ES
import and CommonJS require() bindings are recognized). Go has
tree-sitter-go discovery for MCP tools (mark3labs/mcp-go and the official
modelcontextprotocol/go-sdk), audited by the language: go rules in the MCP
pack. C# has tree-sitter-c-sharp discovery for the official ModelContextProtocol
SDK's [McpServerTool] methods, audited by the language: csharp rules. PHP has
tree-sitter-php discovery for #[McpTool]-attributed methods (official mcp/sdk
and community php-mcp/server), audited by the language: php rules. Rust has
tree-sitter-rust discovery for the official rmcp crate's #[tool]-attributed
methods (descriptions read from the description = "..." arg or the /// doc
comment), audited by the language: rust rules; other Go, .NET, PHP, and Rust
SDKs are recognized as files by Recon but not yet AST-parsed.
The rule schema's language: field gates per-language rule sets.
Scope boundaries
Shortened here. Read the whole README on GitHub.
Signals
- GitHub stars
- 72
- Forks
- 64
- Last commit
- Oct 2026
Advanced
- Delivery
- agent-reliability-analyzer MCP server → your ahel connector (mcp.ahel.ai) → your AI.
- Item type
- mcp-server
- Key
io-github-trustabl-agent-reliability-analyzer- Source
- github.com/trustabl/agent-reliability-analyzer
github.com/trustabl/agent-reliability-analyzer