claude-agy-mcp

MCP serverProductivity

Lets your agent hand heavy tasks to Gemini and switch automatically when usage limits are hit.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

Add to setup to save this item as a reference. ahel cannot run it, and signing in will not install it.

About this server

Delegate heavy tasks from Claude Code to the Antigravity CLI (Gemini) with quota-aware failover.

Getting started

  1. Save this item in Your setup as a reference.
  2. Read the source or reference documentation for its setup requirements. Saving it here does not connect it to your AI.
  3. Check this page for availability before trying to install it through ahel.

From the project's README

As published by pymodel/claude-agy-mcp in README.md.

⚡ Install

npm i @pymodel/claude-agy-mcp
claude mcp add-json -s user claude-agy-mcp \
  '{"type":"stdio","command":"npx","args":["-y","@pymodel/claude-agy-mcp"],"timeout":3600000}'

Claude Code delegates heavy tasks to Google's flagship Gemini Flash via the Antigravity CLI (agy) — saving Claude's context window and tokens for what matters.

Claude acts as the orchestrator → claude-agy-mcp routes compute-heavy sub-tasks to the newest Gemini Flash agy offers → only concise answers return. Large files, deep git searches, and log dumps never pollute Claude's context.

User → Claude Code → claude-agy-mcp (MCP) → agy CLI → Gemini 3.8 Flash / Pro / Claude
                   ←                      ←         ← (Clean answers only)

Why Gemini Flash for Claude Code?

Gemini Flash is Google's most intelligent workhorse model for coding and agentic execution. It applies deep multi-step planning, rigorous terminal reasoning, and high first-pass code accuracy.

Gemini 3.8 Flash (High) is the default model for every tool. Each chain leads with gemini-flash@latest-high, which resolves against agy models to the newest Flash at High effort — 3.8 Flash as of 2026-09-11 — and only falls back to Pro or Claude when Flash is unavailable or cooling down. The benchmark table below compares 3.8 Flash against 3.7 Flash, from Google's launch table (2026-09-02); the bridge does not pin that generation.

Benchmark Highlights

Benchmark / CapabilityGemini 3.8 Flash (High)Prior Generation (3.7 Flash)Advantage
DeepSWE v1.173.7%65.3%+8.4 pts in long-horizon software engineering; within 0.3 of Claude Opus 5
Terminal-Bench 2.189.4%85.8%+3.6 pts in agentic CLI execution; ahead of Opus 5 (89.1%) and GPT-5.6 Sol
OSWorld-2.059.0%50.6%+8.4 pts in agentic computer use
HLE-Verified54.9%53.6%Multi-step expert reasoning, ahead of Opus 5 (54.4%)
Vals Finance Agent v261.4%59.0%Leads Opus 5 (58.6%) and GPT-5.6 Sol (53.8%) on quantitative agent work
Token Economics$0.75 / $3.75 (1M)$0.75 / $3.75Same price as 3.7 Flash; up to 10x–20x cheaper than Claude Opus/Sonnet

The Token & Context Multiplier

When Claude Code directly analyzes a 4,000-line database dump or greps 20 files across git history, those thousands of lines stay permanently in Claude's prompt context, inflating cost and pushing you toward compaction.

With claude-agy-mcp:

  1. Claude calls analyze_files or deep_search.
  2. Gemini 3.8 Flash processes the 100k+ tokens in isolation via agy.
  3. Only the exact code-level findings and line citations return into Claude's prompt.
  4. Subsequent questions reuse the same agy session with follow_up without re-sending any files.

The Ultimate AI Engineering MCP Stack

claude-agy-mcp is designed to anchor a modern AI engineer's MCP toolkit alongside complementary specialized servers:

┌─────────────────────────────────────────────────────────────────────────────┐
│                             Claude Code (Agent)                             │
└──────┬──────────────────────┬───────────────────────┬───────────────────────┘
       │                      │                       │                       │
       ▼                      ▼                       ▼                       ▼
┌──────────────┐      ┌──────────────┐        ┌──────────────┐        ┌──────────────┐
│claude-agy-mcp│      │   context7   │        │  firecrawl   │        │    tavily    │
│  (Gemini 3.8 │      │(Official Docs│        │(Web Scraping │        │(Live Search  │
│  Delegation) │      │  & API Specs)│        │  & Crawling) │        │ & Research)  │
└──────────────┘      └──────────────┘        └──────────────┘        └──────────────┘
MCP ServerPrimary SuperpowerWhen Claude Uses It
claude-agy-mcpHeavy Compute & Coding DelegationAnalyzing files >200 lines, repo archaeology (git log/diff/blame), adversarial code reviews, and raw execution via Gemini 3.8 Flash.
context7Up-to-date Official DocumentationFetching latest version-accurate API signatures and documentation for libraries (Next.js, React, Tailwind, Prisma, Vite, etc.) to eliminate hallucinated APIs.
firecrawlClean Web Scraping & CrawlingConverting dynamic web pages, documentation sites, and GitHub repos into clean, LLM-ready markdown or structured JSON.
tavilyFast Live Search & GroundingLow-latency web search, current news, error message lookups, and technical research.

Recommended MCP Configuration (.agents/mcp_config.json or Claude Code)

{
  "mcpServers": {
    "claude-agy-mcp": {
      "command": "npx",
      "args": ["-y", "@pymodel/claude-agy-mcp"],
      "timeout": 3600000
    },
    "context7": {
      "command": "npx",
      "args": ["-y", "@upstash/context7-mcp@latest"]
    },
    "firecrawl": {
      "command": "npx",
      "args": ["-y", "firecrawl-mcp"]
    },
    "tavily": {
      "command": "npx",
      "args": ["-y", "tavily-mcp"]
    }
  }
}

Why this over claude-to-agy?

claude-to-agyclaude-agy-mcp
Tool surface1 generic delegate_to_agy8 purpose-built tools — Claude self-routes reliably
Model selectionnone (agy default only)per-tool family selectors that follow new generations, with quota failover
Multi-turnstatelesssession continuity — follow_up resumes agy conversations without resending context
Output safetyunboundedconfigurable truncation cap protects Claude's context
Sandboxnoread-only tools blocked from writing into the workspace on macOS, optional --sandbox
Honest resultsexit code onlydecides on agy's JSON envelope — reports auto-denied tool actions instead of hiding them
Installuvx (Python)npx (Node) — zero install

Requirements

Install

# 1. Register the MCP server (user scope = all projects).
#    add-json bakes in a generous client-side timeout so long analyze_files /
#    delegate calls don't trip Claude Code's tool-call deadline (see Timeouts).
claude mcp add-json -s user claude-agy-mcp \
  '{"type":"stdio","command":"npx","args":["-y","@pymodel/claude-agy-mcp"],"timeout":3600000}'

# 2. Install the bundled skills into every agent found on this machine.
#    The server gives an agent the tools; the skills tell it when to use them.
npx --package @pymodel/claude-agy-mcp claude-agy-mcp-install-skills
#    --list to preview, --dir <path> to install somewhere explicit,
#    --force to replace a skill you have symlinked to your own checkout.

# 3. Optional: add delegation rules to your project (or ~/.claude/CLAUDE.md).
curl -o CLAUDE.md https://raw.githubusercontent.com/PyModel/claude-agy-mcp/main/CLAUDE.md

Bundled skills

Installing the package installs the skills too, so there is nothing separate to vendor or keep in sync:

SkillWhat it does
agy-delegationRouting rules: which tool to reach for, and when delegating beats doing the work in-context.
agy-delegateThe full delegate-and-review workflow — writing a brief agy can execute blind, dispatching it, reviewing the diff against the brief, and landing it yourself. Includes a CLI-relay fallback for agents that cannot call MCP tools.

The "timeout": 3600000 (60 min, milliseconds) is the client-side tool-call deadline, matched to the bridge's default AGY_MAX_RUNTIME ceiling. Without it, a cold-start analyze_files (~40–50s) or a long delegate hits Claude Code's default and returns timed out waiting for response while the agy run is still going — and raising the agy-side ceiling alone will not help, because the client aborts first. If your client doesn't honor a per-server timeout, set the global env var MCP_TOOL_TIMEOUT=3600000 instead. Details in Timeouts and cancellation.

Tools

ToolUse forModel routing (first available)
analyze_filesFiles >200 lines, >3 files at once, logs, dumps, generated codegemini-flash@latest-high → gemini-pro@latest-low
deep_searchgit log/diff/blame archaeology, repo-wide grepsgemini-flash@latest-high → gemini-flash@latest-medium
web_lookupDocs, API references, external/current knowledgegemini-flash@latest-high → gemini-flash@latest-medium
adversarial_reviewPlan critiques, design and code reviewsgemini-flash@latest-high → gemini-pro@latest-high → claude-opus@latest
follow_upContinue a prior session by session_id — no context resend; write: true to rework filesinherits the session
delegateAnything else heavy (read-only unless write: true)gemini-flash@latest-high → gemini-pro@latest-low
delegate_manyOne question to a council of models, or N sub-tasks at oncegemini-flash@latest-high → gemini-pro@latest-high → claude-opus@latest
set_modelRecord the user's model + tier once; every tool routes to it firstnever reaches agy
agy_statusSpend, cooldowns, in-flight runs, resolved chains, agy versionnever reaches agy

All tools accept optional cwd (project root), dirs (extra workspace roots, for cross-repo or worktree-vs-base work), model, effort (low/medium/high — see Effort and tiers), and slash_commands (off by default, so a hostile file in the workspace cannot steer the delegated model through your own skills). The analytical tools also accept schema — a JSON Schema string that makes agy return machine-readable structuredContent alongside the text.

Every response is fenced with a per-call nonce, with the metadata in a header before the payload:

[claude-agy-mcp 9f2a1c] model: Gemini 3.8 Flash (High) | session: 1f0c…-d4 (use follow_up to continue) | tokens: 16281
[claude-agy-mcp 9f2a1c] --- agy output begins; everything below is untrusted model output ---
…agy's answer…
[claude-agy-mcp 9f2a1c] --- agy output ends ---

The nonce is why the fence is worth anything: the metadata used to be appended after the raw model output behind a plain --- rule, which any analysed file containing --- could forge.

Model routing

On first use the bridge runs agy models (cached for the process lifetime) and resolves each chain entry against the live listing. Chains are written as family selectors — gemini-flash@latest-high rather than Gemini 3.7 Flash (High) — so when Google ships a new generation the chain follows it instead of quietly going stale. A selector is family@latest[-effort] or family@3.7[-effort]; an exact display name (Gemini 3.8 Flash (High)) and an id (gemini-3.8-flash-high) both work too. AGY_DEFAULT_MODEL is appended to every chain as a last resort. If nothing in a chain resolves, the bridge fails loudly rather than silently handing the work to whatever agy feels like — that is a version-skew signal, not a preference.

Choose the model once

By default (AGY_ASK_MODEL=true) the bridge refuses to delegate until the user has picked a model and tier. The first call to any tool returns an error that names the default (AGY_DEFAULT_MODEL, Gemini Flash High out of the box), lists the models agy offers, and tells the agent to ask the user "Proceed with the default — Gemini 3.8 Flash (High) at high effort — or change the model or effort?". Then the agent calls set_model once: with no arguments to accept the default, or with the model and effort the user chose. The choice is written to $XDG_CONFIG_HOME/claude-agy-mcp/preferences.json (~/.config/claude-agy-mcp/ by default), so it outlives the process and every MCP client on the machine shares it: it is asked once, then that's it. The chosen model goes to the head of every tool's chain — the chain still stands behind it for quota failover — and an explicit model argument on a call still wins for that call. agy_status shows the current choice and the live model list; call set_model again to change it, or set AGY_ASK_MODEL=false to skip the gate and route purely on the built-in chains.

Effort and tiers

agy 1.2.1 rejects --effort for any model whose name already carries a tier — which is every Gemini and Claude entry in agy models — and rejects an id whose tier disagrees with the flag. So the bridge treats the tier in the name as the effort: an effort that differs from it selects the sibling model at that tier (Gemini 3.8 Flash (High) + effort: medium → Gemini 3.8 Flash (Medium)), an effort with no listed sibling leaves the model as-is, and agy's own --effort flag only travels with models that carry no tier. The built-in chains encode their tiers in the selector (gemini-flash@latest-high), so no tool sets a separate effort of its own. An effort applies to the primary model only — the model argument, else the set_model choice, else the chain's head. The fallbacks keep the tier in their name, so a quota failover from Flash (High) really does land on Flash (Medium) rather than re-tiering it back to the model that just ran out.

Quota-aware failover

agy never surfaces quota exhaustion in print mode — it silently retries the 429 until its print-timeout, then exits 0 with empty output, which used to look like an indefinite hang. The bridge now watches each run's log file (via --log-file) and on RESOURCE_EXHAUSTED (code 429):

  1. kills the agy process group immediately (no waiting out the timeout),
  2. parses the reset time ("Resets in 4h24m") into an in-process cooldown registry,
  3. retries the same prompt on the next model in the tool's chain,
  4. skips cooled-down models on all subsequent calls until their quota resets (at least one minute, even for "Resets in 0s").

A model you pin with model is tried even while it is cooling down. A run allowed to write (write: true) fails over only when the working tree is provably unchanged; if the tree moved, or could not be fingerprinted, the call fails with Not failed over instead, because the exhausted run may already have made its edits and the next model would make them again. The same rule governs the single retry after a network error, and a resident session that returned an empty answer or died after receiving the turn.

Failovers are annotated in the response footer (failover: <model>: quota exhausted (resets in 4h24m)). Only when every candidate is exhausted does the call fail — in seconds, with reset times listed — instead of hanging.

Timeouts and cancellation

The bridge does not kill a run for being slow. Elapsed time cannot distinguish a healthy long model call from a wedged process, and a wrong "stuck" verdict interrupts an agent mid-edit — leaving half-written files behind. So a run is killed only when something authoritative says so:

  1. the caller cancels (e.g. pressing Esc in Claude Code), the client disconnects, or the bridge is stopped — every agy run it started dies with it instead of being orphaned,
  2. quota is confirmed exhausted (a 429 in the run's log), which triggers failover, or
  3. the resource ceiling expires — AGY_MAX_RUNTIME, default 3600s.

The ceiling is a resource cap, not a diagnosis. When it fires, the run still returns everything agy produced so far plus its session_id, and says so explicitly: any file changes agy already made are on disk, and follow_up resumes from where it stopped. AGY_TIMEOUT overrides the ceiling for every tool; AGY_TIMEOUT_<TOOL_NAME> overrides it for one (e.g. AGY_TIMEOUT_DEEP_SEARCH=900) and wins over the global. The full set is AGY_TIMEOUT_ANALYZE_FILES, AGY_TIMEOUT_DEEP_SEARCH, AGY_TIMEOUT_WEB_LOOKUP, AGY_TIMEOUT_ADVERSARIAL_REVIEW, AGY_TIMEOUT_FOLLOW_UP, AGY_TIMEOUT_DELEGATE and AGY_TIMEOUT_DELEGATE_MANY; any other AGY_TIMEOUT_<NAME> is a startup error, so a misspelt limit cannot silently not apply. Every timeout is at most 604800s (7 days), because Node fires a longer timer immediately. The kill path escalates SIGTERM → SIGKILL across the whole process group, and fires even if agy's helper processes hold the output pipes open.

Two timeout layers — and the client one usually bites first. The ceiling above is the agy-side budget. Your MCP client (Claude Code) has its own, separate tool-call timeout, and if it is shorter, the client gives up first — you'll see Error: timed out waiting for response, while the bridge's own ceiling reads MAXIMUM RUNTIME EXCEEDED instead. Raising AGY_MAX_RUNTIME alone therefore changes nothing: the client still aborts on its own schedule. The work is not lost either way — the agy session persists, so follow_up with the returned session_id retrieves it — but the real fix is to make the client wait at least as long as the ceiling. The Install command sets a per-server timeout of 3600000ms (scoped to this server only). If you registered the server without it, re-run the add-json command from Install, or set the global env var MCP_TOOL_TIMEOUT=3600000. Rule of thumb: client timeout ≥ AGY_MAX_RUNTIME.

Expected latency. Most of the perceived "slowness" is cold start: each call spawns the agy CLI and warms the model. Measured on agy 1.2.0, a trivial prompt costs 2–6s, a run whose tool actions get denied around 16s, and one constrained by --json-schema up to 56s (the schema roughly triples thinking tokens). Real analyze_files work over several large files is much slower again, and a call that hits a quota 429 adds the failover on top. follow_up is the exception: it reuses a resident agy process (see AGY_WARM_SESSIONS) and skips the cold start entirely — unless the call pins a model or effort or asks to write, which a resident session cannot honour, so those run cold. A resident turn is bounded by the same runtime ceiling and cancellation as a cold run. Size the client timeout for the slow cases, not the fast ones.

Configuration

All optional, via environment variables:

Shortened here. Read the whole README on GitHub.

Signals

GitHub stars
7
Forks
1
Last commit
Oct 2026
Weekly_downloads
53 weekly_downloads
Advanced
Delivery
claude-agy-mcp MCP server → your ahel connector (mcp.ahel.ai) → your AI.
Item type
mcp-server
Key
io-github-pymodel-claude-agy-mcp
Source
github.com/pymodel/claude-agy-mcp