SuperCompress v2

MCP serverAI & models

Lets your agent shrink coding-session context by about two thirds to fit more work in the model's memory.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use SuperCompress v2

About this server

Compress coding-agent context ~64%. Hosted Neural Keep MCP for Cursor/Claude/Codex.

Install SuperCompress v2

The server’s own address, for the clients that take one directly. Or connect ahel once and every client you use reads it from one address, with the account kept on ahel rather than in each client’s config.

  • Claude Code

    claude mcp add --transport http --scope user supercompress-v2 'https://api.supercompress.dev/api/mcp'

    Run it once in your project, then open /mcp to approve any sign-in the server asks for.

  • Claude Desktop

    https://api.supercompress.dev/api/mcp

    Add a custom connector in Settings, paste this address, and approve the sign-in.

  • Cursor

    cursor://anysphere.cursor-deeplink/mcp/install?name=supercompress-v2&config=eyJ1cmwiOiJodHRwczovL2FwaS5zdXBlcmNvbXByZXNzLmRldi9hcGkvbWNwIn0=

    Open the link and Cursor adds the server at that address.

  • ChatGPT

    https://api.supercompress.dev/api/mcp

    In Settings, enable Developer mode, create an MCP app, and paste this address. Your plan and workspace must allow custom apps.

  • Codex

    codex mcp add supercompress-v2 --url 'https://api.supercompress.dev/api/mcp'

    Run it once, then sign in with codex mcp login supercompress-v2 if the server asks for an account.

From the project's README

As published by supercompress/supercompress in README.md.


Why it exists

Every LLM call ships a pile of context: RAG chunks, chat history, tool dumps, logs, JSON. Most of it is irrelevant to the current question — but the model still reads it, and you still pay for it.

Usual “fix”What actually happens
TruncateDeletes the middle. The answer is often in the middle.
SummarizeRewrites evidence. IDs, stack traces, and exact errors get soft.
HopeShip the full dump. Watch the bill climb.

SuperCompress v2 compresses context against the query. It keeps answer-critical lines in their original wording and drops the rest — on our coding-agent benchmark (B5), 64.1% mean context reduction with 24/24 evidence-retention passes.


Product

SuperCompress is a compression layer in front of inference:

  1. Takes a long context + the current query
  2. Segments and scores blocks by relevance to that query
  3. Keeps entities, errors, definitions, nearby dependencies
  4. Returns a smaller prompt + token stats

The query is never compressed — only the surrounding context.

Two paths

Neural v2 (recommended)Compiler (fast/local)
Engine~400M query-aware cross-encoderLightweight local policy
RuntimeHosted Fly CPU (Neural Keep)Local CPU · millisecond-class
Best forHighest-quality keep on agent dumpsLocal preprocessing / speed
BenchmarksLaunch / B5Legacy section

Two products

Coding-agent pluginAPI / Python
ForCursor, Claude Code, Codex, and 40+ agent harnessesApps, RAG, agents, backends
Installnpm i -g supercompress-proxy && npx supercompress setuppip install supercompress
What you getMCP compress_context on big dumpsCompress before every model call
LoginKeep your normal agent loginAPI key from the dashboard

Docs: coding agents (first 5 minutes checklist) · API quickstart

Repo map

PathWhat
packages/proxyCoding-agent plugin (npm)
api/Hosted API + billing
web/Site + docs HTML
supercompress/Python package
docs/REPO_LAYOUT.mdWhat belongs in OSS vs private

Private marketing, outreach, and model training stay out of this repo (see .gitignore + docs/REPO_LAYOUT.md).


Benchmarks & stats

We measure whether required evidence survives (containment), not downstream LLM completion.

Neural v2 launch (hosted ~400M) — B5 coding-agent suite

MetricResult
Mean context cut64.1%
Evidence passes24 / 24
Tokens16,647 → 5,148
Max cut @ ≥99% retention96.6%
B5 latency p50~5.8 s (hosted Neural Keep on Fly CPU)
Public cases (full suite)390
Downstream LLM evalNot yet run

Aggregate mean cut across all 390 cases is ~3.6% — the engine often refuses to over-cut dense needle/QA slices. The 64.1% figure is the coding-agent suite where dumps are noisy.

Raw JSON: launch-benchmark.json · writeup: benchmarks

Legacy compiler (separate product)

CPU / millisecond-class local path. Older held-out compiler numbers (≈58–66% cut, 99.4% gold containment) live under Legacy/compiler on /benchmarks. Do not mix with Neural v2.


Try it

Coding agents (recommended)

npm install -g supercompress-proxy
npx supercompress setup

Links your account, detects agents, installs MCP. Docs: coding agents

Python / HTTP

pip install supercompress
export SUPERCOMPRESS_API_KEY=sc_live_YOUR_KEY
from supercompress.client import SuperCompress

sc = SuperCompress()
result = sc.compress(
    context=long_context,
    query="What failed and how do we fix it?",
)
print(f"{result.original_tokens} → {result.kept_tokens} tokens")
print(result.compressed_text)
curl -X POST https://www.supercompress.dev/api/v1/compress \
  -H "X-API-Key: $SUPERCOMPRESS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"context":"...","query":"What failed?"}'

Or paste a dump into the Arena — no integration required.


How it compares

TruncateSummarizeSuperCompress
Cuts tokens✓✓✓
Uses the query✗weak✓
Keeps original evidencesometimes✗✓
Auditable kept linespartial✗✓

More: vs truncation · vs summarization · vs alternatives


Signals

GitHub stars
103
Forks
22
Last commit
Oct 2026
Weekly downloads
1k
Advanced
Delivery
supercompress MCP server → your ahel connector (mcp.ahel.ai) → your AI.
Item type
mcp-server
Key
io-github-supercompress-supercompress
Source
github.com/supercompress/supercompress
Hosted endpoint
https://api.supercompress.dev/api/mcp