tokcalc

MCP serverAI & models

LLM capacity planning tools to estimate VRAM, compute, latency, and GPU topology for agents.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use tokcalc

Install tokcalc

The server’s own address, for the clients that take one directly. Or connect ahel onceand every client you use reads it from one address, with the account kept on ahel rather than in each client’s config.

  • Claude Code

    claude mcp add --transport http --scope user tokcalc 'https://mcp.tokcalc.app/mcp'

    Run it once in your project, then open /mcp to approve any sign-in the server asks for.

  • Claude Desktop

    https://mcp.tokcalc.app/mcp

    Add a custom connector in Settings, paste this address, and approve the sign-in.

  • Cursor

    cursor://anysphere.cursor-deeplink/mcp/install?name=tokcalc&config=eyJ1cmwiOiJodHRwczovL21jcC50b2tjYWxjLmFwcC9tY3AifQ==

    Open the link and Cursor adds the server at that address.

  • ChatGPT

    https://mcp.tokcalc.app/mcp

    In Settings, enable Developer mode, create an MCP app, and paste this address. Your plan and workspace must allow custom apps.

  • Codex

    codex mcp add tokcalc --url 'https://mcp.tokcalc.app/mcp'

    Run it once, then sign in with codex mcp login tokcalc if the server asks for an account.

From the project's README

As published by stevecrates489-commits/tokcalc in README.md.

The open-source LLM serving capacity planner

Plan your LLM deployment before you rent the GPUs.


Switch model → multi-GPU → long-context capacity planner → Build vs Buy → Reference catalog


tokcalc turns your LLM traffic, context length, latency SLOs, cache behavior, and model choice into a defensible serving topology and cost plan — with transparent formulas and cited benchmarks.

It's the engineering-grade pre-deployment decision layer for LLM inference. Not another static "tokens per second" calculator.

Try it now

tokcalc.dev — no signup, no tracking, no paywall.

Pick a model, a GPU, and a workload. Get an instant capacity plan:

  • Generation speed (tok/s)
  • Time-to-first-token (TTFT) + inter-token latency (ITL)
  • VRAM budget with KV-cache sizing
  • Multi-GPU topology recommendation (Single GPU → TP×2/4/8 → Context Parallel)
  • Monthly cost + break-even vs API pricing
  • Shareable URL — your config encoded in the URL hash, send to colleagues

Why tokcalc?

The market has dozens of "tokens per second" calculators and self-host-vs-API break-even tools (induwara.lk, gigagpu, kickllm, cloudparity, curlscape, profitable.ai). None of them are unified capacity planners.

What tokcalc answers that competitors can't

"Can I serve Qwen 2.5 72B at 128K context on 2× H100 with 20 concurrent users?"

"How many H200s do I need for 1,000 req/min with P95 TTFT < 2s?"

"Does FP8 or AWQ save more money once quality, KV cache, and engine support are included?"

"At what daily volume does an H100 beat GPT-4o pricing?"

"What happens to cost and latency if an agent makes 8 model calls, has 3 tool calls, and its context grows by 5K tokens each turn?"

"Would prefix caching, continuous batching, or PD disaggregation save more for this workload?"

Features

Calculator tab

FeatureWhat it computes
Model fit / VRAMWill the model + KV cache fit in the GPU's memory?
ThroughputDecode tok/s (per-stream) + aggregate (batched) + prefill tok/s
Latency splitTime-to-first-token (= prefill) + inter-token latency (= decode)
Continuous batchingUser-tunable 1.0–4× multiplier (cited 1.5–4× SOSP range)
Reasoning tokensHidden reasoning budget added to billed output (o1 / R1 / Claude thinking)
Prompt cachingSelf-hosted vLLM APC + Anthropic 5m/1h TTL + OpenAI 50%-off cached tokens
Speculative decodingUser-tunable 1.2–4× boost factor
Multi-GPU TP1× → 8× tensor parallel with NVLink efficiency factor
Long-context capacityKV memory + max concurrency + prefill time at 4K → 1M context
Topology recommendationSingle GPU → TP×2 → TP×4 → TP×8 → TP×8 + Context Parallel (RingAttention)
Cost economicsGPU $/hr → $/M output tokens → $/request → monthly cost

Build vs Buy tab

Independent calculator (separate state) that compares:

  • Self-host: model + GPU + quant + utilization + batch → $/M tokens + monthly cost
  • API: 13 providers (OpenAI / Anthropic / Gemini / Groq / DeepSeek / Mistral / Together)
  • Verdict: Self-host cheaper / API cheaper / Not enough volume — with break-even reqs/day

Reference tab

5 sub-tables — fully transparent, every record source-linked where available:

  • Models (35 entries: Llama 4 Scout/Maverick, Qwen 3 family, DeepSeek V3/R1, Pixtral, BGE-M3, ...)
  • GPUs (30 entries: H100/H200/B200/B300, AMD MI300X/MI325X, Intel Gaudi 3, TPU v5p/Trillium, Groq LPU, Cerebras WSE-3, Apple M2/M3/M4 Ultra, ...)
  • Quantization (16 formats: FP16/BF16, GGUF Q2_K→Q8_0, GPTQ, AWQ, EXL2, FP8, NVFP4)
  • API pricing (13 models with input/cached/output + retired/current status)
  • Cloud GPU pricing (all GPUs with $/hr > 0 + typical providers)

The long-context capacity planner (the differentiator)

This is the formula the Perplexity research brief called "the most important tokcalc should visibly expose":

$$ \text{KV bytes/request} = 2 \cdot L \cdot T \cdot H_{\text{kv}} \cdot D_h \cdot B $$

For dense attention, prefill cost grows superlinearly with context length:

$$ \text{prefill FLOPs} = \underbrace{2 \cdot N \cdot T}{\text{linear}} + \underbrace{T^2 \cdot H{\text{kv}} \cdot D_h \cdot L}_{\text{attention}} $$

tokcalc shows you:

  • Max concurrent users at 4K / 8K / 16K / 32K / 64K / 128K / 256K / 512K / 1M context
  • KV memory per request at each context length
  • Prefill time (with superlinear attention correction beyond 32K)
  • Required topology (Single GPU → TP×2/4/8 → TP×8 + Context Parallel)
  • RingAttention citation when CP is needed

The math, transparently

Every number above comes from a formula you can inspect. No black boxes.

$$ \text{decode tok/sec} \approx \frac{\text{HBM BW} \cdot \eta_{\text{mem}} \cdot \text{quant_eff}}{\text{model size}} $$

Where:

  • HBM BW = GPU memory bandwidth (e.g., 3350 GB/s for H100 SXM)
  • η_mem = 0.65 = typical real-world memory utilization (35% overhead)
  • quant_eff = dequantization efficiency multiplier (1.0 for FP16, 1.5 for FP8 on H100, 0.85 for INT4)
  • model size = active_params × bytes_per_param (uses ACTIVE params for MoE, not total)

Refs: PagedAttention paper (arxiv.org/abs/2309.06180)

$$ \text{prefill tok/sec} \approx \frac{\text{GPU FLOPS} \cdot \eta_{\text{compute}}}{2 \cdot \text{active params}} $$

Where:

  • GPU FLOPS = dense FP16/BF16 TFLOPS (sparse values not used)
  • η_compute = 0.50 = typical compute utilization
  • Factor of 2 = one multiply + one add per parameter per token

For long context (>32K), the superlinear attention correction above applies.

$$ \text{aggregate tok/sec} = \text{decode tok/sec} \cdot \text{batch size} \cdot \text{continuous batching multiplier} $$

Critical caveat: There is no universal continuous batching multiplier. vLLM reported 14–24× vs HF Transformers (extreme), 2.2–2.5× vs TGI. SOSP paper finds 2–4× typical vs FasterTransformer/Orca. tokcalc defaults to a conservative 1.5× and lets you tune.

Refs:

For shared prefix of length $T_p$, suffix of length $T_u$, output $O$, cache hit rate $h$:

$$ \text{API input cost} = N \cdot \left[ (1-h) \cdot T_p \cdot P_{\text{write}} + h \cdot T_p \cdot P_{\text{read}} + T_u \cdot P_{\text{input}} \right] $$

Anthropic multipliers (verified 2025-2026):

  • 5-minute cache write: 1.25× base input
  • 1-hour cache write: 2.0× base input
  • Cache read: 0.1× base input (90% savings)

OpenAI: cached input discounted 50%, no separate write fee.

Refs:

$$ \text{total needed} = \text{model weights} + (\text{KV per request} \cdot \text{batch size}) $$

Walk the smallest topology that fits:

  • Single GPU: total ≤ VRAM × 1
  • Tensor Parallel ×2/4/8: total ≤ VRAM × N (weights + KV split evenly)
  • TP×8 + Context Parallel: total > VRAM × 8 — use RingAttention to shard KV across nodes

Refs: RingAttention paper

$$ \text{self-host $/M tokens} = \frac{\text{GPU $/hr}}{3600 \cdot \text{effective tok/s} \cdot \text{utilization}} \cdot 10^6 $$

$$ \text{break-even req/day} = \frac{\text{monthly self-host cost}}{30 \cdot \text{API cost per request}} $$

The decisive term is effective utilization — not peak throughput. A GPU running at 10% utilization pays 10× more per token than the theoretical minimum.

Comparison with adjacent tools

Capabilityinduwara / techfuelhq / pcmasterstudiogigagpu / kickllm / cloudparityHF Open LLM Leaderboard / MLPerftokcalc
Model × GPU × quant tok/s✓somesome✓
Model-fit / VRAMbasicrarerare✓
KV-cache by context + concurrency——implicit✓
Prefill vs decode split (TTFT/ITL)——engine-specific✓
Continuous batching / paged attention——docs only✓
Long-context (128K–1M) planning———✓
Topology recommendation (TP/CP)——partial✓
Prompt-cache economics—partial API only—✓
Reasoning tokens (o1/R1/Claude thinking)———✓
API vs self-host break-evensome✓—✓
Transparent formulas / open sourcemixedusually nomixed✓
Cited benchmark evidence per configrarerare✓ (not planning)✓ (in progress)
Shareable URL per config———✓

Roadmap

Shipped

  • ✅ 35 models, 30 GPUs, 16 quantization formats
  • ✅ Continuous batching, reasoning tokens, prompt caching
  • ✅ TTFT/ITL split, long-context superlinear attention
  • ✅ Long-context capacity planner + topology recommendation
  • ✅ Build-vs-Buy calculator (13 API providers with retired/current status)
  • ✅ Reference catalog (5 sub-tables)
  • ✅ Share URL + localStorage persistence
  • ✅ Dark mode toggle
  • ✅ Plain-English glossary (28 terms with hover tooltips)
  • ✅ OG image + Twitter card + social metadata

Next 30 days

  • ⏳ GitHub Action (tokcalc/plan PR comment)
  • ⏳ MCP server (read-only capacity-planning tools for AI agents)
  • ⏳ 3 SEO landing pages (/compare/h100-vs-h200, /gguf-q4-k-m-vs-q5-k-m, /vllm-vs-sglang)
  • ⏳ i18n: Chinese, Japanese, Korean

Next 90 days

  • ⏳ Workload-trace / SLO capacity planner (prompt/output/arrival distributions, p50/p95 TTFT/ITL)
  • ⏳ P/D disaggregation planner (separate prefill + decode pools)
  • ⏳ Cache-aware economics (prefix-sharing distribution, multi-turn/agent traces)
  • ⏳ Engine-aware presets (vLLM / SGLang / TensorRT-LLM / llama.cpp)
  • ⏳ Versioned price + benchmark provenance system

Long-term

  • 🔮 Agentic workflow calculator (multi-turn + tool calls + growing context)
  • 🔮 Multi-LoRA capacity planner (Punica / S-LoRA economics)
  • 🔮 VLM image-token accounting (per-model patch/tile tokenization)
  • 🔮 Embedding model mode (vectors/sec, separate workload)
  • 🔮 Training/fine-tuning estimator (LoRA / QLoRA / full-SFT FLOPs)
  • 🔮 Energy / carbon per million tokens (region-specific grid intensity)

Open core model

tokcalc is open core — the calculator and catalog are open source; the cloud/data/team features are paid.

AssetLicenseNotes
Source codeApache 2.0This repo. Free to use, modify, distribute
Model/GPU/quant catalogCC0 1.0Public domain data. Anyone can use, no attribution required
Benchmark provenance dataCC-BY-SA 4.0Anyone can use, but must attribute + share-alike
DocumentationCC-BY 4.0Attribution required if copied
"tokcalc" name + logoTrademarkEven without formal registration, common-law rights apply
Cloud SaaS layerProprietaryReal-time pricing API, benchmark DB, team workspaces (coming soon)

Why this structure

  • Trust: Open-source formulas build credibility vs opaque competitors
  • Community: Contributors can submit models, GPUs, quants, benchmarks
  • Defensibility: Trademark + cloud features + URL-share viral loop protect against forks
  • Revenue: Cloud tier funds ongoing development + pricing/benchmark data maintenance

Contributing

We welcome contributions! See CONTRIBUTING.md for:

  • How to add a model (with HuggingFace config.json as source)
  • How to add a GPU (with critical guardrails for B200/B300 null FP16 fields)
  • How to add a quantization format (with measured file sizes as source)
  • How to submit a benchmark (3-tier confidence model)
  • How to improve a formula (cite the source, no magic numbers)

Most-needed contributions

  • 🟢 New models (Qwen 3, Mistral Large 4, Llama 4 variants as they release)
  • 🟢 New GPUs (B300, AMD MI400, Apple M5 Ultra when shipping)
  • 🟢 New quantization formats (mxFP8, MXFP4, BitNet 2)
  • 🟢 Real benchmark data (run vLLM benchmarks and submit with provenance)
  • 🟢 Translations (especially Chinese, Japanese, Korean)

MCP server — use tokcalc from AI agents

tokcalc ships an MCP (Model Context Protocol) server that lets AI agents (Cursor, Claude Desktop, Cline) call tokcalc during design reviews.

6 read-only tools

ToolWhat it does
estimate_capacityVRAM/KV/throughput/latency/cost for one config
compare_gpusRanked GPU comparison for one workload
recommend_topologyTP/CP topology recommendation
estimate_api_vs_self_hostBreak-even analysis
list_modelsDiscover supported model IDs
list_gpusDiscover supported GPU IDs

All tools are read-only — no side effects, no cloud credentials, no deployments.

Install

Add to your Claude Desktop config (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):

{
  "mcpServers": {
    "tokcalc": {
      "command": "npx",
      "args": ["-y", "@tokcalc/mcp-server"]
    }
  }
}

Or run locally:

git clone https://github.com/stevecrates489-commits/tokcalc.git
cd tokcalc
bun install
bun mini-services/mcp-server/index.ts

Example agent prompt

"I need to serve Llama 3.3 70B at 32K context for 50 concurrent users. What GPU topology do you recommend, and how much will it cost per month?"

The agent calls list_models → list_gpus → recommend_topology → estimate_capacity → returns a structured plan with VRAM, throughput, latency, cost, and confidence.

Tech stack

  • Framework: Next.js 16 with App Router
  • Language: TypeScript 5
  • Styling: Tailwind CSS 4 + shadcn/ui (New York)
  • Charts: Recharts
  • State: React hooks (useState + useEffect + useMemo)
  • Theme: next-themes (dark mode default)
  • Database: None (pure client-side, no backend required for core features)

Acknowledgments

tokcalc builds on the work of:

  • vLLM team — PagedAttention, continuous batching (arxiv.org/abs/2309.06180)
  • llama.cpp / ggml-org — GGUF format and quantization variants
  • MLPerf / MLCommons — standardized inference benchmark methodology
  • Stanford HELM — efficiency-aware model evaluation framework
  • Anthropic / OpenAI / Google — published prompt-caching pricing rules
  • NVIDIA / AMD / Intel / Google / Groq / Cerebras — published hardware specs

Every formula has a citation. Every model/GPU/quant entry has a source URL where available. If you spot an unsourced claim, please open an issue.

Star history


Live demo · Documentation · Contributing · License · Code of Conduct

Made with care by the tokcalc community. Apache 2.0 licensed.

Advanced
Delivery
tokcalc MCP server → your ahel connector (mcp.ahel.ai) → your AI.
Catalog kind
mcp-server
Key
io-github-stevecrates489-commits-tokcalc
Source
github.com/stevecrates489-commits/tokcalc.git
Hosted endpoint
https://mcp.tokcalc.app/mcp