Cost and Budget Enforcement

SkillAI & models

Set per-request, per-session, daily, and monthly spend limits, configure rate limiting and circuit breakers, and isolate costs per user or tenant.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Cost and Budget Enforcement skill

What this skill tells your AI

The instructions your AI receives, as published by tylerjrbuell/reactive-agents-ts in apps/docs/skills/cost-budget-enforcement/SKILL.md and read by ahel’s review.

Agent objective

Produce a builder with cost tracking, budget limits, and rate limiting configured so the agent never exceeds defined spending thresholds.

When to load this skill

  • Deploying agents in production with real API costs
  • Building multi-tenant SaaS where per-user cost isolation matters
  • Protecting against runaway agent loops consuming excessive tokens
  • Adding circuit breakers for provider reliability

Implementation baseline

import { ReactiveAgents } from "@reactive-agents/runtime";

const agent = await ReactiveAgents.create()
  .withName("assistant")
  .withProvider("anthropic")
  .withReasoning({ defaultStrategy: "adaptive", maxIterations: 15 })
  .withTools({ allowedTools: ["web-search", "http-get", "checkpoint"] })
  .withCostTracking({
    perRequest: 0.25,   // max $0.25 per LLM call
    perSession: 2.0,    // max $2.00 per agent.run() call
    daily: 10.0,        // max $10.00/day across all sessions
    monthly: 100.0,     // max $100.00/month
  })
  .withRateLimiting({
    requestsPerMinute: 30,
    tokensPerMinute: 50_000,
    maxConcurrent: 3,
  })
  .withCircuitBreaker()   // auto-opens on provider errors; prevents cascading failures
  .build();

Key patterns

withCostTracking() — budget limits

.withCostTracking()
// Enables cost tracking with defaults:
// perRequest: $1.00, perSession: $5.00, daily: $25.00, monthly: $200.00

.withCostTracking({
  perRequest: 0.50,    // hard stop mid-request if cost would exceed this
  perSession: 5.0,
  daily: 25.0,         // daily limit (default $25.00)
  monthly: 200.0,
})

When a budget is exceeded, the agent throws a BudgetExceededError and stops. Daily/monthly budgets reset based on the timezone configured in .withGateway() (if used) or UTC by default.

withRateLimiting() — throughput caps

.withRateLimiting()
// Defaults: 60 RPM, 100,000 TPM, 10 concurrent requests

.withRateLimiting({
  requestsPerMinute: 60,     // max LLM requests per minute
  tokensPerMinute: 100_000,  // max tokens per minute (input + output)
  maxConcurrent: 10,         // max simultaneous in-flight LLM requests
})

Requests that exceed limits are queued (not dropped) — the agent waits for capacity before proceeding.

withCircuitBreaker() — provider reliability

.withCircuitBreaker()
// Default thresholds (open after 5 failures in 60s window, retry after 30s)

.withCircuitBreaker({
  failureThreshold: 5,       // open circuit after N consecutive failures
  windowMs: 60_000,          // failure counting window
  retryAfterMs: 30_000,      // wait before trying half-open probe
})

Circuit breaker states: closed (normal) → open (failing fast) → half-open (probing recovery).

Per-user cost isolation (multi-tenant)

// Create one agent per user/tenant with separate tracking contexts
const userAgent = await ReactiveAgents.create()
  .withProvider("anthropic")
  .withCostTracking({ perSession: 1.0, daily: 5.0 })
  .withName(`user-${userId}`)
  .withSystemPrompt(`You are assisting user ${userId}.`)
  .build();

// Or use per-request context injection:
const result = await agent.run(task, {
  context: { userId, tenantId },   // included in cost tracking metadata
});

Dynamic pricing (LiteLLM / custom providers)

import { createLiteLLMPricingProvider } from "@reactive-agents/llm-provider";

.withDynamicPricing(createLiteLLMPricingProvider())
// Fetches live model prices from LiteLLM pricing API
// Required when using models whose costs are not in the built-in price table

CostTrackingOptions reference

FieldTypeDefaultNotes
perRequestnumber1.00Max USD per single LLM request
perSessionnumber5.00Max USD per agent.run() call
dailynumber20.00Max USD per calendar day
monthlynumber200.00Max USD per calendar month

RateLimiterConfig reference

FieldTypeDefaultNotes
requestsPerMinutenumber60Max LLM requests/minute
tokensPerMinutenumber100_000Max tokens/minute (input + output)
maxConcurrentnumber10Max simultaneous in-flight requests

Pitfalls

  • Budget limits are enforced per-process — multiple processes running the same agent each get their own daily/monthly counters; use an external store for true cross-process budget tracking
  • withCostTracking() with no args is still useful — it enables cost telemetry without enforcing limits (all defaults are generous)
  • withCircuitBreaker() opens on LLM provider errors, not on budget exceeded errors — they are independent systems
  • Rate limiting queues requests rather than dropping them — set maxConcurrent based on your provider's actual concurrency limits to avoid provider-side 429s
  • withDynamicPricing() makes an external HTTP call during build — ensure network access and handle build failures
  • Daily budget resets at midnight UTC by default — to use a different timezone, configure it via .withGateway({ timezone: "America/New_York" })

Signals

GitHub stars
27
Forks
4
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
cost-budget-enforcement
Source
github.com/tylerjrbuell/reactive-agents-ts