LLM FinOps & Dynamic Model Router
SkillAI & modelsLets your agent design cost-saving AI model routing with fallbacks and caching across providers.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the LLM FinOps & Dynamic Model Router skill
About this skill
Expert guide for AI FinOps, dynamic model routing (LiteLLM, Portkey), cost optimization, and latency-based fallback architectures / Panduan ahli untuk AI FinOps, routing model dinamis, optimasi biaya, dan arsitektur fallback.
What this skill tells your AI
The instructions your AI receives, as published by roedyrustam/vibes-plug in skills/llm-finops-router/SKILL.md and read by ahel’s review.
English | Bahasa Indonesia
English
Orchestration & Integration
Connects and orchestrates with ai-llm-integration-expert, vercel-ai-sdk-expert, multi-agent-orchestration, and cloud-hosting-expert.
Purpose
To optimize the operational costs, reliability, and latency of multi-model LLM architectures using dynamic routing, fallbacks, and FinOps analytics.
Key Technologies
- LiteLLM: For standardized API access and basic routing.
- Portkey: For AI Gateway, observability, caching, and robust routing rules.
- Martian / RouteLLM: For intelligent routing based on task complexity.
Architectural Guidelines
- Cost-Aware Routing: Direct simple classification or text extraction tasks to smaller models (e.g., Gemini Flash-Lite, Llama 3 8B), and complex reasoning tasks to frontier models (Claude 3.5 Sonnet/Opus, Gemini 3.1 Pro).
- Resilience & Fallbacks: Automatically fall back to secondary providers if the primary provider hits rate limits or experiences downtime.
- Semantic Caching: Implement caching at the gateway level to return immediate responses for similar queries, reducing token costs by up to 30-50%.
Bahasa Indonesia
Integrasi Orkestrasi
Terhubung dan mengorkestrasi bersama ai-llm-integration-expert, vercel-ai-sdk-expert, multi-agent-orchestration, dan cloud-hosting-expert.
Tujuan
Mengoptimalkan biaya operasional, keandalan, dan latensi arsitektur LLM multi-model menggunakan routing dinamis, fallback, dan analitik FinOps.
Teknologi Utama
- LiteLLM: Untuk akses API standar dan routing dasar.
- Portkey: Untuk AI Gateway, observabilitas, caching, dan aturan routing yang kuat.
- Martian / RouteLLM: Untuk routing cerdas berdasarkan kompleksitas tugas.
Panduan Arsitektur
- Routing Sadar Biaya: Arahkan tugas klasifikasi sederhana ke model yang lebih kecil (mis. Gemini Flash-Lite), dan tugas penalaran kompleks ke model frontier (Claude 3.5 Sonnet, Gemini 3.1 Pro).
- Ketahanan & Fallback: Otomatis mundur ke penyedia cadangan jika penyedia utama terkena rate limit atau mengalami downtime.
- Semantic Caching: Implementasikan caching di tingkat gateway untuk mengembalikan respons instan pada kueri yang serupa, mengurangi biaya token hingga 30-50%.
Signals
- GitHub stars
- 69
- Forks
- 13
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Key
llm-finops-router- Source
- github.com/roedyrustam/vibes-plug