omni-inference
SkillAI & modelsThe core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the omni-inference skill
About this capability
AI gateway with enhanced compliance for China
What this skill tells your AI
The instructions your AI receives, as published by jaccen/airoute in skills/omni-inference/SKILL.md and read by ahel’s review.
Overview
The core OpenAI-compatible inference endpoints: chat completions, embeddings, images, audio (TTS/STT), moderations, rerank, and the Responses API. The primary integration surface for AI agents.
Authentication
All requests require a valid Bearer token or session cookie. Obtain a token via POST /api/auth/login or configure REQUIRE_API_KEY=false for local development.
Endpoints
POST /api/v1/chat/completions
Create chat completion
OpenAI-compatible chat completions endpoint. Routes to configured providers.
curl -X POST https://localhost:20128/api/v1/chat/completions \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
GET /api/v1/ws
Chat completion over WebSocket (handshake + upgrade)
OpenAI-compatible chat over a WebSocket connection. GET with ?handshake=1 returns the connection descriptor (auth path, message protocol and live-event channels) as JSON; a plain GET without an Upgrade returns 426 Upgrade Required. After upgrading, the client exchanges JSON frames — {type:"request", id, payload:{model, messages}} to start a completion and {type:"cancel", id} to abort it. A separate live channel (default port LIVE_WS_PORT=20129, path /live) streams dashboard events on the requests, combo and credentials topics with a 15s heartbeat. Requires an API key.
curl https://localhost:20128/api/v1/ws \
-H "Authorization: Bearer $AIRoute_TOKEN"
POST /api/v1/providers/{provider}/chat/completions
Create chat completion (provider-specific)
Routes to a specific provider by name.
curl -X POST https://localhost:20128/api/v1/providers/{provider}/chat/completions \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/api/chat
Ollama-compatible chat endpoint
Provides compatibility with Ollama's /api/chat format.
curl -X POST https://localhost:20128/api/v1/api/chat \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/messages
Create message (Anthropic-compatible)
Anthropic Messages API endpoint. Routes to Claude providers.
curl -X POST https://localhost:20128/api/v1/messages \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/messages/count_tokens
Count tokens for a message
curl -X POST https://localhost:20128/api/v1/messages/count_tokens \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/responses
Create response (OpenAI Responses API)
OpenAI Responses API endpoint.
curl -X POST https://localhost:20128/api/v1/responses \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/embeddings
Create embeddings
curl -X POST https://localhost:20128/api/v1/embeddings \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/providers/{provider}/embeddings
Create embeddings (provider-specific)
curl -X POST https://localhost:20128/api/v1/providers/{provider}/embeddings \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/images/generations
Generate images
curl -X POST https://localhost:20128/api/v1/images/generations \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/providers/{provider}/images/generations
Generate images (provider-specific)
curl -X POST https://localhost:20128/api/v1/providers/{provider}/images/generations \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/audio/speech
Generate speech audio
Text-to-speech endpoint. Routes to configured TTS providers.
curl -X POST https://localhost:20128/api/v1/audio/speech \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/audio/transcriptions
Transcribe audio
Audio-to-text transcription endpoint.
curl -X POST https://localhost:20128/api/v1/audio/transcriptions \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/moderations
Create moderation
Content moderation endpoint. Routes to configured moderation providers.
curl -X POST https://localhost:20128/api/v1/moderations \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/rerank
Rerank documents
Document reranking endpoint.
curl -X POST https://localhost:20128/api/v1/rerank \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
GET /api/v1
API v1 root endpoint
Returns basic API info and status.
curl https://localhost:20128/api/v1 \
-H "Authorization: Bearer $AIRoute_TOKEN"
GET /api/v1/providers/{provider}/models
List models for a specific provider
Returns only models for the selected provider with provider prefix removed from each model id.
curl https://localhost:20128/api/v1/providers/{provider}/models \
-H "Authorization: Bearer $AIRoute_TOKEN"
POST /api/v1/ocr
Document OCR
Mistral OCR–compatible document OCR endpoint. Accepts a JSON body referencing a document/image and returns extracted text. Success responses carry the X-AIRoute-* cost-telemetry headers.
curl -X POST https://localhost:20128/api/v1/ocr \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
POST /api/v1/audio/translations
Translate audio to English
OpenAI Whisper–compatible audio translation (multipart/form-data). Unlike /api/v1/audio/transcriptions, output is always English regardless of the source language. Success responses carry the X-AIRoute-* cost-telemetry headers.
curl -X POST https://localhost:20128/api/v1/audio/translations \
-H "Authorization: Bearer $AIRoute_TOKEN"
-H "Content-Type: application/json" \
-d '{}'
GET /api/v1/providers/suggested-models
Suggested media models
Read-only server-side proxy to the public HuggingFace Hub models search API, used by the dashboard to suggest models for a media provider kind without exposing an HF token client-side. Never accepts or returns credentials.
curl https://localhost:20128/api/v1/providers/suggested-models \
-H "Authorization: Bearer $AIRoute_TOKEN"
GET /api/v1/provider-plugin-manifest
Provider plugin manifest
Returns the manifest describing installed provider plugins.
curl https://localhost:20128/api/v1/provider-plugin-manifest \
-H "Authorization: Bearer $AIRoute_TOKEN"
Payloads
See the full OpenAPI specification at GET /api/openapi/spec or docs/openapi.yaml for detailed request/response schemas.
Chat completions
Requires AIRoute_URL and AIRoute_KEY. See entry-point SKILL for setup.
Endpoints
POST $AIRoute_URL/v1/chat/completions— OpenAI formatPOST $AIRoute_URL/v1/messages— Anthropic Messages formatPOST $AIRoute_URL/v1/responses— OpenAI Responses API
Discover
curl $AIRoute_URL/v1/models | jq '.data[].id'
Combos (e.g. auto, cost-optimized, subscription) auto-fallback through multiple providers.
OpenAI format example
curl -X POST $AIRoute_URL/v1/chat/completions \
-H "Authorization: Bearer $AIRoute_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"messages": [{"role": "user", "content": "Refactor this function"}],
"stream": true
}'
Anthropic format example
curl -X POST $AIRoute_URL/v1/messages \
-H "Authorization: Bearer $AIRoute_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"max_tokens": 4096,
"messages": [{"role": "user", "content": "Hi"}]
}'
Tool use
Supports OpenAI tools array and Anthropic tools block. Tool results
auto-compressed via RTK (47 filters: git-diff, grep, test-jest, terraform-plan,
docker-logs, etc.) — 20-40% token savings. Disable per-request with
X-AIRoute-Rtk: off header.
Reasoning / thinking
Anthropic extended thinking and OpenAI Responses reasoning blocks are forwarded verbatim. Cached automatically via reasoning cache.
Errors
401→ invalid API key400 invalid_model→ model not in registry; check/v1/models503 circuit_open→ provider circuit breaker tripped; retry later or use combo429 rate_limited→ honorRetry-After; consider using a combo for auto-fallback
Image generation
Requires AIRoute_URL and AIRoute_KEY. See entry-point SKILL for setup.
Endpoints
POST $AIRoute_URL/v1/images/generations— Text-to-imagePOST $AIRoute_URL/v1/images/edits— Image edit (mask)POST $AIRoute_URL/v1/images/variations— Variations
Discover
curl $AIRoute_URL/v1/models/image | jq '.data[]'
Returns { id, owned_by, sizes:[...], capabilities:[...] } per model.
Generate example
curl -X POST $AIRoute_URL/v1/images/generations \
-H "Authorization: Bearer $AIRoute_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "dall-e-3",
"prompt": "a red bicycle on a wet street, photoreal",
"n": 1,
"size": "1024x1024",
"response_format": "b64_json"
}'
Response: { created, data: [{ url? or b64_json, revised_prompt }] }
Errors
400 invalid_size→ not supported by this model; check/v1/models/image400 content_policy_violation→ blocked by provider safety503→ provider unavailable; try another model in/v1/models/image
Text-to-speech
Requires AIRoute_URL and AIRoute_KEY. See entry-point SKILL for setup.
Endpoint
POST $AIRoute_URL/v1/audio/speech— returns binary audio (mp3/opus/wav/flac)
Discover
curl $AIRoute_URL/v1/models/tts | jq '.data[]'
Each entry includes voices:[...] for the available voice names per provider.
Example
curl -X POST $AIRoute_URL/v1/audio/speech \
-H "Authorization: Bearer $AIRoute_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "Hello from AIRoute.",
"voice": "alloy",
"response_format": "mp3"
}' --output speech.mp3
Voices
Voice names vary by provider. Check /v1/models/tts — each entry has voices:[...].
Common OpenAI voices: alloy, echo, fable, onyx, nova, shimmer.
Errors
400 invalid_voice→ voice not supported by this model400 input_too_long→ input exceeds model character limit503→ provider unavailable; try another model in/v1/models/tts
Speech-to-text
Requires AIRoute_URL and AIRoute_KEY. See entry-point SKILL for setup.
Endpoints
POST $AIRoute_URL/v1/audio/transcriptions— multipart upload, returns textPOST $AIRoute_URL/v1/audio/translations— transcribe + translate to English
Discover
curl $AIRoute_URL/v1/models/stt | jq '.data[]'
Example
curl -X POST $AIRoute_URL/v1/audio/transcriptions \
-H "Authorization: Bearer $AIRoute_KEY" \
-F "file=@audio.mp3" \
-F "model=whisper-1" \
-F "response_format=verbose_json"
Response: { text, language, duration, segments?:[{ start, end, text }] }
Supported formats
Audio: mp3, mp4, mpeg, mpga, m4a, wav, webm.
Response formats: json, text, srt, verbose_json, vtt.
Errors
400 invalid_file_format→ unsupported audio format400 file_too_large→ exceeds provider limit (usually 25MB)503→ provider unavailable; try another model in/v1/models/stt
Embeddings
Requires AIRoute_URL and AIRoute_KEY. See entry-point SKILL for setup.
Endpoint
POST $AIRoute_URL/v1/embeddings
Discover
curl $AIRoute_URL/v1/models/embedding | jq '.data[]'
Each entry: { id, owned_by, dimensions, max_input_tokens }.
Example
curl -X POST $AIRoute_URL/v1/embeddings \
-H "Authorization: Bearer $AIRoute_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-large",
"input": ["first text", "second text"],
"encoding_format": "float"
}'
Response: { data:[{ embedding:[...], index }], usage:{ prompt_tokens, total_tokens } }
Batch input
input accepts a string or array of strings (up to provider batch limit, typically 2048 items).
Errors
400 input_too_long→ input exceedsmax_input_tokensfor this model400 invalid_encoding_format→ usefloatorbase64503→ provider unavailable; try another model in/v1/models/embedding
Web search
Requires AIRoute_URL and AIRoute_KEY. See entry-point SKILL for setup.
Endpoint
POST $AIRoute_URL/v1/web/search— unified search format
Discover
curl $AIRoute_URL/v1/models/web | jq '.data[] | select(.kind == "webSearch")'
Example
curl -X POST $AIRoute_URL/v1/web/search \
-H "Authorization: Bearer $AIRoute_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tavily/search",
"query": "AIRoute github latest release",
"max_results": 5,
"include_answer": true
}'
Response: { answer?, results:[{ url, title, content, score }] }
Parameters
| Field | Type | Description |
|---|---|---|
model | string | Provider model from /v1/models/web |
query | string | Search query |
max_results | number | Max results (default: 5) |
include_answer | boolean | Include AI-synthesized answer |
search_depth | string | basic or advanced (Tavily) |
Errors
400 query_too_long→ shorten the search query503→ provider unavailable; try another model in/v1/models/web
Web fetch
Requires AIRoute_URL and AIRoute_KEY. See entry-point SKILL for setup.
Endpoint
POST $AIRoute_URL/v1/web/fetch
Discover
curl $AIRoute_URL/v1/models/web | jq '.data[] | select(.kind == "webFetch")'
Example
curl -X POST $AIRoute_URL/v1/web/fetch \
-H "Authorization: Bearer $AIRoute_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jina/reader",
"url": "https://anthropic.com",
"format": "markdown"
}'
Response: { url, title, markdown, links?:[...], images?:[...] }
Parameters
| Field | Type | Description |
|---|---|---|
model | string | Provider from /v1/models/web (e.g. jina/reader, firecrawl/scrape) |
url | string | URL to fetch |
format | string | markdown (default), html, text |
Errors
400 invalid_url→ URL must be http/https403 blocked→ provider blocked by target site; try a different model503→ provider unavailable; try another model in/v1/models/web
Signals
- GitHub stars
- 32
- Forks
- 8
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
omni-inference-jaccen- Source
- github.com/jaccen/airoute