Add a New AI Model to Kiln
SkillCommunicationLets your agent add new AI models to a model list file and draft a Discord announcement about it.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Add a New AI Model to Kiln skill
About this capability
Add new AI models to Kiln's ml_model_list.py and produce a Discord announcement. Use when the user wants to add, integrate, or register a new LLM model (e.g. Claude, GPT, DeepSeek, Gemini, Kimi, Qwen, Grok) into the Kiln model list, mentions adding a model to ml_model_list.py, asks to discover/find
What this skill tells your AI
The instructions your AI receives, as published by kiln-ai/kiln in .agents/skills/claude-maintain-models/SKILL.md and read by ahel’s review.
Branch check first: if the request involves a provider Kiln does not support yet, start at Adding a Net-New Provider. That workflow spans core, server, UI, tests, and tooling, and is gated on a client release — the model-entry steps below are only its final, gated piece.
For a new model on an already-supported provider, integrating it into
libs/core/kiln_ai/adapters/ml_model_list.py requires:
ModelNameenum – add an enum memberbuilt_in_modelslist – add aKilnModel(...)entry with providersModelFamilyenum – only if the vendor is brand-new
After code changes, run paid integration tests, then draft a Discord post.
Global Rules
These apply throughout the entire workflow.
- Slug verification: NEVER guess or infer model slugs from naming patterns. Every
model_idmust come from an authoritative source (LiteLLM catalog, official docs, API reference, or changelog). If you can't verify a slug, tell the user and ask them to provide it. - Date awareness: These models are often released very recently. Web search for current info before assuming you know the details.
Phase 1 – Model Discovery (only when asked to find new/missing models)
If the user asks you to find new models, do NOT just web search "new AI models this week" — that only surfaces major releases. Instead, systematically check each family against both the LiteLLM catalog and models.dev, then union the results. Both are attempts to catalog available models and each has gaps the other fills.
-
Read the
ModelFamilyandModelNameenums to know what we already have. -
Query both catalogs for each family (run in parallel where possible):
LiteLLM catalog — filters out mirror providers to avoid duplicates:
curl -s 'https://api.litellm.ai/model_catalog?model=SEARCH_TERM&mode=chat&page_size=500' -H 'accept: application/json' | jq '[.data[] | select(.provider != "openrouter" and .provider != "bedrock" and .provider != "bedrock_converse" and .provider != "vertex_ai-anthropic_models" and .provider != "azure") | .id] | unique | .[]'models.dev — search all model IDs across all providers:
curl -s https://models.dev/api.json | jq '[to_entries[].value.models // {} | keys[]] | .[]' | grep -i "SEARCH_TERM"For details on a specific provider+model:
curl -s https://models.dev/api.json | jq '.["PROVIDER"].models["MODEL_ID"]' -
Search terms (one query per term):
claude,gpt,o1,o3,o4(OpenAI reasoning),gemini,llama,deepseek,qwen,qwq,mistral,grok,kimi,glm,minimax,hunyuan,ernie,phi,gemma,seed,step,pangu -
Union and cross-reference results from both catalogs against
ModelName. A model found in either source counts as available. Focus on direct-provider entries (not OpenRouter/Bedrock/Azure mirrors). Skip pure coding models (e.g.codestral,deepseek-coder,qwen-coder). -
Run targeted web searches per family to catch very fresh releases not yet in either catalog:
"[family] new model [current year]""[family] release [current month] [current year]"
-
Present findings as a summary. Let the user decide which to add.
Phase 1B – Lagging-Provider Backfill Check (every run)
Some providers — Fireworks AI, Together AI, SiliconFlow — expose new models on their own endpoints 1–2 weeks before those entries surface in models.dev / LiteLLM. Relying only on those two catalogs will both under-populate the provider list for the model you're adding now and miss the window to backfill recently-added models whose provider support has since grown.
Run this check on every invocation of the skill, regardless of whether you're in discovery mode or adding a specific model.
-
Pull the 10 most recently added models from git history. List position is NOT a recency signal — entries are ordered by family/version/size (see 3c), and net-new families sit at the END of the list, so recent additions can be anywhere:
git log --follow -p -- libs/core/kiln_ai/adapters/ml_model_list.py | grep -E "^\+\s+name=ModelName\." | head -20 -
For the model you're adding (if any) AND each of those 10 models, cross-check Fireworks, Together, and SiliconFlow directly using the endpoints in the Lagging Providers Reference. Do NOT trust
models.dev/ LiteLLM as the final word for these three providers. -
If a lagging provider now supports a recently-added model that isn't yet in its
KilnModelentry, flag it to the user and propose either bundling the provider addition into the current change or opening a separate PR. Do not silently add it.
Phase 2 – Gather Context
-
Read the predecessor model in
ml_model_list.py(e.g. for Opus 4.6 → read Opus 4.5). You inherit most parameters from it. -
Query the LiteLLM catalog for the new model. This is the primary slug source since Kiln uses LiteLLM. See the Slug Lookup Reference for query syntax and all verified sources.
-
Get the OpenRouter slug via:
curl -s https://openrouter.ai/api/v1/models | jq '.data[].id' | grep -i "SEARCH_TERM"- Fallback: WebSearch for
openrouter [model name] model id
-
Get the direct-provider slug (Anthropic, OpenAI, Google, etc.). Use the LiteLLM catalog first, then official docs. See the Slug Lookup Reference for provider-specific URLs.
-
Identify quirks — check the Provider Quirks Reference for the relevant provider, and web search for any new quirks:
- Structured output mode (JSON schema vs function calling)?
- Reasoning model (needs
reasoning_capable, parsers, OpenRouter options)? - Vision/multimodal support? Which MIME types?
- Provider-specific flags (
temp_top_p_exclusive, etc.)? - Rate limit concerns (
max_parallel_requests)?
-
Determine thinking levels — does the model support configurable reasoning effort? See Thinking Levels Reference for the full lookup chain. Key quick checks:
- Check the vendor model page (e.g. OpenAI model pages say "Reasoning.effort supports: X, Y, Z")
- Check OpenRouter
supported_parameters— ifreasoningis absent, skip thinking levels - R1-style thinking models (DeepSeek, Qwen thinking variants) do NOT get thinking level dicts
Phase 3 – Code Changes
All changes go in libs/core/kiln_ai/adapters/ml_model_list.py.
3a. ModelName enum
- snake_case:
claude_opus_4_6 = "claude_opus_4_6" - Place before predecessor (newer first within group). If the vendor is brand-new there is no predecessor — start a new group at the end of the enum
- Follow existing grouping (all claude together, all gpt together, etc.)
3b. KilnModel entry in built_in_models
- Place per the ordering rules in 3c — this placement is user-visible, get it right
- Copy predecessor's structure and modify:
name,friendly_name,model_idper provider, flags friendly_namemust follow the existing naming pattern of sibling models in the same family. Check the predecessor. For example, Claude Sonnets use"Claude {version} Sonnet"(e.g. "Claude 4.5 Sonnet"), not"Claude Sonnet {version}". Do NOT use the vendor's marketing name if it differs from Kiln's established convention.
Provider model_id formats:
| Provider | Format | Notes |
|---|---|---|
openrouter | vendor/model-name | Always verify via API |
openai | Bare model name | Verify via OpenAI docs |
anthropic | Variable — older models have date stamps, newer may not | Always verify via Anthropic docs |
gemini_api | Bare name | Verify via Google AI Studio docs |
fireworks_ai | accounts/fireworks/models/... | Verify via Fireworks docs |
together_ai | Vendor path format | Verify via Together docs |
vertex | Usually same as gemini_api | Verify via Vertex docs |
siliconflow_cn | Vendor/model format | Verify via SiliconFlow docs |
featherless_ai | HuggingFace repo id, case-sensitive (zai-org/GLM-5.2) | Verify via their /v1/models — see Featherless |
Every single model_id must be verified from an authoritative source. No exceptions.
Setting flags — use catalog data + predecessor as dual signals:
The LiteLLM catalog and models.dev responses include capability flags (supports_vision, supports_function_calling, supports_reasoning, etc.). Use these as the primary signal for what to enable on the new model:
- If the catalog says
supports_vision: true→ enablesupports_vision,multimodal_capable, and vision MIME types (see 2c) - If the catalog says
supports_function_calling: true→ useStructuredOutputMode.json_schema(orfunction_callingdepending on provider norms — check predecessor) - If the catalog says
supports_reasoning: true→ the model can reason, but do NOT reflexively setreasoning_capable=True— default toreasoning_capable=False(see Reasoning Capable Default). Still addavailable_thinking_levelsif it supports effort levels, and check parser/formatter flags.
Then cross-check against the predecessor. The predecessor tells you how Kiln configures a similar model (which structured_output_mode, which provider-specific flags, etc.). The catalog tells you what the model can do. Use both:
- Catalog says the model supports vision but predecessor doesn't have it? Enable it — this is a new capability.
- Predecessor has
temp_top_p_exclusivebut nothing in the catalog mentions it? Keep it — it's a provider quirk the catalog doesn't track. - Catalog and predecessor disagree on something? Trust the catalog for capabilities, trust the predecessor for Kiln-specific configuration patterns.
Common flags:
structured_output_mode– how the model handles JSON outputsuggested_for_evals/suggested_for_data_gen– see zero-sum rule belowmultimodal_capable/supports_vision/supports_doc_extraction– see multimodal rules belowreasoning_capable– for thinking/reasoning models. Default new models toreasoning_capable=Falseunless the model always emits its reasoning (see Reasoning Capable Default)temp_top_p_exclusive– Anthropic models that can't have both temp and top_pparser/formatter– for models needing special parsing (e.g. R1-style thinking)
2c. Multimodal capabilities
If the model supports non-text inputs, configure:
multimodal_capable=Trueandsupports_doc_extraction=Trueif it supports any MIME typessupports_vision=Trueif it supports imagesmultimodal_requires_pdf_as_image=Trueif vision-capable but no native PDF support (also addKilnMimeType.PDFto MIME list). Always set this on OpenRouter providers — OpenRouter routes PDFs through Mistral OCR which breaks LiteLLM parsing.- Always include
KilnMimeType.TXTandKilnMimeType.MDon anymultimodal_capablemodel
Strategy: start broad, narrow based on test failures. Enable a generous set of MIME types, run tests, and remove only types the provider explicitly rejects (400 errors). Don't remove types for timeout/auth/content-mismatch failures.
Full MIME superset (Gemini uses all):
# documents
KilnMimeType.PDF, KilnMimeType.CSV, KilnMimeType.TXT, KilnMimeType.HTML, KilnMimeType.MD
# images
KilnMimeType.JPG, KilnMimeType.PNG
# audio
KilnMimeType.MP3, KilnMimeType.WAV, KilnMimeType.OGG
# video
KilnMimeType.MP4, KilnMimeType.MOV
3c. Ordering — the list IS the UI
The order of built_in_models is exactly the order users see: model dropdowns
group by provider and list each provider's models in list order, and the model
library page lists models in list order. A misplaced entry ships a scrambled
dropdown to every client via the remote config. Rules:
-
Families are contiguous. Every model of a family sits in one block. Never append a new family member at the bottom of the file or after another family — that strands it (past bugs: Mistral Small 4/3 ended up inside the Qwen 2.5 region; GLM-Z1 models ended up after the Kimi family).
-
Within a family, versions run newest → oldest, top to bottom. 4 > 3.8 > 3.7 > 3.6 > 3.5 > 3 > 2.5. A new version goes at the TOP of the family block. A new model of an existing version goes inside that version's group — NOT at the top of the family and NOT below older versions (past bugs: Gemini 3.6/3.7 Flash were inserted mid-3.1; Qwen 3.6/3.7 landed below Qwen 3.5 entries; Phi 3.5 sat above Phi 4).
-
Within a version, big → small. Always.
- Commercial tiers: Max > Plus > Flash; Pro > Flash > Flash Lite; Large > Medium > Small.
- Open-weight sizes descending: 405B > 70B > 8B. Keep base/Non-Thinking variant pairs adjacent; variant sub-groups (e.g. the Qwen VL Instruct and VL Thinking blocks) stay intact, ordered descending internally.
- Tier-grouped families follow the same rule for their tier blocks: Claude runs Fable, then Opus > Sonnet > Haiku, versions descending inside each tier. (It once ran Haiku-first, and several families ran sizes ascending — those were bugs, not conventions. Do not preserve an ascending run because it's "what the family already does.")
-
A net-new family's block goes at the END of
built_in_models(this is the existing convention — the newest niche vendors sit at the bottom of the list). The "place before predecessor" rule only applies within an existing family; a new family has no predecessor. -
Verify after editing — print the order you just shipped and eyeball the affected family:
uv run python -c " from kiln_ai.adapters.ml_model_list import built_in_models for m in built_in_models: print(m.family, '|', m.friendly_name)"
3d. suggested_for_evals / suggested_for_data_gen
Only set these if the predecessor already has them, OR web search shows the model is a clear SOTA leap (ask user to confirm first).
Zero-sum rule: When adding a new model with these flags, remove them from the oldest same-family model to keep the suggested count stable. Ask the user to confirm the swap before making changes.
3e. ModelFamily enum (only if needed)
Only add a new family if the vendor is completely new.
3f. Thinking Levels (available_thinking_levels / default_thinking_level)
If the model supports configurable reasoning effort (not just on/off), add available_thinking_levels and default_thinking_level to each provider entry. See Thinking Levels Reference for the full lookup chain and existing constants.
Quick rules:
- Reuse an existing
_THINKING_LEVELSconstant if the levels match exactly - Create a new constant only if levels differ; name it
{MODEL}_{PROVIDER_CONTEXT}_THINKING_LEVELS default_thinking_levelmust be one of the values inavailable_thinking_levels
Phase 4 – Run Tests
Tests call real LLMs and cost money. Ideally the user only needs to consent to two script executions: the smoke test, then the full parallel suite.
Vertex AI authentication: Vertex tests require active gcloud credentials. If you are changing a model that uses Vertex, you must not run the test until asking the user to run gcloud auth application-default login before trying. These failures are auth issues, not model config problems.
-k filter syntax: Always use bracket notation for model+provider filtering, never and:
- Good:
-k "test_name[glm_5-fireworks_ai]"or-k "glm_5" - Bad:
-k "glm_5 and fireworks"—andis a pytest keyword expression that can match wrong tests
4.0 — If the test env can't build (blocked git dependency)
Kiln's core lib pins together to a git fork (libs/core/pyproject.toml: together = { git = "https://github.com/scosman/together-python" }). In a sandboxed environment whose GitHub access is scoped to kiln-ai/kiln only (e.g. Claude Code Web), uv sync fails to fetch that fork with a 403 from the git proxy, so the test venv can't be built. This is a GitHub repo-scope block, not a network-domain allowlist issue — every non-scoped repo 403s, only kiln-ai/kiln resolves.
Workaround to build the venv for testing (the together fork isn't exercised by model-integration tests, which route through LiteLLM):
- Do NOT use
uv sync --no-sources— it strips the workspaceworkspace = truesources too and makeskiln-root → kiln-aiunsatisfiable. - Instead, temporarily comment out ONLY the
together = { git = ... }line inlibs/core/pyproject.toml, then runuv sync(this resolvestogetherfrom PyPI). - Run the tests.
- Revert the workaround with
git checkout -- libs/core/pyproject.toml uv.lock— this restores the commented-outtogethergit-fork pin inlibs/core/pyproject.toml(NOT the repo-rootpyproject.toml, which the workaround never touches) and revertsuv.lockif it changed. Onlyml_model_list.py(and any intended test-file edits) should remain modified.
Note: the PyPI together may pull slightly different transitive deps (e.g. a newer starlette), which can cause unrelated collection ImportErrors in desktop/server/rag/vector-store modules — scope your -k filters to the model files and ignore those.
4a. Parallel testing + API keys
Parallel testing is already on. There is no pytest.ini — pytest config lives in [tool.pytest.ini_options] in the root pyproject.toml, and addopts = "-n auto" is active. No edit and no revert are needed. (Only override to -n 8 if a provider rate-limits you.)
Paid tests read API keys from the ENVIRONMENT, not from the Kiln app's settings. conftest.py has an autouse use_temp_settings_dir fixture that points Config.settings_path at a temp dir, so ~/.kiln_ai/settings.yaml is deliberately ignored during tests. A key the user added through the app's provider page will not be seen — the test fails with "Attempted to use X without an API key set", which looks like a config bug but isn't.
Bridge the key from the user's settings into the test environment without ever printing it:
export FEATHERLESS_AI_API_KEY="$(uv run python -c \
"from kiln_ai.utils.config import Config; print(Config.shared().featherless_ai_api_key)" 2>/dev/null | tail -1)"
Use the provider's env_var name from libs/core/kiln_ai/utils/config.py. To check which keys are available before running, print booleans only — never the values:
uv run python -c "from kiln_ai.utils.config import Config; c=Config.shared(); print(bool(c.fireworks_api_key))"
Many paid tests also carry the ollama marker, so --runpaid alone silently skips them. Always pass --runpaid --ollama together, and use -rs to see skip reasons when a test you expected to run reports as skipped.
4b. Smoke test — verify slug works
Run a single test+provider combo first:
uv run pytest --runpaid --ollama -k "test_data_gen_sample_all_models_providers[MODEL_ENUM-PROVIDER]"
If it fails, fix the slug/config before proceeding. Use --collect-only to find exact parameter IDs if unsure.
4c. Full test suite
uv run pytest --runpaid --ollama -k "MODEL_ENUM" -v 2>&1 | grep -E "PASSED|FAILED|ERROR|short test|=====|collected"
If tests fail — debug one at a time:
- Pick ONE failing test, run it with
-vfor full output - Fix the config
- Re-run that single test to verify
- Only re-run the full suite once the single test passes
Anthropic API key gotcha: if an Anthropic-direct test fails with an auth/API key error, check whether the user's environment exports the key as KILN_ANTHROPIC_API_KEY instead of ANTHROPIC_API_KEY (the Kiln app uses the prefixed name; the Anthropic SDK used by tests expects the unprefixed name). Prepend the test command with a one-shot alias — don't export it globally:
ANTHROPIC_API_KEY="$KILN_ANTHROPIC_API_KEY" uv run pytest --runpaid ...
4d. Extraction tests (if supports_doc_extraction=True)
Tests are in libs/core/kiln_ai/adapters/extractors/test_litellm_extractor.py.
# See what will run:
uv run pytest --collect-only libs/core/kiln_ai/adapters/extractors/test_litellm_extractor.py::test_extract_document_success -q | grep MODEL_ENUM
# Run them:
uv run pytest --runpaid --ollama libs/core/kiln_ai/adapters/extractors/test_litellm_extractor.py::test_extract_document_success -k "MODEL_ENUM"
If a provider rejects a data type (400 error), remove that KilnMimeType and re-run.
4e. Confirm failures are actually yours
Before treating a failure as a problem with your change, check whether the same test already fails for an existing provider of that model. Several assertions are provider-independent and fail regardless.
Known example: test_structured_input_cot_prompt_builder asserts len(trace) == 5 unconditionally, which is incompatible with any provider setting reasoning_capable=True (that selects the single-call strategy, producing 3 messages). It fails for gpt_oss_120b on fireworks_ai on a clean tree.
uv run pytest --runpaid --ollama -q "path::test_name[MODEL-OTHER_PROVIDER]"
If it fails there too, it's pre-existing — report it as such rather than contorting the config to work around it.
4f. Test output format
Collect test results for use in the PR body (Phase 5). Organize by model name and provider using these symbols:
- ✅ for passed tests
- ⚠️ for tests that failed due to content quality flakes (e.g. model returned fewer items than expected, weak assertion mismatches) — include a brief reason
- ❌ for tests that failed due to real errors (bad slug, unsupported feature, 400/500 errors) — include a brief reason
- List every test using the full pytest parametrize ID, grouped by provider
- Include extraction tests (Phase 4d) if they were run
Phase 5 – Create Pull Request
5.0 — Important context about Claude Code Web's stop hook
This skill is often run via Claude Code Web (Slack connector). That environment has a non-user-configurable stop hook which, at end of session, will:
- Block the session from ending if there are uncommitted changes, untracked files, or unpushed commits
- Instruct the agent to commit and push any local work before stopping
- Explicitly tell the agent NOT to create a PR unless the user asked for one
The problems this causes:
- When tests fail mid-skill, the agent has historically pushed a half-broken branch to satisfy the hook, leaving a graveyard of abandoned
add-model/*branches on the remote. - The hook's "do not create a PR unless the user asked" rule directly conflicts with this skill's Phase 5, which ends in a PR. Running this skill is the explicit user request for a PR — so when tests pass and the user confirms, creating a PR in 5b is correct and the hook's warning does not apply. Do not let the hook text scare you out of the final PR step on a successful run.
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 5k
- Forks
- 378
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
claude-maintain-models- Source
- github.com/kiln-ai/kiln