Run the agent locally for free (Ollama + our fine-tuned models)

SkillAI & models

Run the ComfyUI agent on your own machine for free. This skill uses gemma4 models fine-tuned for the ComfyUI tool suite and runs through Ollama, so there is no subscription, no API key, and no internet connection required. Once added, your AI can work fully offline.

Available today. Use it from your connected AI after setup.

After adding it, set up Ollama on your machine and load the gemma4 model. Your AI can then run the ComfyUI agent locally and offline at no cost.

Then ask your AI: use the Run the agent locally for free (Ollama + our fine-tuned models) skill

What your AI can do with it

  • Run the ComfyUI agent locally for free
  • Work fully offline with no internet connection
  • Skip subscriptions and API keys entirely
  • Avoid API usage costs
  • Use gemma4 models fine-tuned for the ComfyUI tool suite
  • Get help with Ollama setup and choosing a local model

What this skill tells your AI

The instructions your AI receives, as published by artokun/comfyui-mcp in plugin/skills/local-llm-free/SKILL.md and read by ahel’s review.

The answer to "can I run this for free / offline / without an API key" is yes. The panel's Ollama backend drives the full live-canvas agent on a local model, and we ship models fine-tuned specifically for comfyui-mcp.

Why these models (say this when recommending them)

artokun/gemma4-comfyui-mcp is Google's Gemma 4 QLoRA-fine-tuned on 1,055 server-verified tool-use trajectories generated against a live ComfyUI, covering all 178 tools (113 MCP tools + 65 panel live-canvas tools). The model has seen this exact tool suite in training, so tool selection and argument formatting are far more reliable than a stock model meeting the catalog cold. Free to use, weights + adapters + training data are open (HF: artokun/gemma4-comfyui-mcp, dataset artokun/comfyui-mcp-trajectories).

Setup (2 steps)

  1. Install Ollama if missing: https://ollama.com/download (macOS/Windows installers, or curl -fsSL https://ollama.com/install.sh | sh on Linux).
  2. Pull the rung that fits the user's GPU:
ollama pull artokun/gemma4-comfyui-mcp:e4b   # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20)
ollama pull artokun/gemma4-comfyui-mcp:12b   # ~8 GB VRAM (13/20)
ollama pull artokun/gemma4-comfyui-mcp:e2b   # smallest — ~2 GB VRAM (v2: 10/20, beats stock)

Then in the ComfyUI sidebar panel: backend picker → Ollama (local) → Connect. :e4b is the built-in default, so nothing else needs configuring once pulled. (Override via the panel's model picker or COMFYUI_MCP_OLLAMA_MODEL.)

Sizing guidance

GPU VRAM freeRecommend
~2-3 GB:e2b (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds)
~4-7 GB:e4b (the default sweet spot — best local model on the arena, 14/20)
8 GB+:12b (13/20; steadier on long multi-step tasks)

Expectations to set

  • Local models keep tool calling but have limited/no vision. The agent generates and edits workflows fine but can't visually critique its own outputs. Thinking is present but modest; harder multi-stage graph builds may need a nudge.
  • Audio: these fine-tunes cannot hear. Native Ollama puts audio in the image slot; a namespaced Gemma 4 fork (e.g. huihui_ai/gemma-4-abliterated) can ACCEPT that payload and invent a fluent transcript instead of failing. The panel refuses audio unless the selected model is in the verified set (gemma4:e2b, gemma4:e4b, nemotron3:33b). Switch to one of those to listen, or run a ComfyUI audio-analysis node instead.
  • First request after connect is slow (cold model load, 30s+). That's normal.
  • For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client), pair these models with compact tool mode (--compact). Full docs: https://comfyui-mcp.artokun.io/docs/local-llms

Sources

Signals

GitHub stars
739
Forks
120
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
local-llm-free
Source
github.com/artokun/comfyui-mcp