Model requirements
Be upfront with yourself about the model you’re bringing. The minimum spec for the full experience is a model with tool calling + thinking + vision:
Hosted models that fit the full spec change monthly — check your provider’s
model card for the three capabilities rather than trusting a list. As of
mid-2026: Xiaomi MiMo-V2.5 (vision + tools + long context) fits the full
spec cheaply; DeepSeek-V3.x / GLM / MiniMax class models have strong tool
calling + thinking but text-only variants lose the vision loop; small local
models (below) generally keep tool calling and drop the rest.
Compact tool mode
The full surface is 37 tools with rich JSON schemas (~200 KB, roughly 50k tokens, pertools/list). Most non-Claude harnesses inject every registered
schema straight into the model’s context — fine for frontier models, fatal
for a 4B local one. Compact tool mode registers exactly three
meta-tools and keeps the real catalog behind them:
The model’s loop is:
list_tools → pick → describe_tool → call_tool.
Schemas enter context one tool at a time. The meta-tools are deliberately
forgiving of small-model quirks: args may be an object or a JSON-encoded
string, common field aliases (tool_name, arguments) are accepted, and
validation errors come back with the expected schema so the model can
self-correct instead of dying on an opaque protocol error.
Compact is opt-in — the direct surface is the default, so a small model
needs one of these (the flag wins over the env var):
--full is still
accepted and is now a no-op.
Auto-selection: keyed on the model, not the provider
In the panel’s local-LLM backends (Ollama / LM Studio / llama.cpp / OpenAI-compatible), when you haven’t chosen a mode, the model picks one:- a model whose id carries a parameter count at or above 70B
(
llama3.3:70b,gpt-oss:120b,mixtral:8x22b) gets the full surface; - anything smaller stays compact;
- a model id with no readable parameter count (
moonshotai/kimi-k2.5) is treated as unknown, not as small, and gets the documented compact fallback.
COMFYUI_MCP_TOOL_MODE=full
forces the full surface onto a 4B model; COMFYUI_MCP_TOOL_MODE=compact
forces the router onto a 405B one. Auto-selection only fills the gap where
nothing was chosen.
The 70B threshold is deliberately conservative: it is the only figure anyone
has actually asserted about this axis, so nothing is promoted on a guess.
COMFYUI_MCP_FULL_SURFACE_MIN_PARAMS_B=30 lowers it if you want to find where
your hardware’s ceiling really is.
The system prompt follows the mode. The compact prompt tells the model it
has six tools and routes ComfyUI through call_tool; when the full surface is
selected that is simply untrue, so the full-mode prompt says the ComfyUI tools
are advertised directly and keeps the router description for panel_* only.
Auto-selecting full while denying the tools exist would be worse than the
default it replaced.
The active mode and its reason are printed on the backend’s ready line,
e.g. Tool mode: compact — chosen for this MODEL: "qwen3:4b" is ~4B parameters, below the 70B full-surface threshold…, so the lever is never
invisible again.
This auto-selection covers the panel’s local-LLM lane. The Codex / Gemini /
Grok / Copilot HTTP lane stays pinned to compact for a different reason —
their own tool budgets crowd out the
panel_* tools otherwise — and the
standalone MCP server’s default is unchanged.Audio input
On theollama backend (native /api/chat), audio reaches the model only
where the model actually reports it can hear. Before sending, the backend asks
POST /api/show for that model’s capabilities:
audio capability, the attachment is refused
out loud — with the capability list the server reported and a pull command
for a model that can hear — rather than being dropped into the request where
the model would answer from your text alone. The same applies to a file that
isn’t an audio format, or that is present but zero bytes.
Delivering the bytes is not quite the whole job. Measured live against
gemma4:e2b: with the WAV demonstrably in context (555 prompt tokens,
/api/show reporting audio), the model still answered “I do not have the
capability to transcribe audio — my functions are limited to operating
ComfyUI”. The panel system prompt casts it as a graph operator and a small
model reasons itself out of a sense it actually has. So a turn whose audio was
capability-checked and attached also carries a short note telling the model the
audio is there and that it should answer from what it hears. With that note the
same model transcribed correctly on four runs out of four.
On the OpenAI-compatible backends (LM Studio, llama.cpp, OpenRouter,
custom) there is no capability endpoint to ask. Audio is sent as an
input_audio content part and the turn carries an explicit “I cannot confirm
the model actually receives them” line. Refusing would deny audio to every
endpoint that simply has no capability API; a guard that cannot run is not a
verdict — but it is also not a confirmation, and the wording says so.
See Backends → Audio input
for what every other provider does.
One-command setup
comfyui-mcp setup <agent> writes the server entry into the harness’s own
config file (merging with whatever is already there — existing servers,
comments in YAML, all preserved):
--compact / --full override the per-agent default,
--comfyui-url <url> embeds your ComfyUI target (local, LAN, or RunPod proxy
URL), --dry-run prints the merged config instead of writing it.
Hermes Agent
~/.hermes/config.yaml (add it by hand if you prefer):
/reload-mcp (or restart Hermes). Hermes prefixes the tools, so
the model sees mcp_comfyui_list_tools, mcp_comfyui_describe_tool, and
mcp_comfyui_call_tool — three definitions in context instead of two hundred.
On a frontier model (via Nous Portal / OpenRouter) you can rerun setup with
--full and optionally use Hermes’ own tools.include allowlist. Compact is
the right default for anything smaller.comfyui skill that drives ComfyUI over raw
REST from Python scripts. It works, but it predates this server — the MCP
route gets you workflow authoring/validation, model + custom-node management,
installer packs, queue control, and self-diagnosis. Disable the skill if the
agent keeps reaching for it instead of the MCP tools.
OpenClaw
~/.openclaw/openclaw.json:
Copilot CLI
~/.copilot/mcp-config.json:
--compact if you’re routing Copilot at a smaller model).
Check it with /mcp show inside copilot.
Our fine-tuned local models (free, recommended)
If you want to run the agent locally for free, start here. We fine-tuned the Gemma 4 family specifically for comfyui-mcp: QLoRA-trained on 1,055 server-verified tool-use trajectories synthesized against a live ComfyUI — covering the full 178-tool surface (113 MCP + 65 panel tools) — so the model knows this exact tool suite natively instead of meeting it cold.
Every rung now beats its stock base. The
:e2b v2 retrain (dual-view
training: direct tool calls AND the deployed router envelope) fixed v1’s
call_tool format regression — zero malformed envelopes across the verdict
runs. Sizing guidance stands: :e4b is the sweet spot (only ~1.5 GB more
than e2b and +4 on the arena); :e2b is now a legitimate pick for tight
VRAM; :12b buys steadiness on long multi-step tasks, not raw score.
The panel’s Ollama backend defaults to
:e4b — pick Ollama (local)
in the backend picker and it just works once the model is pulled. No account,
no API key, no per-token cost.
Vision caveat: these tags are text-only today. The current GGUF export
dropped Gemma 4’s multimodal projector
(#642), so the default
driver can’t see — get_image (action:"view") critique loops,
get_workflow (action:"from_image"), and visual comparison are unavailable,
exactly the “flies blind” row in the capability
table above. Tool calling, generation, and graph
management are unaffected. The fix on our side is shipping the mmproj with
the next publish, not retraining. If the vision loop matters to you today,
pick a vision-capable model per Model requirements
instead of the default.
Context window: the tags ship a 65,536-token window baked in, and the
orchestrator defers to it (stock models get 16K). The architecture supports up
to 128K (:e2b/:e4b) and 256K (:12b) — raise it with
COMFYUI_MCP_OLLAMA_NUM_CTX=131072 if you have the VRAM (KV cache grows with
the window). If the agent starts “forgetting” mid-conversation, watch the
orchestrator log: it warns when a turn fills ≥85% of the window. Weights, LoRA adapters, and the training
pipeline are open: artokun/gemma4-comfyui-mcp
(dataset: artokun/comfyui-mcp-trajectories).
LM Studio
The panel speaks LM Studio natively: pick LM Studio in the backend picker and the orchestrator drives its local server (http://127.0.0.1:1234/v1,
override with COMFYUI_MCP_LMSTUDIO_HOST). Setup is two clicks: install from
lmstudio.ai, then Developer → Start Server with a
tool-calling model loaded. The model picker mirrors whatever the server
offers; with no default set, the first served model is adopted automatically.
The orchestrator manages the full lifecycle hands-off: it auto-starts the
server when needed, JIT-loads your model, frees its VRAM while a ComfyUI
render runs (chat is held and answered when the render finishes), unloads the
outgoing model on a model switch, and releases everything when you switch to
a different provider.
Our fine-tuned GGUFs work here too — search artokun/gemma4-comfyui-mcp in
LM Studio’s model downloader and grab a model-q4_k_m.gguf. Expect the same
JIT cold-load pause on the first message that Ollama has (30s+ is normal).
llama.cpp (llama-server)
Running rawllama.cpp? Pick llama.cpp in the backend picker — the
orchestrator drives llama-server’s OpenAI-compatible endpoint
(http://127.0.0.1:8080/v1, override with COMFYUI_MCP_LLAMACPP_HOST):
-c) — the agent warns if
the server runs under 16K (the tool payload needs it). Tool calling is on by
default in current builds; older builds need --jinja (the panel detects a
tools-incapable server at connect and says exactly that). The single loaded
model is adopted automatically — no picking needed.
On a single-GPU box a local llama-server (or llama-swap in front of it)
joins the same VRAM handoff as Ollama and LM Studio: while a ComfyUI render
runs, your chat is held and answered the moment the render finishes. Since
llama-server has no unload API (and llama-swap swaps models upstream on
demand), the handoff is hold-only — nothing is explicitly unloaded or warmed.
A remote COMFYUI_MCP_LLAMACPP_HOST is someone else’s GPU and is never
gated. The handoff is on by default for all three local backends; opt out with
COMFYUI_MCP_PAUSE_LOCAL_ON_GEN=0 (the legacy
COMFYUI_MCP_OLLAMA_PAUSE_ON_GEN=0 is still honored).
Custom endpoint (any OpenAI-compatible server)
Anything that speaks/v1/chat/completions — vLLM, DeepSeek, Together, Azure
OpenAI, a llama-server on another box, your company’s gateway — plugs in as
the Custom endpoint provider:
- ComfyUI Settings → Comfy MCP Agent → Custom endpoint → set the
Endpoint base URL (include the
/v1, e.g.http://192.168.1.20:8000/v1). - If the server needs a key: Set API key… — a masked input; the key is
stored
0600by the orchestrator in~/.comfyui-mcp, never in ComfyUI settings or chat. - Pick Custom endpoint in the backend picker and Connect.
/v1/models; single-model servers are
adopted automatically, or set a Default model id explicitly for endpoints
that don’t list models. Env escape hatches: COMFYUI_MCP_CUSTOM_BASE_URL,
COMFYUI_MCP_CUSTOM_MODEL, COMFYUI_MCP_CUSTOM_API_KEY. The model must
support tool calling.
Ollama & local models — the LLM Arena
Any MCP harness that talks to Ollama (or an OpenAI-compatible endpoint) can drive compact mode with a local model. Two repeatable harnesses ship in the repo:npm run test:local-llm (quick single-model check) and
node scripts/llm-arena.mjs — the ComfyUI LLM Arena, which runs a field
of models through an identical task set against a live ComfyUI and verifies
every outcome against the server, never the model’s claims.
Local-tier scores on the full 10-scenario ladder (RTX 4090, ComfyUI 0.27,
temperature 0 — see the Arena page for the task ladder and the
all-tier leaderboard including frontier and hosted models):
Takeaways: the qwen3/gemma4 class solidly clears single-tool tasks (health,
installed models, registry search, queue) and picks up points on the harder
bands, but multi-stage graph composition (one graph with two piped outputs, a
staged two-stage img2img pipeline) is still frontier/B-tier territory.
llama3.1:8b’s tool-format discipline collapses on this catalog (it
hallucinates tool names and prints tool-call JSON as text). Gemma 4 shipped
native function calling across the family (Ollama ≥ v0.20);
e4b or larger
is the sweet spot.
Remember the capability ladder above: these small models keep tool calling
but have limited/no vision and thinking, so they can generate and manage
workflows yet can’t visually critique results.
The sidebar panel on a local model
The panel agent gains an Ollama backend alongside Claude / ChatGPT / Gemini: pick Ollama (local) in the backend picker and the orchestrator drives your live graph with a local model — no account, no API key, fully offline. The model sees the 6-tool router (the 3 compact comfyui meta-tools pluspanel_list_tools / panel_describe_tool /
panel_call_tool for the live canvas), so even a 4B model isn’t drowned in
schemas. Default model: artokun/gemma4-comfyui-mcp:e4b — our gemma4
fine-tune, trained on this
exact tool suite (supersedes stock gemma4:e4b, the previous Arena best);
override with COMFYUI_MCP_OLLAMA_MODEL or the panel’s model picker, which
lists whatever you’ve pulled locally. Expect honest trade-offs versus frontier
backends: slower turns (especially the first, while the model loads), no
vision, no conversation rollback.
What you get (and don’t)
Any MCP client gets the full tool surface — generation, workflow authoring, models, custom nodes, queue, diagnostics — in either tool mode. The Claude Code plugin extras (skills, slash commands, hooks, installer packs, the sidebar panel agent) are plugin features and don’t travel to other harnesses. Thelist_tools catalog is designed to carry enough orientation
that agents without that knowledge layer can still find their way.
Troubleshooting
- The model calls
call_toolwith a stringifiedargs— supported; the server parses JSON-encoded strings automatically. - The model invents tool names — unknown names return close-match
suggestions plus a pointer back to
list_tools. - Wrong/missing parameters — the error includes the tool’s JSON Schema; capable models self-correct on the next attempt.
- The model answers from the catalog without running anything — a known small-model failure mode; nudge it (“catalog entries are tool names, not data — run the tool with call_tool”).
- ComfyUI not reachable — compact mode only changes tool registration; connection config is identical to every other setup (see Configuration).