artokun/comfyui-mcp/plugin/skills/troubleshooting/SKILL.md
troubleshooting
Common ComfyUI errors and fixes. OOM, missing nodes, dtype mismatches, black images, and debugging strategies
- Source repository stars
- 690
- Declared platforms
- 0
- Static risk flags
- 1
- Last source update
- 2026-08-28
- Source checked
- 2026-08-28
Decision brief
What it does: where it fits
Render completes but looks WRONG (artifacts, wrong subject/pose/color, a ControlNet/mask/LoRA not taking, a refiner degrading it)? That's not an error. Use the debug-render skill (listpacks with action: "skillread", name: "debug-render") to localize the bad stage with run-to-nod…
Not for
- Render completes but looks WRONG (artifacts, wrong subject/pose/color, a ControlNet/mask/LoRA not taking, a refiner degrading it)? That's not an error. Use the debug-render skill (listpacks with action: "skillread", nam…
Compatibility matrix
Platform support, with evidence labels
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
Inspect first. Install second.
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/artokun/comfyui-mcp --skill "plugin/skills/troubleshooting"Inspect the Agent Skill "troubleshooting" from https://github.com/artokun/comfyui-mcp/blob/3a950ff3d32d6ed219c921974d7fb1cd7aa5eba3/plugin/skills/troubleshooting/SKILL.md at commit 3a950ff3d32d6ed219c921974d7fb1cd7aa5eba3. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
What the source asks the agent to do
- 01
Workflow Failed — Get Details
The response includes: - status.statusstr: "success" or "error" - status.messages: Timestamped execution messages - outputs: Node outputs (images, etc.) - Error traceback for failed nodes
status.statusstr: "success" or "error"status.messages: Timestamped execution messagesoutputs: Node outputs (images, etc.) - 02
Error Diagnosis Strategy
When a workflow fails, follow this approach:
Get the error. Use gethistory(action="diagnose") to retrieve the execution result with the full traceback, plus any missing models/nodesCheck logs. Use getsystemstats (action:"logs") with keyword filters like "error", "warning", "traceback"Identify the failing node. The history response includes the nodeid and nodetype that failed - 03
Out of Memory (OOM)
The GPU does not have enough VRAM to hold the model weights, intermediate tensors, and latent images at the same time. Common triggers: - High resolution images (2048x2048+) - Multiple models loaded at the same time - FP32 precision models on limited VRAM - Video generation (LTX…
High resolution images (2048x2048+)Multiple models loaded at the same timeFP32 precision models on limited VRAM - 04
Error Pattern
Review the “Error Pattern” section in the pinned source before continuing.
Review and apply the “Error Pattern” source section. - 05
Root Cause
The GPU does not have enough VRAM to hold the model weights, intermediate tensors, and latent images at the same time. Common triggers: - High resolution images (2048x2048+) - Multiple models loaded at the same time - FP32 precision models on limited VRAM - Video generation (LTX…
High resolution images (2048x2048+)Multiple models loaded at the same timeFP32 precision models on limited VRAM
Permission review
Static risk signals and limitations
Network access
The documentation includes network, browsing, or remote request actions.
download_model({ action: "download", url: "...", target_subfolder: "checkpoints" })Evidence record
Why each signal appears
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 96/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 690 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Provenance and original SKILL.md
- Repository
- artokun/comfyui-mcp
- Skill path
- plugin/skills/troubleshooting/SKILL.md
- Commit
- 3a950ff3d32d6ed219c921974d7fb1cd7aa5eba3
- License
- MIT
- Collected
- 2026-08-28
- Default branch
- main
View the original SKILL.md
ComfyUI Troubleshooting Guide
Render completes but looks WRONG (artifacts, wrong subject/pose/color, a ControlNet/mask/LoRA not taking, a refiner degrading it)? That's not an error. Use the debug-render skill (
list_packswithaction: "skill_read",name: "debug-render") to localize the bad stage with run-to-node (panel_runto_node_id) by previewing intermediate steps. This guide is for runs that fail with an error, OOM, or missing node.
Error Diagnosis Strategy
When a workflow fails, follow this approach:
- Get the error. Use
get_history(action="diagnose")to retrieve the execution result with the full traceback, plus any missing models/nodes - Check logs. Use
get_system_stats (action:"logs")with keyword filters like"error","warning","traceback" - Identify the failing node. The history response includes the
node_idandnode_typethat failed - Cross-reference inputs. Use
create_workflow (action:"node_info")to verify the failing node's expected input schema - Check models. Use
list_local_modelsto verify all referenced model files exist
Out of Memory (OOM)
Error Pattern
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate X MiB.
GPU 0 has a total capacity of 24.00 GiB of which X MiB is free.
Or:
RuntimeError: CUDA error: out of memory
Root Cause
The GPU does not have enough VRAM to hold the model weights, intermediate tensors, and latent images at the same time. Common triggers:
- High resolution images (2048x2048+)
- Multiple models loaded at the same time
- FP32 precision models on limited VRAM
- Video generation (LTXV, AnimateDiff) with many frames
- Large batch sizes
Fixes (in order of preference)
- Reduce resolution. Drop to the model's native resolution (512 for SD 1.5, 1024 for SDXL/Flux)
- Use FP8/FP16 quantized models. FP8 Flux models use ~8GB vs ~24GB for FP16
- Search for FP8 variants:
download_model({ action: "search", query: "flux fp8" })or the same with"sdxl fp8"
- Search for FP8 variants:
- Launch flags (the VRAM ladder). Offload via ComfyUI CLI flags:
--lowvramoffloads text encoders / model parts to CPU--novramis extreme offload, the go-to for long video (LTX 2 / WAN) OOM--cache-nonecaches nothing (lowest RAM/VRAM); combine with--novram--reserve-vram Nreserves N GB so the GPU stops spilling into slow shared VRAM (Windows); typical2to4--disable-smart-memoryforces offload to RAM when a run gets stuck or OOMs intermittently- Full matrix and recipes:
comfyui-launch-flags
- Free VRAM between generations. ComfyUI should auto-manage, but restarting clears leaked memory
- Use tiled VAE decoding. For high-resolution images, tile the VAE decode step
- Node:
VAEDecodeTiledinstead ofVAEDecode - Breaks the image into tiles, decodes each separately, and stitches them together
- Node:
- Reduce batch size. Set batch_size to 1 in
EmptyLatentImage - Avoid multiple models. Don't load two full checkpoints at the same time; use one checkpoint and LoRAs instead
- For LTXV/video: always use FP8 quantized video models on 24GB cards
VRAM Estimates
| Model | FP32 | FP16 | FP8 |
|---|---|---|---|
| SD 1.5 | ~4GB | ~2GB | ~1GB |
| SDXL | ~12GB | ~6GB | ~3GB |
| Flux Dev | ~48GB | ~24GB | ~12GB |
| Flux Schnell | ~48GB | ~24GB | ~12GB |
| LTXV | ~20GB+ | ~10GB+ | ~6GB |
Launch Flags — VRAM / Cache / Attention / Precision
ComfyUI's startup flags tune the speed↔VRAM tradeoff. Match them to the detected
GPU (the panel orchestrator reports VRAM/GPU/torch/sage in its env block; pick the
tier from there). Set them on the process that launches ComfyUI (or the
--panel-orchestrator / connect command's ComfyUI, not the agent).
VRAM mode (pick ONE by card size)
| Flag | Card | Behavior |
|---|---|---|
--gpu-only | 16GB+ | Everything (CLIP/VAE/UNet) stays on GPU — fastest, max VRAM |
--highvram | 12–16GB | Models stay resident in GPU after use, no CPU offload |
--normalvram | 8–12GB | Default balance — unload to CPU RAM when idle |
--lowvram | 6–8GB | Split the UNet, aggressive CPU offload — slower |
--novram | 4–6GB | Extreme split/offload — for OOM even on lowvram, or long videos |
--cpu | <4GB / no GPU | CPU only (very slow) |
--reserve-vram N (GB) leaves headroom for the OS and other apps. Bump it if you
OOM intermittently mid-run (VAE decode / audio round-trips spike).
Cache (RAM vs re-run speed)
| Flag | Effect |
|---|---|
--cache-classic | Default aggressive caching (fastest re-runs, most RAM) |
--cache-lru N | Keep the last N node results (bounded RAM) |
--cache-ram N | Cap cache to N GB of headroom |
--cache-none | No caching — minimal RAM, re-runs every node |
Attention (speed vs compatibility)
| Flag | Notes |
|---|---|
--use-sage-attention | Recommended — fast + efficient (needs SageAttention + Triton; see triton-sageattention) |
--use-flash-attention | Very fast on supported GPUs |
--use-pytorch-cross-attention | PyTorch 2.x native — best compatibility |
--use-split-cross-attention | Lower VRAM, slower |
--use-quad-cross-attention | Sub-quadratic optimization |
| (omit) | Auto-selects xFormers if available |
Precision (UNet)
| Flag | Effect |
|---|---|
--fp16-unet | Half precision, ~50% VRAM |
--bf16-unet | BFloat16, good balance (newer GPUs) |
--fp8_e4m3fn-unet | 8-bit float, max savings (newest GPUs) |
Typical recipes:
- RTX 4090/5090 (24 to 32GB):
--gpu-only --use-sage-attention --cache-classic - 12 to 16GB:
--highvram --use-sage-attention(or--fp8_e4m3fn-unetfor big models) - 8GB:
--normalvram --use-sage-attention --cache-lru 20 - 6GB:
--lowvram --use-split-cross-attention --cache-none - OOM on long video:
--novram --reserve-vram 2
Device Mismatch
Error Pattern
RuntimeError: Expected all tensors to be on the same device, but found at least
two devices, cuda:0 and cpu!
Root Cause
A tensor on the CPU is combined with a tensor on the GPU. This usually happens when:
- A custom node doesn't move tensors to the correct device
- Model offloading placed parts of the model on CPU
- A node produces CPU tensors while downstream expects GPU tensors
Fixes
- Check if the error occurs with a specific custom node. Update or replace that node
- If using
--lowvramor--cpu, some nodes may not support CPU offloading - Restart ComfyUI to reset device state
- Check if a custom node has a newer version that fixes device handling
Missing Nodes
Error Pattern
Cannot find node class 'NodeClassName'
Or in the execution response:
"error": {"type": "node_not_found", "message": "Cannot find node class 'X'"}
Root Cause
The workflow references a node type that is not installed. This happens when:
- A custom node pack is not installed
- A custom node pack is installed but failed to load (import error)
- The node was renamed or removed in a pack update
Fixes
- Search for the node pack:
search_custom_nodes(action="search", query="NodeClassName") - Install via ComfyUI Manager or the registry
- Check logs for import errors:
Import errors often reveal missing Python dependenciesget_system_stats (action:"logs")(keyword="import") get_system_stats (action:"logs")(keyword="error") - Install missing Python dependencies. If the custom node requires a pip package:
pip install missing-package - Restart ComfyUI after installing any custom node. Nodes are loaded at startup
NaN Tensor Errors
Error Pattern
RuntimeError: Input contains NaN
Or images come out as solid gray/noise with NaN warnings in logs.
Root Cause
Numerical instability during the diffusion process. Common triggers:
- CFG scale too high. Values above 15-20 can cause numerical overflow
- Corrupted model weights. Damaged download or incompatible merge
- FP16 overflow. Some operations overflow at half precision
- Incompatible LoRA. A LoRA trained for a different base model
Fixes
- Lower CFG. Try CFG 7.0 for SD 1.5/SDXL, 1.0 for Flux
- Use FP32 VAE. Some VAEs produce NaN in FP16. Switch to
vae-ft-mse-840000-ema-pruned.safetensors(FP32) - Remove LoRAs. Test without LoRAs to isolate the cause
- Re-download the model. Hash verification can detect corrupted files
- Check LoRA compatibility. The LoRA must match the base model family
Dtype Mismatches
Error Pattern
RuntimeError: expected scalar type Float but found Half
Or:
RuntimeError: expected scalar type Half but found Float
Or:
RuntimeError: Input type (float) and bias type (c10::Half) should be the same
Root Cause
A model component expects one precision (FP32/FP16) but receives another. Most common with:
- VAE precision mismatch (FP16 model + FP32 VAE or vice versa)
- Mixed-precision LoRAs
- Custom nodes that force a specific dtype
Fixes
- Use a separate VAE. Load an explicit FP32 VAE instead of the checkpoint's built-in VAE
- Node:
VAELoaderwithvae-ft-mse-840000-ema-pruned.safetensors
- Node:
- Match precision. If the model is FP16, use FP16-compatible nodes throughout
- Force FP32 VAE decode. Some node packs offer
VAEDecodeFP32nodes - Check ComfyUI settings. The
--force-fp32flag forces everything to FP32 (uses more VRAM)
CLIP Token Overflow
Error Pattern
No explicit error. The prompt is truncated at 77 tokens without warning, and details mentioned late in the prompt are ignored.
Symptoms
- Later parts of long prompts have no effect on the image
- Adding more descriptive text doesn't change the output
- Removing early tokens suddenly makes later tokens work
Fixes
- Use a BREAK token. Split the prompt at natural boundaries:
subject description, pose, clothing, setting BREAK lighting, style, quality, camera angle - Use CLIPTextEncodeSDXL. SDXL's dual-CLIP processes two 77-token chunks
- Prioritize important tokens. Put the most important descriptors first
- Use fewer filler words. Remove articles and prepositions where possible
- Use embeddings. Condense complex concepts into single tokens with textual inversions
Black Images
Error Pattern
No error in the execution. The workflow "succeeds" but produces completely black or near-black images.
Root Causes and Fixes
| Cause | Diagnosis | Fix |
|---|---|---|
denoise = 0 | Check KSampler inputs | Set denoise to 1.0 for txt2img, 0.5-0.8 for img2img |
cfg = 0 | Check KSampler inputs | Set CFG to 7.0 (SD 1.5), 1.0 (Flux) |
steps = 0 | Check KSampler inputs | Set steps to 20+ (standard) or 4+ (turbo) |
| Wrong VAE | VAE doesn't match model | Use the correct VAE for the model family |
| Empty prompt | CLIPTextEncode has empty text | Add a text prompt |
| Wrong scheduler | Incompatible scheduler/sampler combo | Try "normal" scheduler with "euler" sampler |
| Seed collision | Extremely rare | Change the seed value |
| FP16 VAE overflow | VAE decode produces black | Use FP32 VAE or VAEDecodeTiled |
Quick Diagnostic Checklist
- Check
denoise> 0 (should be 1.0 for txt2img) - Check
cfg> 0 (should be 7.0 for SD 1.5, 1.0 for Flux) - Check
steps> 0 (should be 20 for standard, 4 for turbo) - Verify the positive prompt is not empty
- Try a different seed
- Try a known-working sampler/scheduler combo:
euler+normal
Connection Type Errors
Error Pattern
Output type 'IMAGE' doesn't match input type 'LATENT'
Or:
Required input 'model' of type 'MODEL' but got connection of type 'CLIP'
Root Cause
Connecting the wrong output slot of a node to an incompatible input. Often caused by using the wrong output index.
Fixes
- Check output indices. Use
create_workflow (action:"node_info")to verify the exact output orderCheckpointLoaderSimpleoutputs: 0=MODEL, 1=CLIP, 2=VAE- Getting index wrong:
["1", 0]gives MODEL,["1", 1]gives CLIP
- Verify connection format.
["nodeId", outputIndex], where node ID is a string and index is an integer - Check data type flow. The pipeline must follow the correct type chain:
MODEL → KSampler CLIP → CLIPTextEncode → CONDITIONING → KSampler LATENT → KSampler → LATENT → VAEDecode → IMAGE VAE → VAEDecode, VAEEncode
Model Loading Errors
Error Pattern
FileNotFoundError: [Errno 2] No such file or directory: 'models/checkpoints/model.safetensors'
Or:
SafetensorError: Error reading file: invalid header
Or:
RuntimeError: PytorchStreamReader failed reading zip archive
Root Causes
- File not found. Model file doesn't exist at the referenced path
- Corrupted download. Incomplete or damaged file
- Wrong format. File is not a valid safetensors/pickle/checkpoint format
Fixes
- Verify the model exists:
list_local_models({ action: "list", model_type: "checkpoints" }) - Check the exact filename. Model names in workflows must match the filename exactly (case-sensitive)
- Re-download. If hash mismatch or corruption:
download_model({ action: "download", url: "...", target_subfolder: "checkpoints" }) - Check file size. A 1KB safetensors file is corrupted; re-download
- Verify subfolder. Models must be in the correct subfolder (
checkpoints/,loras/,vae/, etc.)
Torch / CUDA Version Errors
Error Pattern
RuntimeError: CUDA error: no kernel image is available for execution on the device
Or:
ImportError: cannot import name 'xxx' from 'torch'
Or:
AssertionError: Torch not compiled with CUDA enabled
Root Cause
PyTorch and CUDA version incompatibility, usually after:
- Updating PyTorch without matching CUDA toolkit
- Installing a custom node that downgrades/changes PyTorch
- Using pip install that pulls a CPU-only PyTorch
Fixes
- Check current versions:
get_system_stats() # Shows PyTorch version and CUDA version - Verify CUDA availability. In Python:
torch.cuda.is_available() - Reinstall PyTorch with CUDA. Visit pytorch.org for the correct install command matching your CUDA version
- Pin PyTorch version. After fixing, avoid running
pip installcommands that might change PyTorch - Use ComfyUI's bundled venv. ComfyUI Desktop ships with a pre-configured Python environment
ComfyUI Desktop vs CLI Differences
Key Differences
| Aspect | ComfyUI Desktop | ComfyUI CLI |
|---|---|---|
| Default port | 8000 | 8188 |
| Python | Embedded (bundled) | System/venv Python |
| Install location | AppData/Local/Programs/ComfyUI/ | Wherever you cloned it |
| Custom nodes | Documents/ComfyUI/custom_nodes/ | ./custom_nodes/ in repo |
| Models | Documents/ComfyUI/models/ | ./models/ in repo |
| Config | extra_model_paths.yaml for shared paths | Same |
| Updates | Auto-updater in the app | git pull |
Common Issues
- Wrong port. MCP tools default to 8188; if using Desktop, configure for port 8000
- Path confusion. Desktop separates user data from application files
- Custom node pip installs. Desktop's embedded Python may not be on PATH; install within the venv
Error-Specific Debugging Commands
Workflow Failed — Get Details
get_history(action="list") # Most recent execution
get_history(action="list", prompt_id="abc-123") # Specific execution
get_history(action="diagnose") # Why the last run failed
The response includes:
status.status_str: "success" or "error"status.messages: Timestamped execution messagesoutputs: Node outputs (images, etc.)- Error traceback for failed nodes
Check Server Health
get_system_stats() # GPU info, VRAM, Python/PyTorch versions
queue(action="list") # Running and pending jobs
get_system_stats (action:"logs")(max_lines=50, keyword="error") # Recent error logs
Verify Node Availability
create_workflow(action="node_info", node_type="KSampler") # Check specific node
create_workflow(action="node_info", node_type="ControlNetApply") # Verify custom nodes loaded
Verify Models
list_local_models({ action: "list", model_type: "checkpoints" }) # Installed checkpoints
list_local_models({ action: "list", model_type: "loras" }) # Installed LoRAs
list_local_models({ action: "list", model_type: "controlnet" }) # Installed ControlNets
Quick Reference: Error to Fix
| Error Message (partial) | Most Likely Fix |
|---|---|
CUDA out of memory | Reduce resolution, use FP8 model; VRAM ladder --lowvram → --novram --cache-none → --reserve-vram N (launch flags) |
Expected all tensors on same device | Update custom node, restart ComfyUI |
Cannot find node class | Install the node pack, restart ComfyUI |
Input contains NaN | Lower CFG, use FP32 VAE, remove LoRAs |
expected scalar type Float but found Half | Use FP32 VAE, or --force-fp32 |
No such file or directory (model) | Check filename, re-download model |
invalid header (safetensors) | Re-download — file is corrupted |
CUDA error: no kernel image | Reinstall PyTorch with matching CUDA version |
| Black images, no error | Check denoise > 0, cfg > 0, steps > 0, prompt not empty |
| Image looks garbled/noisy | Wrong model+VAE combo, wrong sampler settings |
Connection refused on port 8188 | ComfyUI not running, or using Desktop (port 8000) |
Prompt outputs failed validation | Node inputs don't match schema — check create_workflow (action:"node_info") |
Sources
- Official: none found as a vendor error catalog. Launch-flag names cross-check against
comfyui-launch-flags(upstream cli_args.py). - Empirical: error→fix table from observed ComfyUI failures.
Frequently asked questions
What to verify before installation and use
What does the troubleshooting source document cover?
Render completes but looks WRONG (artifacts, wrong subject/pose/color, a ControlNet/mask/LoRA not taking, a refiner degrading it)? That's not an error. Use the debug-render skill (listpacks with action: "skillread", name: "debug-render") to localize the bad stage with run-to-nod…
How do I install troubleshooting?
The source record exposes this install command: npx skills add https://github.com/artokun/comfyui-mcp --skill "plugin/skills/troubleshooting". Inspect the command and pinned source before running it.
Which permission-related actions were detected?
Static rules flagged network in the source; the page lists the matching lines and excerpts.