ComfyUI QwenVL
ComfyUI-QwenVL custom node: Integrates the Qwen-VL series, including Qwen2.5-VL and the latest Qwen3-VL, with GGUF support for advanced multimodal AI in text generation, image understanding, and video analysis.
Install / Use
npx skills add 1038lab/ComfyUI-QwenVLInstalls into whichever agent you are using.
README
QwenVL for ComfyUI
The ComfyUI-QwenVL custom node integrates the powerful Qwen-VL series of vision-language models (LVLMs) from Alibaba Cloud, including the latest Qwen3-VL and Qwen2.5-VL, plus GGUF backends and text-only Qwen3 support. This advanced node enables seamless multimodal AI capabilities within your ComfyUI workflows, allowing for efficient text generation, image understanding, and video analysis.
📰 News & Updates
-
2026/02/08: v2.1.1 Fixed compatibility for Transformers 4.x and 5.x [Update]
-
2026/02/05: v2.1.0 Added SageAttention support with per-GPU architecture optimization, improved FP8 model handling, and automatic attention mode selection. [Update]
- SageAttention Support: New attention mode with per-GPU optimized kernels (SM80, SM89, SM90, SM120)
- Improved FP8 Handling: Better support for pre-quantized FP8 models with automatic SDPA fallback
- Smart Attention Selection: Auto mode now tries Sage → Flash → SDPA for optimal performance
- Progress Bar: Added ComfyUI progress bar for model loading and generation stages
- Better Memory Management: Improved cache clearing when changing attention modes or quantization
-
2025/12/22: v2.0.0 Added GGUF supported nodes and Prompt Enhancer nodes. [Update]
[!IMPORTANT]
Install llama-cpp-python before running GGUF nodes instruction
- 2025/11/10: v1.1.0 Runtime overhaul with attention-mode selector, flash-attn auto detection, smarter caching, and quantization/torch.compile controls in both nodes. [Update]
- 2025/10/31: v1.0.4 Custom Models Supported [Update]
- 2025/10/22: v1.0.3 Models list updated [Update]
- 2025/10/17: v1.0.0 Initial Release
- Support for Qwen3-VL and Qwen2.5-VL series models.
- Automatic model downloading from Hugging Face.
- On-the-fly quantization (4-bit, 8-bit, FP16).
- Preset and Custom Prompt system for flexible and easy use.
- Includes both a standard and an advanced node for users of all levels.
- Hardware-aware safeguards for FP8 model compatibility.
- Image and Video (frame sequence) input support.
- "Keep Model Loaded" option for improved performance on sequential runs.
- Seed parameter for reproducible generation.
✨ Features
- Standard & Advanced Nodes: Includes a simple QwenVL node for quick use and a QwenVL (Advanced) node with fine-grained control over generation.
- Prompt Enhancers: Dedicated text-only prompt enhancers for both HF and GGUF backends.
- Preset & Custom Prompts: Choose from a list of convenient preset prompts or write your own for full control.
- Multi-Model Support: Easily switch between various official Qwen-VL models.
- Automatic Model Download: Models are downloaded automatically on first use.
- Smart Quantization: Balance VRAM and performance with 4-bit, 8-bit, and FP16 options.
- Hardware-Aware: Automatically detects GPU capabilities and prevents errors with incompatible models (e.g., FP8).
- Reproducible Generation: Use the seed parameter to get consistent outputs.
- Memory Management: "Keep Model Loaded" option to retain the model in VRAM for faster processing.
- Image & Video Support: Accepts both single images and video frame sequences as input.
- Robust Error Handling: Provides clear error messages for hardware or memory issues.
- Clean Console Output: Minimal and informative console logs during operation.
- SageAttention Support: GPU-optimized attention mechanism with per-architecture kernels (Ampere, Ada, Hopper, Blackwell).
- Progress Bar: Visual feedback during model loading and generation stages.
- Intelligent Cache Management: Automatically clears VRAM when changing attention modes or quantization settings.
🚀 Installation
-
Clone this repository to your ComfyUI/custom_nodes directory:
cd ComfyUI/custom\_nodes git clone https://github.com/1038lab/ComfyUI-QwenVL.git -
Install the required dependencies:
cd ComfyUI/custom\_nodes/ComfyUI-QwenVL pip install \-r requirements.txt -
Restart ComfyUI.
Optional: SageAttention Support
For optimal performance on supported GPUs, install SageAttention:
pip install sageattention
🧭 Node Overview
Transformers (HF) Nodes
- QwenVL: Quick vision-language inference (image/video + preset/custom prompts).
- QwenVL (Advanced): Full control over sampling, device, and performance settings.
- QwenVL Prompt Enhancer: Text-only prompt enhancement (supports both Qwen3 text models and QwenVL models in text mode).
GGUF (llama.cpp) Nodes
- QwenVL (GGUF): GGUF vision-language inference.
- QwenVL (GGUF Advanced): Extended GGUF controls (context, GPU layers, etc.).
- QwenVL Prompt Enhancer (GGUF): GGUF text-only prompt enhancement.
🧩 GGUF Nodes (llama.cpp backend)
This repo includes GGUF nodes powered by llama-cpp-python (separate from the Transformers-based nodes).
- Nodes:
QwenVL (GGUF),QwenVL (GGUF Advanced),QwenVL Prompt Enhancer (GGUF) - Model folder (default):
ComfyUI/models/llm/GGUF/(configurable viagguf_models.json) - Vision requirement: install a vision-capable
llama-cpp-pythonwheel that providesQwen3VLChatHandler/Qwen25VLChatHandler
See docs/LLAMA_CPP_PYTHON_VISION_INSTALL.md
🗂️ Config Files
- HF models:
hf_models.jsonhf_vl_models: vision-language models (used by QwenVL nodes).hf_text_models: text-only models (used by Prompt Enhancer).
- GGUF models:
gguf_models.json - System prompts:
AILab_System_Prompts.json(includes both VL prompts and prompt-enhancer styles).
📥 Download Models
The models will be automatically downloaded on first use. If you prefer to download them manually, place them in the ComfyUI/models/LLM/Qwen-VL/ directory.
HF Vision Models (Qwen-VL)
| Model | Link | | :---- | :---- | | Qwen3-VL-2B-Instruct | Download | | Qwen3-VL-2B-Thinking | Download | | Qwen3-VL-2B-Instruct-FP8 | Download | | Qwen3-VL-2B-Thinking-FP8 | Download | | Qwen3-VL-4B-Instruct | Download | | Qwen3-VL-4B-Thinking | Download | | Qwen3-VL-4B-Instruct-FP8 | Download | | Qwen3-VL-4B-Thinking-FP8 | Download | | Qwen3-VL-8B-Instruct | Download | | Qwen3-VL-8B-Thinking | Download | | Qwen3-VL-8B-Instruct-FP8 | Download | | Qwen3-VL-8B-Thinking-FP8 | Download | | Qwen3-VL-32B-Instruct | Download | | Qwen3-VL-32B-Thinking | Download | | Qwen3-VL-32B-Instruct-FP8 | Download | | Qwen3-VL-32B-Thinking-FP8 | Download | | Qwen2.5-VL-3B-Instruct | Download | | Qwen2.5-VL-7B-Instruct | Download |
HF Text Models (Qwen3)
| Model | Link | | :---- | :---- | | Qwen3-0.6B | Download | | Qwen3-4B-Instruct-2507 | Download | | qwen3-4b-Z-Image-Engineer | Download |
GGUF Models (Manual Download)
| Group | Model | Repo | Alt Repo | Model Files | MMProj | | :-- | :-- | :-- | :-- | :-- | :-- | | Qwen text (GGUF) | Qwen3-4B-GGUF | Qwen/Qwen3-4B-GGUF | | Qwen3-4B-Q4_K_M.gguf, Qwen3-4B-Q5_0.gguf, Qwen3-4B-Q5_K_M.gguf, Qwen3-4B-Q6_K.gguf, Qwen3-4B-Q8_0.gguf | | | Qwen-VL (GGUF) | Qwen3-VL-4B-Instruct-GGUF | Qwen/Qwen3-VL-4B-Instruct-GGUF | | Qwen3VL-4B-Instruct-F16.gguf, Qwen3VL-4B-Instruct-Q4_K_M.gguf, Qwen3VL-4B-Instruct-Q8_0.gguf | mmproj-Qwen3VL-4B-Instruct-F16.gguf | | Qwen-VL (GGUF) | Qwen3-VL-8B-Instruct-GGUF | Qwen/Qwen3-VL-8B-Instruct-GGUF | | Qwen3VL-8B-Instruct-F16.gguf, Qwen3VL-8B-Instruct-Q4_K_M.gguf, Qwen3VL-8B-Instruct-Q8_0.gguf | mmproj-Qwen3VL-8B-Instruct-F16.gguf | | Qwen-VL (GGUF) | Qwen3-
Related Skills
gortex
1.1kHigh-performance code-intelligence engine for AI agents and IDE, supports 257 languages, multi repositories, based on graph, with access via CLI, MCP Server, and API. AI coding agents teammate - expose only needed information, cutting token usage up to 50x. 100% local.
techrogue
TechRogue – Roguelike technical quiz for engineers. Usage: /techrogue | /techrogue build | /techrogue settings
cc-switch
125.5kA cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io
cc-switch
125.6kA cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

