Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Use the comfyui minimax h3 workflow in ComfyUI to produce open-weight videos with native stereo audio — from text, image, or reference inputs, up to 2K at 24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit

Seedance2.0
The Future of AI Video Is Here.

Free AI Image
Truly Free AI Image Generator

Happy Horse 1.1

Gemini Omni
Gemini Omni Video Generator

Seedance 2.1
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
Why the comfyui minimax h3 Workflow Is a Game-Changer
The comfyui minimax h3 workflow integrates MiniMax's omni-modal generation model into ComfyUI with open weights. It understands text, images, video, and audio together, generating clips with native stereo audio in a single pass. Output reaches up to 2K resolution at 24fps for 15 seconds, with node-level control over all parameters.
- Synced Stereo Audio in Every ClipSpeech, effects, and music are generated alongside the footage in a single MP4, perfectly synced thanks to the comfyui minimax h3 pipeline.
- Full Control with Open WeightsRun the comfyui minimax h3 model on your own hardware, adjusting resolution, duration, and all diffusion settings freely without API restrictions.
- Generate from Any Mixed ReferencesFeed in text, images, clips, or audio references simultaneously to lock a character, style, motion, camera angle, or voice through the comfyui minimax h3 nodes.
A Simple 3-Step Guide to the comfyui minimax h3 Workflow
Follow these three steps to start producing open-weight videos with native sound using the comfyui minimax h3 workflow.
Key Capabilities of the comfyui minimax h3 Workflow
With three built-in ComfyUI templates, open-weight multimodal generation, native stereo audio, reference-driven control, and optional Sage Attention acceleration, the comfyui minimax h3 workflow gives you a complete local video production stack.
Three Built-in Workflow Templates
The comfyui minimax h3 template pack includes text-to-video, image-to-video, and reference-to-video examples, ready to run for each generation mode.
All-in-One Multimodal Understanding
The comfyui minimax h3 model processes text, images, video, and audio together in a single context, blending all reference types into one generation.
Generate From Rich References
Lock in a character, style, motion, camera movement, or voice from references — up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.
Crisp Text and Brand Elements
Spelled-out words and brand graphics appear clearly with the comfyui minimax h3 model, following instructions that describe reference relationships in everyday language.
Faster Generation with Sage Attention
Add the Patch Sage Attention KJ node to the comfyui minimax h3 workflow to roughly double speed with barely any quality drop.
Flexible Resolution and Duration Grid
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapping to the model's 32-pixel grid and 17-frame-per-block duration at 24fps.
Frequently Asked Questions about comfyui minimax h3
Answers to common queries about using the comfyui minimax h3 workflow and MiniMax H3 within ComfyUI.
What exactly is the comfyui minimax h3 workflow?
It’s ComfyUI’s official integration of MiniMax H3, an omni-modal generation model released with open weights. This workflow creates video with native stereo audio from text, images, video, and audio references in one forward pass.
What resolution and frame rate can I expect?
The comfyui minimax h3 workflow delivers up to 2K at 24fps for around 15 seconds. Its native canvas uses a 768px short edge, limited to 768x1344 pixels and rounded to multiples of 32.
What generation modes come with the workflow?
The comfyui minimax h3 template library includes three examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) for locking character, style, motion, camera, or voice.
Can it really produce audio along with video?
Absolutely — the comfyui minimax h3 model generates native stereo audio, including voice, effects, and music, modeled with the video in one pass and synced into a single MP4.
What do I need to begin using it?
Upgrade ComfyUI to 0.30.0+, go to Template Library > Video, select a comfyui minimax h3 workflow, then follow the pop-up to fetch models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Are there ways to make generation faster?
Yes — install SageAttention and KJNodes, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double rendering speed.
Begin Creating Today with the comfyui minimax h3 Workflow
Run MiniMax H3 on your own ComfyUI setup with native stereo audio, open weights, and total parameter control — T2V, I2V, and R2V workflows are ready whenever you are.
