Start Generating with comfyui minimax h3
Describe your scene, and the comfyui minimax h3 workflow will generate video with synced audio
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Use the comfyui minimax h3 workflow in ComfyUI to produce open-weight videos with native stereo audio — from text, image, or reference inputs, up to 2K at 24fps.

All Tools

Discover our comprehensive AI-powered animation toolkit

Why the comfyui minimax h3 Workflow Is a Game-Changer

The comfyui minimax h3 workflow integrates MiniMax's omni-modal generation model into ComfyUI with open weights. It understands text, images, video, and audio together, generating clips with native stereo audio in a single pass. Output reaches up to 2K resolution at 24fps for 15 seconds, with node-level control over all parameters.

  • Synced Stereo Audio in Every Clip
    Speech, effects, and music are generated alongside the footage in a single MP4, perfectly synced thanks to the comfyui minimax h3 pipeline.
  • Full Control with Open Weights
    Run the comfyui minimax h3 model on your own hardware, adjusting resolution, duration, and all diffusion settings freely without API restrictions.
  • Generate from Any Mixed References
    Feed in text, images, clips, or audio references simultaneously to lock a character, style, motion, camera angle, or voice through the comfyui minimax h3 nodes.

A Simple 3-Step Guide to the comfyui minimax h3 Workflow

Follow these three steps to start producing open-weight videos with native sound using the comfyui minimax h3 workflow.

Key Capabilities of the comfyui minimax h3 Workflow

With three built-in ComfyUI templates, open-weight multimodal generation, native stereo audio, reference-driven control, and optional Sage Attention acceleration, the comfyui minimax h3 workflow gives you a complete local video production stack.

Three Built-in Workflow Templates

The comfyui minimax h3 template pack includes text-to-video, image-to-video, and reference-to-video examples, ready to run for each generation mode.

All-in-One Multimodal Understanding

The comfyui minimax h3 model processes text, images, video, and audio together in a single context, blending all reference types into one generation.

Generate From Rich References

Lock in a character, style, motion, camera movement, or voice from references — up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.

Crisp Text and Brand Elements

Spelled-out words and brand graphics appear clearly with the comfyui minimax h3 model, following instructions that describe reference relationships in everyday language.

Faster Generation with Sage Attention

Add the Patch Sage Attention KJ node to the comfyui minimax h3 workflow to roughly double speed with barely any quality drop.

Flexible Resolution and Duration Grid

The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapping to the model's 32-pixel grid and 17-frame-per-block duration at 24fps.

FAQ

Frequently Asked Questions about comfyui minimax h3

Answers to common queries about using the comfyui minimax h3 workflow and MiniMax H3 within ComfyUI.

1

What exactly is the comfyui minimax h3 workflow?

It’s ComfyUI’s official integration of MiniMax H3, an omni-modal generation model released with open weights. This workflow creates video with native stereo audio from text, images, video, and audio references in one forward pass.

2

What resolution and frame rate can I expect?

The comfyui minimax h3 workflow delivers up to 2K at 24fps for around 15 seconds. Its native canvas uses a 768px short edge, limited to 768x1344 pixels and rounded to multiples of 32.

3

What generation modes come with the workflow?

The comfyui minimax h3 template library includes three examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) for locking character, style, motion, camera, or voice.

4

Can it really produce audio along with video?

Absolutely — the comfyui minimax h3 model generates native stereo audio, including voice, effects, and music, modeled with the video in one pass and synced into a single MP4.

5

What do I need to begin using it?

Upgrade ComfyUI to 0.30.0+, go to Template Library > Video, select a comfyui minimax h3 workflow, then follow the pop-up to fetch models from the Hugging Face Comfy-Org/MiniMax-H3 repository.

6

Are there ways to make generation faster?

Yes — install SageAttention and KJNodes, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double rendering speed.

Begin Creating Today with the comfyui minimax h3 Workflow

Run MiniMax H3 on your own ComfyUI setup with native stereo audio, open weights, and total parameter control — T2V, I2V, and R2V workflows are ready whenever you are.