Feedback
AI Ad Video Example
Loading...
minimax h3 video model
Turn a text prompt into a 15-second 2K clip with synchronized audio through the minimax h3 video model — one multimodal engine for text, images, video, and sound.
All Tools
Discover our comprehensive AI-powered animation toolkit

Seedance2.0
The Future of AI Video Is Here.

Free AI Image
Truly Free AI Image Generator

Happy Horse 1.1

Gemini Omni
Gemini Omni Video Generator

Seedance 2.1
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
Why Teams Prefer the minimax h3 video model
Built by MiniMax and offered as a Day 0 partner on fal.ai, the minimax h3 video model is an open-weight omni-modal engine. It accepts text, images, video, and audio in a single prompt, outputs 2K footage with native stereo sound up to 15 seconds, and enables targeted edits, crisp text and UI rendering, and up to 12 reference inputs per call.
- A Single Context for All Media TypesFeed it up to 9 images, 3 video clips, and 3 audio tracks at once — the minimax h3 video model fuses subject identity, acting, camera movement, and audio into one consistent output.
- Built-In Stereo SoundEach minimax h3 video model result includes original music, speech, foley, and room tone matched to the cut — plus voice transfer and cloning from supplied audio references.
- Pinpoint Local EditsSwap a product, alter signage, replace dialogue, or turn daylight into night — the minimax h3 video model changes only the specified area while the rest of the frame remains untouched.
Start with the minimax h3 video model API in 3 Steps
Generate 2K video with synced audio using the minimax h3 video model API in just three steps.
Key Features of the minimax h3 video model
With three API endpoints, one shared multimodal context, native stereo sound, targeted editing, crisp text rendering, and usage-based pricing, the minimax h3 video model forms a full 2K video production workflow on fal.ai.
Three Endpoints for Every Workflow
The minimax h3 video model provides text-to-video, image-to-video with first/last-frame control, and reference-to-video endpoints for all production styles.
Up to 12 Inputs in One Request
Mix 9 images, 3 video clips, and 3 audio tracks — the minimax h3 video model extracts subject identity, acting, camera movement, framing, and cut rhythm from those references.
Sharp Text and Interface Generation
Create crisp captions, lower thirds, end cards, and brand logos, and bring real UI to life — websites, game menus, HUDs, and kinetic typography using the minimax h3 video model.
Massive Prompt Capacity
Include an entire storyboard in one call — the minimax h3 video model accepts prompts up to 7,000 characters for complete scene direction.
High-Resolution 2K Output at 24fps
Produce 2K footage with a 1440-pixel short edge, up to 15 seconds at 24 frames per second, and six aspect ratios plus an adaptive mode via the minimax h3 video model.
Usage-Based, No Subscription Required
Access the minimax h3 video model through serverless, pay-as-you-go pricing — no minimum spend, no subscription, and full commercial rights to generated work.
Common Questions About the minimax h3 video model
Answers to frequent questions about running the minimax h3 video model on fal.ai.
What exactly does the MiniMax H3 video model do?
It's an open-weight, omni-modal generation model from MiniMax, available on fal.ai as a Day 0 ecosystem partner. The same model handles text, images, video, and audio together, creating 2K clips with native stereo sound up to 15 seconds.
Which API endpoints can I call with the MiniMax H3 video model?
The minimax h3 video model exposes three API endpoints: text-to-video, image-to-video with optional first/last-frame control, and reference-to-video for locking subjects, style, motion, camera direction, and voice from source materials.
What resolutions and durations are available?
The minimax h3 video model generates 2K video (1440-pixel short edge) at 24fps, runs from 5 to 15 seconds, and supports aspect ratios of 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, along with adaptive mode.
Does the MiniMax H3 video model produce audio?
Yes. Each minimax h3 video model output includes native stereo audio — original music, dialogue, foley, and ambience aligned to the cut — plus voice transfer and cloning from reference recordings.
How many reference files can I provide?
You can supply up to 12 files in total: 9 reference images, 3 reference video clips of 2-15 seconds each, and 3 reference audio tracks of 2-15 seconds each. For the minimax h3 video model, audio needs at least one accompanying image or video.
Can I use generated content in commercial projects?
Absolutely. Content produced through the fal.ai API with the minimax h3 video model can be used commercially, subject to fal.ai's terms of service.
Start Creating 2K Videos with the minimax h3 video model
In one request, the minimax h3 video model lets you make 2K video with built-in stereo audio — combine multimodal references, precise edits, and pay-per-use API pricing on fal.ai.
