Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Harness the comfyui minimax h3 workflow to generate open-weight video with sound: provide a prompt, a photo, or a clip and receive up to 2K footage at 24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Why Choose the comfyui minimax h3 Pipeline for Video Generation
MiniMax H3 arrives in ComfyUI as an open-weight, omni-modal model, and the comfyui minimax h3 workflow uses it to process text, stills, footage, and sound in one unified context. The output carries two-channel audio — speech, effects, and music — generated simultaneously with the visuals. Clips scale up to 2K at 24fps for about 15 seconds, with every node-level parameter exposed for fine-tuning.
- Audio Rendered in the Same PassDialogue, sound effects, and background music are created together with the footage and delivered in one MP4 — frame-matched to the action, so no separate editing tool is required.
- Full Local ControlSince the model weights are open, you can run the comfyui minimax h3 workflow entirely on your own hardware. Resolution, duration, and every diffusion parameter are fully adjustable with no API dependency.
- Blend Multiple Input TypesFeed text alongside images, footage, or audio references into the nodes, and the output anchors character identity, visual style, motion, camera work, or voice — all in a single run.
Three Steps to Run the comfyui minimax h3 Workflow
Load a preset, connect open-weight checkpoints, and generate sound-ready video in three steps with the comfyui minimax h3 workflow.
Core Capabilities of the comfyui minimax h3 Workflow
The comfyui minimax h3 workflow combines three ready-to-use ComfyUI templates, open-weight multimodal generation, audio baked into the render, reference-driven control, and an optional Sage Attention speed-up — a full-fledged local video toolset.
Three Native Workflow Templates
The comfyui minimax h3 template library ships with text-to-video, image-to-video, and reference-to-video examples, each covering one generation mode out of the box.
Omni-Modal Context
The comfyui minimax h3 model understands text, images, video, and audio together in a single context, combining all reference types in one generation.
Reference-Driven Generation
Lock a character's identity, a style, a motion, a camera move, or a voice from reference materials — up to 9 images, 3 videos, and 3 audio clips via the comfyui minimax h3 R2V node.
Accurate Text & Brand Rendering
Spelled-out text and brand elements render cleanly with the comfyui minimax h3 model, with instruction following that describes reference relationships in natural language.
Sage Attention Speedup
Roughly double generation speed with minimal quality loss by adding the Patch Sage Attention KJ node to the comfyui minimax h3 workflow.
Resolution & Duration Grid
The comfyui minimax h3 Resolution Selector computes width and height from aspect ratio and megapixels, snapped to the model's 32-multiple grid and 17-frame-per-block duration at 24fps.
Frequently Asked Questions for the comfyui minimax h3 Workflow
Everything you need to know about running MiniMax H3 inside ComfyUI, from model setup to audio generation.
What exactly is the comfyui minimax h3 workflow?
It is ComfyUI's native integration of MiniMax H3, MiniMax's general-purpose omni-modal generation model released as open weights. The workflow generates video with native stereo audio from text, images, video, and audio references in a single forward pass.
What output quality does it support?
The comfyui minimax h3 workflow outputs up to 2K resolution at 24fps for about 15 seconds. Its native canvas uses a 768px short edge, capped at 768x1344 pixels and rounded to a multiple of 32.
Which generation modes are included?
The comfyui minimax h3 template library ships with three examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks in character, style, motion, camera, or voice.
Does it generate audio?
Yes — the comfyui minimax h3 model produces native stereo audio including voice, sound effects, and music, modeled together with the video in one pass and synced in a single MP4 file.
How do I get started?
Update ComfyUI to version 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 workflow, and follow the pop-up to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Can I speed up generation?
Yes — install SageAttention and the KJNodes custom nodes, then add a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double generation speed.
Begin Building Videos with the comfyui minimax h3 Nodes Today
Produce sound-integrated films entirely on your own machine using the comfyui minimax h3 workflow — open weights, adjustable parameters, and T2V, I2V, and R2V templates all ready.
