Create Videos with comfyui minimax h3
Use the comfyui minimax h3 pipeline to produce clips with audio built in — start with a prompt below.
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Run comfyui minimax h3 in ComfyUI to create clips with synced stereo audio — from text, photos, or footage, up to 2K at 24fps.

All Tools

Discover our comprehensive AI-powered animation toolkit

Key Advantages of Running comfyui minimax h3 in ComfyUI

When you run comfyui minimax h3 inside ComfyUI, you're using MiniMax's open-weight omni-modal model. It processes text, visuals, video, and audio in one shared context, then renders footage with sound — dialogue, effects, and score — baked into a single pass. Clips can reach 2K at 24fps for roughly 15 seconds, and every node is fully adjustable.

  • Built-In Stereo Audio
    Audio is synthesized together with the picture — voices, effects, and music are merged into one MP4, all generated in a single run of the comfyui minimax h3 workflow.
  • Full Local Control
    Because comfyui minimax h3 is open-weight, you can run it on your own machine and adjust resolution, duration, and diffusion settings without hitting API limits.
  • Flexible Reference Mixing
    You can feed text, stills, footage, and audio hints into one generation, and the comfyui minimax h3 nodes lock in a character, style, movement, camera motion, or voice.

Operating the comfyui minimax h3 Workflow in Three Simple Steps

Follow this setup to begin creating local, open-weight videos with synchronized sound through comfyui minimax h3.

What the comfyui minimax h3 Pipeline Brings to ComfyUI

This pipeline bundles three ready-to-run ComfyUI templates, open-weight multimodal generation, built-in stereo audio, reference-based control, and optional Sage Attention acceleration. Together, comfyui minimax h3 gives you a full local video production toolkit.

A Trio of Built-In Workflow Templates

The comfyui minimax h3 set includes text-to-video, image-to-video, and reference-to-video examples, each covering a different generation mode right out of the box.

Unified Multimodal Understanding

With comfyui minimax h3, text, images, video, and audio are all parsed together in the same context, letting you combine every reference type in a single shot.

Guided Creation from References

You can lock in a character's look, an art style, a motion, a camera movement, or a voice from up to 9 images, 3 clips, and 3 audio files via the comfyui minimax h3 reference node.

Sharp Text and Logo Rendering

The comfyui minimax h3 model renders spelled-out words and logos cleanly, and it follows instructions that describe relationships between references in natural language.

Faster Output with Sage Attention

By dropping the Patch Sage Attention KJ node into the comfyui minimax h3 workflow, you can roughly double generation speed with only a minimal quality hit.

Granular Size and Length Control

The comfyui minimax h3 resolution selector calculates width and height from aspect ratio and megapixels, snapped to the model's 32-pixel grid and 17-frame-per-block duration at 24fps.

FAQ

Common Queries About the comfyui minimax h3 Workflow

Everything you need to know about using the MiniMax H3 model with ComfyUI via comfyui minimax h3 templates.

1

How does the comfyui minimax h3 workflow fit into ComfyUI?

This is ComfyUI's official setup for MiniMax H3, an open-weight, all-in-one multimodal model from MiniMax. With comfyui minimax h3, you can produce videos with synchronized sound from text, stills, footage, and audio samples in one forward pass.

2

What resolution and frame rate can I expect?

This workflow delivers clips up to 2K at 24fps, lasting around 15 seconds. The default canvas has a 768px short edge, limited to 768x1344 and rounded to a multiple of 32.

3

What generation modes does the template cover?

The template collection offers three ready workflows: text-to-video (T2V), image-to-video (I2V) with optional start/end frame guidance, and reference-to-video (R2V) which fixes a character, style, movement, camera, or voice.

4

Is audio actually generated?

Yes, it does. The comfyui minimax h3 model creates stereo audio—voices, effects, and music—together with the visuals in a single pass, and the result is one MP4 with perfectly synced sound.

5

What's the quickest way to begin?

Simply update ComfyUI to 0.30.0 or newer, navigate to Template Library > Video, select a comfyui minimax h3 workflow, and follow the prompt to grab the models from the Comfy-Org/MiniMax-H3 Hugging Face repo.

6

Is there a way to make it run faster?

Absolutely. Install SageAttention and KJNodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 pipeline. That can almost double the rendering speed.

Start Producing with comfyui minimax h3 Now

Launch comfyui minimax h3 directly in ComfyUI for local, open-weight video creation with sound included. All three pipelines — text, image, and reference — are set up and waiting for you.