comfyui minimax h3 Video Studio
With the comfyui minimax h3 workflow, your videos naturally include synchronized stereo audio.
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Turn text, images, or references into open-weight video with stereo audio inside ComfyUI using the comfyui minimax h3 workflow — resolution up to 2K at 24fps.

All Tools

Discover our comprehensive AI-powered animation toolkit

Why Adopt the comfyui minimax h3 Workflow

The comfyui minimax h3 workflow embodies MiniMax's omni-modal generation model with open weights inside ComfyUI. A single context jointly processes text, images, video, and audio, then renders video plus native stereo audio — voice, effects, and music synthesized in one forward pass. Output can reach approximately 15 seconds at up to 2K 24fps, with every parameter exposed for node-level tuning.

  • Built-In Stereo Sound
    Dialogue, ambience, and music are synthesized together with the picture and packed into one MP4 — perfectly aligned through the comfyui minimax h3 processing chain.
  • Local Weight Flexibility
    Run the comfyui minimax h3 model on your own hardware, adjusting resolution, clip length, and diffusion parameters without relying on external API quotas.
  • Multi-Format Reference Handling
    Feed text, images, video, or audio references into a single generation, using comfyui minimax h3 nodes to fix a character, aesthetic, motion, camera angle, or voice.

Getting Started with the comfyui minimax h3 Workflow

Produce open-weight video with synced audio in just three steps using the comfyui minimax h3 workflow.

Capabilities of the comfyui minimax h3 Workflow

The comfyui minimax h3 workflow bundles three native ComfyUI templates, open-weight multimodal generation, native stereo audio, reference-based control, and optional Sage Attention acceleration — a complete local video production stack.

Three Ready-Made Node Templates

The comfyui minimax h3 library provides text-to-video, image-to-video, and reference-to-video templates, giving you one straightforward path for every mode.

Unified Omni-Modal Context

With the comfyui minimax h3 model, text, images, video, and audio are understood collectively in one context, allowing all reference types to be combined in one pass.

Reference-Driven Customization

Lock a character's identity, a style, movement, camera motion, or a voice from supplied materials — up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.

Precise Text and Brand Rendering

The comfyui minimax h3 model renders spelled text and brand assets clearly, following instructions that describe reference connections with natural language.

Sage Attention Acceleration

Adding the Patch Sage Attention KJ node to your comfyui minimax h3 workflow can roughly double render speed while preserving visual quality.

Resolution and Duration Grid

The comfyui minimax h3 Resolution Selector scales width and height from aspect ratio and megapixels, snapping to the model's 32-multiple grid and 17-frame-per-block duration at 24fps.

FAQ

comfyui minimax h3 — Common Questions

Frequently asked questions about running the MiniMax H3 model in ComfyUI.

1

What exactly is the comfyui minimax h3 workflow?

It is ComfyUI's built-in setup for MiniMax H3, MiniMax's omni-modal generation model released with open weights. The workflow creates video with native stereo audio from text, images, video, and audio references in a single forward pass.

2

What output quality does this workflow support?

The comfyui minimax h3 workflow can deliver up to 2K resolution at 24fps for about 15 seconds. Its native canvas has a 768px short edge, capped at 768x1344 pixels and rounded to a multiple of 32.

3

Which generation modes are included?

The comfyui minimax h3 template library includes three examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks in character, style, motion, camera, or voice.

4

Can it generate audio together with video?

Yes — the comfyui minimax h3 model produces native stereo audio including voice, sound effects, and music, synthesized alongside the video in one pass and synced into a single MP4 file.

5

How do I begin using this workflow?

Update ComfyUI to version 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 workflow, and follow the pop-up to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.

6

Is there a way to increase generation speed?

Yes — install SageAttention and the KJNodes custom nodes, then add a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double generation speed.

Begin Creating with the comfyui minimax h3 Workflow

Run MiniMax H3 locally in ComfyUI with native stereo audio, open weights, and full parameter control — text-to-video, image-to-video, and reference-to-video workflows ready to go.