🎬 Video · Free

MiniMax H3 — Local AI Video with Native Audio

MiniMax H3 is the first open weights model in MiniMax's Hailuo video line, and it's a big one — a general purpose, omni-modal generator that understands text, images, video, and audio together and produces video with native stereo audio in a single pass, up to 2K resolution, 24fps, and around 15 seconds a clip. Voice, sound effects, and music are modeled jointly with the video, not layered on afterward. It was open sourced on August 3, 2026, with native ComfyUI support merged the same day.

✓ 100% free — download the workflow and run it locally.
MiniMax H3 — Local AI Video with Native Audio preview
⬇ Download workflow (.json) ▶ Watch tutorial

Requirements

Model downloads

Custom nodes

Setup steps

  1. Download the 5 model file(s) listed below and place each one into ComfyUI/models/text_encoders, ComfyUI/models/unet, ComfyUI/models/vae.
  2. Install the 7 custom node(s) listed below via ComfyUI Manager, then fully restart ComfyUI.
  3. Download the workflow .json from the button above and drag it onto the ComfyUI canvas.
  4. Open each loader node and confirm the model files are selected — a fresh install often defaults to the wrong one.
  5. Run the workflow. If a node shows red, update ComfyUI and restart before troubleshooting anything else.

Notes & tips

More resources

Want the 1-click version?

Skip the manual model downloads and node installs — get a ready-to-run installer, or unlock all 75+ with Local Lab Pro.