🎬 Video · Free

MiniMax H3 Audio Lip Sync

Turn a single photo into a video of that person singing or speaking your audio — with lip movement that matches almost perfectly. This is a MiniMax H3 image-to-video workflow I put together for audio-driven lip sync: feed it a reference image and a reference audio track, and H3 animates the subject to the audio, mouth movement and all.

✓ 100% free — download the workflow and run it locally.
⬇ Download workflow (.json)

Requirements

Model downloads

Custom nodes

Setup steps

  1. Download the 4 model file(s) listed below and place each one into ComfyUI/models/diffusion_models, ComfyUI/models/text_encoders, ComfyUI/models/vae.
  2. Install the 5 custom node(s) listed below via ComfyUI Manager, then fully restart ComfyUI.
  3. Download the workflow .json from the button above and drag it onto the ComfyUI canvas.
  4. Open each loader node and confirm the model files are selected — a fresh install often defaults to the wrong one.
  5. Run the workflow. If a node shows red, update ComfyUI and restart before troubleshooting anything else.

Notes & tips

More resources

Want the 1-click version?

Skip the manual model downloads and node installs — get a ready-to-run installer, or unlock all 75+ with Local Lab Pro.