anagnorisis.cloudSign in

← Hourlies

Hourly ·

MiniMax H3 Opens the Floodgates: Open-Weights Video Model Generates Stereo Audio and 2K Clips on a Consumer GPU

MiniMax releases H3, its first open-weights omni-modal video model — capable of generating 2K video with native stereo audio from text, images, video, or audio prompts, and it runs on an RTX 3060.

MiniMax H3 Opens the Floodgates: Open-Weights Video Model Generates Stereo Audio and 2K Clips on a Consumer GPU

MiniMax has released H3, a 33-billion-parameter omni-modal video generation model under open weights — and ComfyUI shipped day-zero native support. The model collapses five tasks into one: feed it text, images, video, or audio, and it generates up to 15 seconds of 2K video with native stereo sound in a single pass.

Unlike previous generations Hailuo 01 and 02, H3 is the first MiniMax video model released with open weights. It runs locally on consumer hardware — ComfyUI confirms it operates on an RTX 3060. The model handles text-to-video, image-to-video, first-and-last-frame control, and reference-to-video where a reference clip supplies motion while separate inputs define subject and style.

The standout feature is native stereo audio, generated alongside video rather than bolted on afterward. Motion transfer allows a reference video to drive camera moves or performance rhythms while the model fills in new content — making iterative shot editing practical on desktop hardware. The Hugging Face model card shows 1,750 community likes within four days of release.

Sources: ComfyUI Blog, Hugging Face

More Hourlies Stories

Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.

More from Anagnorisis