anagnorisis.cloudSign in

← Hourlies

Hourly ·

MiniMax H3 Brings Open-Weights Video Generation With Native Stereo Audio

MiniMax's third-generation H3 model ships with open weights, native stereo audio generation, and 2K video output — optimized to run locally on a single RTX 3060.

MiniMax H3 Brings Open-Weights Video Generation With Native Stereo Audio
Image: Sora / OpenAI, Public domain (license)

MiniMax has released H3, its third-generation video model, as open weights — a rare move in a space dominated by API-walled models. The omni-modal system generates video with real stereo sound in a single pass, up to 2K resolution and 15 seconds per clip, accepting text, images, video, or audio as input.

The model's native audio capability is the headline differentiator. Unlike previous approaches that bolt sound on as a post-processing step, H3 generates stereo audio simultaneously with video in the same inference pass. Motion transfer — where a reference video supplies movement while subject and style come from elsewhere — collapses what used to be five separate tasks into one model call.

ComfyUI shipped day-zero support with significant optimization work under the hood. The engineering team pruned modulation weights, which account for roughly 40% of total parameters, and implemented efficient int8 convolution quantization. The result cuts peak VRAM usage from 123.6 GB in full precision down to 42.5 GB. Combined with dynamic VRAM offloading, the full 2K video model runs on consumer hardware like the RTX 3060.

The weights are available on Hugging Face under the MiniMaxAI organization, alongside ComfyUI workflows for text-to-video, image-to-video, and reference-to-video generation.

Sources: ComfyUI Blog, Hugging Face

More Hourlies Stories

Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.

More from Anagnorisis