Hourly ·
MiniMax H3 Brings Open-Weights Video Generation With Native Stereo Audio
MiniMax's third-generation H3 model ships with open weights, native stereo audio generation, and 2K video output — optimized to run locally on a single RTX 3060.
MiniMax has released H3, its third-generation video model, as open weights — a rare move in a space dominated by API-walled models. The omni-modal system generates video with real stereo sound in a single pass, up to 2K resolution and 15 seconds per clip, accepting text, images, video, or audio as input.
The model's native audio capability is the headline differentiator. Unlike previous approaches that bolt sound on as a post-processing step, H3 generates stereo audio simultaneously with video in the same inference pass. Motion transfer — where a reference video supplies movement while subject and style come from elsewhere — collapses what used to be five separate tasks into one model call.
ComfyUI shipped day-zero support with significant optimization work under the hood. The engineering team pruned modulation weights, which account for roughly 40% of total parameters, and implemented efficient int8 convolution quantization. The result cuts peak VRAM usage from 123.6 GB in full precision down to 42.5 GB. Combined with dynamic VRAM offloading, the full 2K video model runs on consumer hardware like the RTX 3060.
The weights are available on Hugging Face under the MiniMaxAI organization, alongside ComfyUI workflows for text-to-video, image-to-video, and reference-to-video generation.
Sources: ComfyUI Blog, Hugging Face
MiniMax H3 带来原生双声道音频的开权重视频生成
MiniMax的第三代H3模型出厂自带开放权重,原生立体声音频生成和2K视频输出优化以[K 在单个RTX 3060上本地运行。
← 小时精选 小时 · 2026-08-04 04:00 UTC 《MiniMax H3》带来原生立体声音频的开[K 源权重视频生成 MiniMax 的第三代 H3 模型配备了开源权重、原生立体声音频生成以[K 及 2K 视频输出——专为在单个 RTX 3060 上本地运行优化。 图片:Sora / OpenAI,公[K 共领域(许可) MiniMax 发布了其第三代视频模型《MiniMax H3》, 这一模型配备了[K 开源权重、原生立体声音频生成以及 2K 视频输出 ——专门针对在单个 RTX 3060 上本[K 地运行进行了优化。 图片:Sora / OpenAI,公共领域(许可)
More Hourlies Stories
Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.
