Hourly ·
MiniMax H3 Opens the Floodgates: Open-Weights Video Model Generates Stereo Audio and 2K Clips on a Consumer GPU
MiniMax releases H3, its first open-weights omni-modal video model — capable of generating 2K video with native stereo audio from text, images, video, or audio prompts, and it runs on an RTX 3060.
MiniMax has released H3, a 33-billion-parameter omni-modal video generation model under open weights — and ComfyUI shipped day-zero native support. The model collapses five tasks into one: feed it text, images, video, or audio, and it generates up to 15 seconds of 2K video with native stereo sound in a single pass.
Unlike previous generations Hailuo 01 and 02, H3 is the first MiniMax video model released with open weights. It runs locally on consumer hardware — ComfyUI confirms it operates on an RTX 3060. The model handles text-to-video, image-to-video, first-and-last-frame control, and reference-to-video where a reference clip supplies motion while separate inputs define subject and style.
The standout feature is native stereo audio, generated alongside video rather than bolted on afterward. Motion transfer allows a reference video to drive camera moves or performance rhythms while the model fills in new content — making iterative shot editing practical on desktop hardware. The Hugging Face model card shows 1,750 community likes within four days of release.
Sources: ComfyUI Blog, Hugging Face
MiniMax H3 Opens the Floodgates:开权重视频模型在消费级GPU上生成立体音频和2[1D[K 2K短片
MiniMax发布H3,其首款开源权重多模态视频模型——能够从文本、图像、视频或音频提[K 示生成2K分辨率的视频,并带有原生立体声,运行在RTX 3060上。
← 小时精选 小时精选 · 2026-08-04 12:00 UTC MiniMax H3 开放闸门:开源权重多[K 模态视频模型能生成立体音频和2K视频片段,使用消费级GPU MiniMax 推出了其首款开[K 源权重多模态视频模型——H3。该模型能够从文本、图像、视频或音频提示中生成具有原[K 生立体声的2K视频,并且它可以在RTX 3060上运行。MiniMax已发布H3,这是一个包含[K 3.3亿参数的开源权重多模态视频模型
More Hourlies Stories
Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.
