AI视频

MiniMax H3全模态视频模型开源
2K音视频同步+国产芯片Day-0适配

MiniMax H3 Open-Sources Universal Video Model
2K AV Sync + Full Domestic Chip Day-0 Support

2K分辨率、15秒音视频同步生成、11种语言、视频编辑能力全球第一——MiniMax H3今日开源,华为昇腾、摩尔线程等9家芯片厂商同日完成适配。

2K resolution, 15-second synchronized audio-video generation, 11 languages, world's #1 video editing — MiniMax H3 open-sources today, with nine chip vendors including Huawei Ascend and Moore Threads completing Day-0 adaptation.

No.027 2026.08.03 约 5 分钟阅读 ~5 min read

AI视频生成的开源格局,今天被一家中国公司重新定义了。

8月3日,MiniMax正式开源新一代通用视频模型MiniMax H3。这不是又一个"文生视频"模型——H3是一个通用型全模态生成系统,能统一理解文本、图像、视频、音频输入,并生成最高2K分辨率、最长15秒、带原生立体声音频的视频。视频编辑能力据评测排名全球第一。

全模态架构:不只是生成视频

H3的系统设计值得细看。它不是一个单一模型,而是由三个模块组成的完整生成管线:H3-Context-IR负责把复杂多模态输入提炼成模型可理解的中间表示;H3-Base根据中间表示生成768p音视频;H3-Regenerate-2K把768p结果连同原始上下文送回模型,重新生成2K高清输出。这个"先生成再精炼"的架构既保证了生成质量,又充分利用了原始上下文的细节。

输入方式极其灵活:支持0-2张图像作为首尾帧控制视频走向,最多9张图像、3段视频、3段音频作为参考输入,混合输入合计最多12个文件。输出支持21:9到9:16的所有主流宽高比,24FPS,32kHz立体声。稳定支持11种语言,包括中、英、日、韩、法、德等。

这意味着什么?意味着你可以用一张产品图生成广告视频,用首尾帧控制视频的开始和结束画面,用参考视频指导运镜风格,用参考音频匹配背景音乐和音效——所有这一切在一个统一框架里完成,不需要在多个工具之间拼接。

国产芯片Day-0全阵容适配

H3开源最值得注意的一个细节是:开源首日,华为昇腾、摩尔线程、沐曦、海光信息、昆仑芯、天数智芯、壁仞科技七家国产芯片厂商,加上AMD和Intel,全部完成了Day-0适配。国内用户和开发者当天就可以在国产算力上跑H3。

这件事的象征意义远大于技术意义。一年前,一个前沿开源模型发布,国产芯片往往要等几周甚至几个月才能跑起来;今天,9家芯片厂商(其中7家国产)在发布当天完成适配。这说明中国AI生态的"模型-芯片"协同已经从"事后追赶"进化到了"同步首发"。

摩尔线程基于MTT S5000智算卡和MUSA软件栈完成适配;壁仞、海光也在第一时间宣布支持。加上同日适配的华为昇腾,国产AI算力栈对前沿多模态模型的支持能力已经成型。

需要注意的是,H3采用的是MiniMax H3 Community License而非MIT协议,商业使用有一定限制。但对研究者和开发者来说,一个视频编辑能力全球第一的模型开放权重,已经足够让AI视频创作的门槛再次大幅降低。从Seedance 2.5到H3,中国AI视频模型正在以周为单位迭代,把Sora曾经的领先优势一点点吃掉。

明天见。

The open-source AI video generation landscape was redefined today by a Chinese company.

On August 3, MiniMax officially open-sourced its new universal video model MiniMax H3. This is not another "text-to-video" model — H3 is a universal multimodal generation system that uniformly understands text, image, video, and audio inputs, and produces videos up to 2K resolution, up to 15 seconds long, with native stereo audio. Video editing capability is rated #1 globally by benchmarks.

Universal Multimodal Architecture: Not Just Video Generation

H3's system design deserves attention. It's not a single model but a complete three-module generation pipeline: H3-Context-IR distills complex multimodal inputs into an intermediate representation the model can work with; H3-Base generates 768p audio-video from that IR; H3-Regenerate-2K feeds the 768p output back with original context to regenerate at 2K. This "generate-then-refine" architecture balances quality and detail preservation.

Input flexibility is impressive: 0-2 images as first/last frame controls, up to 9 images, 3 videos, 3 audio clips as references, all mixed up to 12 files. Output supports every mainstream aspect ratio from 21:9 to 9:16, 24FPS, 32kHz stereo. Eleven languages are stably supported including Chinese, English, Japanese, Korean, French, German, and more.

What does this mean? You can generate an ad video from a product image, control start/end frames with first/last frame inputs, guide camera movement with reference video, match background music with reference audio — all in one unified framework, no stitching between tools.

Full Domestic Chip Day-0 Support

The most notable detail of H3's launch: on open-source day one, seven domestic Chinese chip vendors — Huawei Ascend, Moore Threads, MetaX, Hygon, Kunlunxin, Tianshu Zhixin, Biren Technology — plus AMD and Intel, all completed Day-0 adaptation. Chinese users and developers could run H3 on domestic compute the same day it launched.

The symbolic significance outweighs the technical. A year ago, when a frontier open-source model dropped, domestic chips often took weeks or months to catch up. Today, nine chip vendors (seven domestic) deliver Day-0 support. This means China's AI ecosystem "model-chip" coordination has evolved from "catch-up after the fact" to "simultaneous launch."

Moore Threads adapted on the MTT S5000 compute card with MUSA software stack; Biren and Hygon also announced support on day one. With Huawei Ascend also in the Day-0 lineup, the domestic AI compute stack's support for frontier multimodal models has taken shape.

Note that H3 uses the MiniMax H3 Community License rather than MIT, with some commercial-use restrictions. But for researchers and developers, a world-#1 video editing model opening its weights is enough to dramatically lower AI video creation barriers yet again. From Seedance 2.5 to H3, Chinese AI video models are iterating weekly, eating away at whatever lead Sora once held.

See you tomorrow.

7家国产芯片厂商Day-0适配——中国AI生态的'模型-芯片'协同已经从'事后追赶'到'同步首发'。

—— Dawn Vision编辑部

Seven domestic chip vendors delivering Day-0 support — China's AI ecosystem 'model-chip' coordination has moved from 'catch-up' to 'simultaneous launch.'

— The Dawn Vision Editorial Desk
MiniMax H3 · 视频生成 · 全模态 · 2K视频 · 开源 · Day-0适配 · 国产芯片 · 华为昇腾 · 摩尔线程 · 多模态生成
MiniMax H3 · video generation · universal multimodal · 2K video · open source · Day-0 support · domestic chips · Huawei Ascend · Moore Threads · multimodal generation
Sources · 信源 Sources
  • IT之家/凤凰科技 - MiniMax H3正式开源
  • 东方财富网 - 多家芯片厂商完成Day-0适配
  • HuggingFace - MiniMax-H3开源仓库

本文基于 Dawn Vision 认知引擎处理的 10 个源信号生成,经编辑部人工审核。素材来源:IT之家、凤凰科技、东方财富网、HuggingFace。

This article was generated by the Dawn Vision cognitive engine processing 10 source signals, with human editorial review. Sources: IT Home, Phoenix Tech, Eastmoney, HuggingFace.