写2000字Prompt描述运镜,不如先花5分钟搭个3D白模。
8月20日,MiniMax正式发布多模态创作Agent工作台MiniMax Design,围绕7月31日开源的原生多模态视频模型H3构建。最引人注目的功能叫"3D导演台"——你可以在一个简化的3D空间里摆角色、调姿态、定机位,从不同视角确认构图和人物关系,然后把这个3D预演结果作为参考输入给H3生成最终视频。AI视频创作第一次有了"先彩排再开机"的专业工作流。
熟悉AI视频的人都知道过去的痛点:你写了一大段精准到焦段和光圈的Prompt,生成出来的视频可能人物穿模、机位乱跳、构图完全不是你想要的,你能做的只有一遍遍重新生成抽卡。运气好抽十次能有一次满意的,运气不好抽几十次都不行。3D导演台的思路很直接:把最容易出错的空间关系用3D方式预先确定下来,AI只需要在确定的框架里做渲染和细节填充,而不是连机位构图都要靠猜。
从"模型"到"Harness":AI视频的工具链进化
MiniMax把Design称为H3开源后首个官方Harness,这个词很关键。Harness本意是"马具/挽具",在AI语境里指的是围绕模型运行的一整套工程机制——任务规划、上下文管理、工具调用、素材组织、多步骤协作、错误纠正,也就是如何把裸模型能力组织成可落地的生产工具。
过去半年AI视频领域的竞争主要在"模型"层面:谁生成的视频更清晰、谁的动作更连贯、谁的时长更长、谁的一致性更好。Sora、Kling、Hunyuan Video、H3——各家在模型效果上你追我赶,但给创作者的工具体验普遍还停留在"输入Prompt,等待生成"的原始阶段。专业创作者需要的不是一个更会猜Prompt的模型,而是一整套能接入自己工作流的生产工具。
Design的思路是Agent化的。你不需要写复杂的Prompt,只需要用自然语言描述创作目标,Agent会自动理解需求、拆解任务、调用文字/图片/视频/音频模型处理素材,然后完成编辑、剪辑、字幕添加等后期工作。对于已经在用ComfyUI做节点式工作流的专业创作者,Design还支持接入原有本地工作流,可以云端运行也可以本地部署,Agent通过对话协助调整节点参数。
AI视频的"Blender时刻"还要多久?
量子位在相关报道里提到一个判断:AI视频创作正在进入"预演时代"。这个判断很准确。3D导演台解决的是AI视频生产中最大的不确定性——空间和构图——把"抽卡"变成"可规划的创作"。
但这还只是开始。真正的专业视频创作是一个极其复杂的流程:剧本、分镜、角色设计、场景搭建、动画、灯光、渲染、剪辑、调色、音效、配乐——AI现在能参与的环节还很有限。3D导演台相当于把"分镜"这个环节从文字升级到了3D预演,但后面还有很长的路要走。
好消息是这条路的方向已经清晰了。就像Blender把3D建模、动画、渲染、后期整合到一个开源平台里一样,AI视频领域也会出现自己的"Blender时刻"——一个统一的、开放的、可扩展的创作平台,把模型能力和专业工作流无缝结合。MiniMax Design是往这个方向走的一步,但它是闭源的SaaS产品;真正的"AI视频Blender"大概率会出现在开源社区。
"从'抽卡碰运气'到'先预演再生成',AI视频正在经历PS时代之前的暗房到数码时代的跨越。"—— 一位AI视频创作者
Design目前覆盖的场景包括KOC内容、UGC短视频、电商种草、带货视频、教育科普和创意短片。这些都是对生产成本敏感、对产出速度要求高的场景,正好匹配AI视频当前的能力边界。MiniMax相关负责人表示,Design的目标是把H3的模型能力进一步组织到商业内容生产流程中,同时持续吸收开源社区的工作流经验。
H3是7月31日发布的全模态生成模型,支持文本、图像、视频、音频统一理解和生成,能输出带原生双声道音频的视频。从模型发布到推出官方创作工作台,MiniMax只用了20天——这个速度比很多人预想的快。随着模型能力继续进化、工具链持续完善,AI视频从"玩具"变成"生产工具"的拐点,可能比大多数人预期的来得更早。
明天见。
Skip the 2000-word prompt describing camera moves; spend five minutes building a 3D animatic instead.
On August 20, MiniMax officially launched MiniMax Design, a multimodal creative Agent workbench built around H3, the natively multimodal video model open-sourced on July 31. The standout feature is called the "3D Director Console" — you can position characters, adjust poses, and set camera angles in a simplified 3D space, confirm composition and spatial relationships from multiple viewpoints, then feed that 3D previsualization as a reference to H3 for final video generation. For the first time, AI video creation has a professional "rehearse before shooting" workflow.
Anyone familiar with AI video knows the pain: you write a hyper-specific prompt down to focal length and aperture, but the generated video has clipping characters, random camera jumps, composition nothing like what you wanted — and your only option is to regenerate repeatedly like pulling a slot machine. Get lucky, one in ten is usable; get unlucky, dozens of attempts go nowhere. The 3D Director Console is conceptually simple: lock down the most error-prone spatial relationships in 3D first, so the AI only needs to render and fill detail within a defined framework rather than guessing at camera and composition from scratch.
From "Model" to "Harness": Toolchain Evolution in AI Video
MiniMax calls Design the first official Harness for H3 since open-sourcing — and "harness" is the key word. In AI, a harness is the engineering mechanism wrapped around a model: task planning, context management, tool invocation, asset organization, multi-step orchestration, error correction — essentially, how you organize raw model capability into a production-ready tool.
The past six months of AI video competition have largely been at the "model" layer: who generates sharper video, smoother motion, longer clips, better consistency. Sora, Kling, Hunyuan Video, H3 — everyone chases each other on model quality, but the creator tool experience universally remains at the primitive "prompt in, wait for generation" stage. Professional creators don't need a model that's better at guessing prompts; they need a complete production tool that plugs into their existing workflow.
Design takes an Agent-native approach. You don't write complex prompts; describe your creative goal in natural language, and the Agent automatically interprets requirements, decomposes tasks, calls text/image/video/audio models to process assets, then handles editing, cutting, and subtitling in post. For pro creators already using ComfyUI for node-based workflows, Design supports importing existing local workflows, running in cloud or local mode, with the Agent assisting via conversation to adjust node parameters.
How Far Until AI Video's "Blender Moment"?
QbitAI's coverage offered a sharp framing: AI video creation is entering the "previsualization era." That's accurate. The 3D Director Console solves AI video's biggest uncertainty — space and composition — turning "lottery generation" into "plannable creation."
But this is only the beginning. Real professional video production is extraordinarily complex: screenwriting, storyboarding, character design, set building, animation, lighting, rendering, editing, color grading, sound design, scoring — AI can currently participate in a limited subset of stages. The 3D Director Console effectively upgrades "storyboarding" from text to 3D previs, but there's a long road ahead.
The good news is the direction is clear. Just as Blender integrated 3D modeling, animation, rendering, and post into a single open platform, AI video will see its own "Blender moment" — a unified, open, extensible creation platform that seamlessly combines model capability with professional workflows. MiniMax Design is a step in that direction, but it's a closed SaaS product; the real "Blender of AI video" will likely emerge from open source.
"From 'lottery-ticket prompting' to 'previsualize-then-generate,' AI video is crossing the equivalent of the darkroom-to-digital leap that preceded the Photoshop era."— An AI video creator
Design currently covers scenarios including KOC content, UGC short video, e-commerce seeding, shoppable video, educational content, and creative shorts. These are all cost-sensitive, fast-turnaround use cases that match AI video's current capability envelope. A MiniMax representative stated that Design aims to further organize H3's model capabilities into commercial content production workflows while continuously incorporating open-source community workflow patterns.
H3 is a full-modal generative model released July 31 that supports unified understanding and generation of text, images, video, and audio, with output including native dual-channel audio. From model release to official creative workbench: MiniMax did it in 20 days — faster than many expected. As model capabilities continue to evolve and toolchains mature, the inflection point where AI video shifts from "toy" to "production tool" may arrive sooner than most expect.
See you tomorrow.
From 'lottery-ticket prompting' to 'previsualize-then-generate,' AI video is making the darkroom-to-digital leap.
— An AI video creator
MiniMax Design · H3模型 · 3D导演台 · AI视频 · 预演时代 · 多模态创作 · Agent工作台 · Harness · AI设计 · 创作工具
MiniMax Design · H3 model · 3D Director Console · AI video · previsualization era · multimodal creation · Agent workbench · harness · AI design · creative tools