以后你跟AI说"帮我把这条采访视频剪个1分钟短视频,加字幕加BGM留最精彩的部分",它可能真的能直接给你出成片。
7月初,开源社区发布了AI MediaKit CLI + Skill,一套让AI Agent可以直接调用的音视频处理工具集。简单说,它把剪视频、加字幕、转格式、混音频、加特效这些专业音视频软件的功能,封装成了Agent可以直接调用的API和Skill——你的Cursor、Claude Code或者其他AI Agent,现在真的能"动手"剪视频了。
从"生成视频"到"处理视频"
过去一年AI视频领域的竞争焦点是"生成":Sora、可灵、Veo、Pika这些工具在比拼谁生成的视频画质更高、时长更长、角色更一致。但对大多数内容创作者来说,视频生成只是工作流的一部分——更多时候你需要处理的是已经拍好的素材:剪采访、做vlog、加字幕、配BGM、调颜色、做字幕条。
AI MediaKit CLI瞄准的就是这个"后半段"市场。它提供的能力包括:视频裁剪和拼接、自动语音识别加字幕、背景音乐混音、格式转换、简单特效和转场、关键帧提取、内容理解和自动高光剪辑。更重要的是,这些能力不是给人在图形界面里点按钮用的,而是设计给Agent调用的命令行工具和Skill——AI可以直接写命令调用这些功能,不需要人操作鼠标。
这意味着什么?举个例子:你录了一期2小时的播客,以前你需要把音频导进剪映、自动识别字幕、手动校对、把精彩片段挑出来、剪成3条适合短视频平台的片段、加片头片尾、导出——整个过程至少两三个小时。现在你可以直接给Agent下指令:"帮我把这个播客音频剪成3条60秒的短视频片段,选最有信息量的部分,加字幕,配个轻快的BGM,导出竖屏格式"——Agent自己调用AI MediaKit的工具,可能十几分钟就给你搞定了。
Agent工具链正在覆盖所有创作领域
AI MediaKit不是一个孤立的产品,它代表了一个清晰的趋势:Agent的工具链正在从编程领域扩展到所有创意生产领域。
在编程领域,Agent已经能调用文件读写、命令行、Git、测试框架等工具完成完整的开发闭环;在设计领域,Figma的MCP插件让Agent能直接操作设计稿;在内容创作领域,NotebookLM的AI Clips能把文档自动生成TikTok风格短视频;现在AI MediaKit补上了音视频处理这块拼图。当Agent能调用的工具越来越多、越来越专业,它能完成的任务也就越来越复杂。
这和NotebookLM做短视频、Meta做vibe coding游戏应用Pocket是同一个方向:AI创作工具的终极形态,不是给人提供一个更好用的软件,而是给Agent提供一套足够强大的工具集,让人只需要说想要什么,Agent负责把东西做出来。
当然,现在的AI MediaKit还在早期阶段,功能相对基础,和Premiere、Final Cut Pro这些专业软件的能力还差得远。但方向是对的——专业软件几十年积累的功能,正在被一点点封装成Agent可以调用的工具。就像AI编程一开始也只能补全几行代码,现在已经能完成完整的功能开发了。
对内容创作者来说,这是最好的时代也是最坏的时代。工具越来越强大,一个人就能干以前一个团队的活;但竞争也会越来越激烈——当生产门槛被无限拉低,真正稀缺的不再是"能把视频做出来"的能力,而是"知道要做什么视频"的审美和判断力。技术会越来越便宜,好品味永远值钱。
In the future, when you tell AI “help me cut this interview into a 1-minute short, add subtitles and BGM, keep the best parts,” it might actually deliver a finished video directly.
In early July, the open-source community released AI MediaKit CLI + Skill, a set of audio/video processing tools that AI Agents can invoke directly. Simply put, it packages the capabilities of professional audio/video software — cutting video, adding subtitles, format conversion, audio mixing, adding effects — into APIs and Skills that Agents can call directly. Your Cursor, Claude Code, or other AI Agent can now literally “get its hands dirty” editing video.
From “Generating Video” to “Processing Video”
Over the past year, AI video competition centered on “generation”: tools like Sora, Kling, Veo, and Pika competing on who could produce higher quality, longer, more character-consistent video. But for most content creators, video generation is only one part of the workflow — more often you need to process already-shot footage: cutting interviews, making vlogs, adding subtitles, adding BGM, color grading, making lower thirds.
AI MediaKit CLI targets exactly this “back half” market. Its capabilities include: video trimming and concatenation, automatic speech recognition for subtitles, background music mixing, format conversion, simple effects and transitions, keyframe extraction, content understanding, and automatic highlight editing. Most importantly, these capabilities aren't designed for humans clicking buttons in a GUI — they're command-line tools and Skills designed for Agent invocation. AI can directly write commands to call these functions without human mouse operations.
What does this mean? For example: you record a 2-hour podcast. Previously you'd need to import audio into CapCut, auto-generate subtitles, manually proofread, pick out highlight segments, cut them into three short-form clips suitable for short video platforms, add intro/outro, and export — the whole process taking at least two to three hours. Now you can simply instruct an Agent: “Cut this podcast audio into three 60-second short video clips, pick the most informative parts, add subtitles, put in an upbeat BGM, export in vertical format” — and the Agent calls AI MediaKit tools itself, potentially finishing in a dozen minutes.
Agent Toolchains Are Covering All Creative Domains
AI MediaKit isn't an isolated product; it represents a clear trend: Agent toolchains are expanding from programming into all creative production domains.
In programming, Agents can already invoke file I/O, command line, Git, test frameworks, and other tools to complete full development loops; in design, Figma's MCP plugin lets Agents directly manipulate design files; in content creation, NotebookLM's AI Clips can auto-generate TikTok-style short videos from documents; now AI MediaKit fills in the audio/video processing piece of the puzzle. As Agents can invoke more numerous and more specialized tools, the tasks they can complete grow in complexity.
This is the same direction as NotebookLM making shorts and Meta's vibe-coded gaming app Pocket: the ultimate form of AI creation tools isn't providing humans with better software — it's providing Agents with a powerful enough toolkit that humans only need to say what they want, and Agents handle making it.
Of course, AI MediaKit is still in early stages with relatively basic functionality, far from the capabilities of professional software like Premiere or Final Cut Pro. But the direction is right — decades of accumulated professional software features are being incrementally packaged into Agent-callable tools. Just as AI coding started with completing a few lines of code and can now complete full feature development.
For content creators, this is the best of times and the worst of times. Tools grow more powerful and a single person can do what used to take a team; but competition intensifies too — when production barriers are lowered to near-zero, what's truly scarce is no longer the ability to ‘make a video,’ but the taste and judgment to know ‘what video to make’. Technology gets cheaper; good taste is always valuable.
所有创意工作流的终点都是:你说想要什么,AI帮你做出来——前提是你真的知道自己想要什么。
—— Dawn Vision编辑部
The endpoint of every creative workflow is: you say what you want, AI makes it for you — provided you actually know what you want.
— Dawn Vision Editorial
AI MediaKit · AI video · audio/video editing · Agent tools · multimodal creation · video editing automation · AI toolchain · content creation