AI 视频

Sora 正式开放
视频生成跨过门槛

Sora Goes Live:
AI Video Crosses the Line

不再是 waitlist,不再是 demo——Sora 向 Plus 用户全面开放,国产三强同期火力全开。AI 视频生成在 2026 年夏天正式从玩具变成工具,内容创作的成本曲线被再次砍断。

No more waitlist, no more demos -- Sora is fully open to Plus users, and China's big three are firing on all cylinders the same week. In summer 2026, AI video generation officially goes from toy to tool, and the cost curve of content creation gets slashed once again.

No.001 2026.06.24 约 5 分钟阅读 ~5 min read

等待了两年的靴子,终于落地了。

6 月 25 日,OpenAI 正式向所有 ChatGPT Plus 用户开放 Sora 视频生成功能。不是 waitlist 排队,不是邀请制内测,不是精心挑选的 demo——是每个月付 20 美元的普通用户,打开 ChatGPT 就能用的真正产品。

这一天的意义不亚于当年 ChatGPT 首次开放:它标志着 AI 视频生成正式跨过了"可用门槛"。过去两年里,AI 视频工具一直停留在"生成几秒钟的猎奇片段、社交媒体上看热闹"的阶段。而今天的 Sora,已经是一个可以支撑真实生产流程的工具。

Sora 带来了什么?

正式版 Sora 的参数表,每一个数字都在改写行业认知:

支持 1080p 分辨率、最长 60 秒的视频生成——这已经覆盖了绝大多数短视频广告的时长需求。更关键的是镜头语言控制:用户可以指定推、拉、摇、移、跟等运镜方式,AI 会按照电影语言的逻辑生成流畅的镜头运动,而不是之前那种"画面在动但没有镜头感"的僵硬抖动。

角色一致性是另一个质变。过去 AI 视频最大的痛点之一是"同一个人在镜头里脸一直在变",根本没法用于叙事。正式版 Sora 支持参考图锁定角色,生成的人物在整个 60 秒里保持五官、服装、体态的高度一致——这意味着 AI 终于能拍"有人物的故事"了,而不只是风景和空镜。

此外,Sora 支持文本+图像混合输入:你可以上传一张产品图,让 AI 围绕这张图生成视频;也可以上传分镜草图,让 AI 按照你的视觉风格生成动态画面。这个能力对广告和电商行业是核弹级的——因为它意味着"用 AI 改片"成为可能,而不只是"用 AI 凭空生成"。

国产三强:视频赛道中国更猛

Sora 开放的同一周,中国 AI 视频团队也密集放出大招,节奏几乎像是约好了一样。

快手旗下的可灵 AI 日活突破 1000 万,成为全球日活最高的 AI 视频产品。可灵的优势在于对中文语境和中国用户需求的理解——生成中式场景、亚洲面孔、本土化剧情的质量在很多评测中已经超过 Sora。更重要的是,可灵深度嵌入快手生态,用户生成视频后可以直接发布到快手平台,形成了"生成-分发-变现"的闭环。

字节跳动的即梦 AI 同期宣布支持 4K 分辨率视频生成,是全球首个量产 4K AI 视频的产品。4K 意味着 AI 生成的视频第一次达到了专业播出标准——不仅能发抖音,还能上电视广告、能投户外大屏、能进院线贴片。

Vidu(生数科技)则选择了一条差异化路线:主打实时视频生成。用户在对话框里输入文字,视频几乎同时开始播放,延迟控制在 3 秒以内。这种"实时交互"的体验,把视频生成从"渲染等待"变成了"对话式创作"——想象一下,你跟 AI 说"把光调暗一点""让主角转身""背景换成海边",视频即时变化,就像跟导演说话一样。

一个值得注意的现象是:在大模型文本赛道,中国团队始终处于追赶状态;但在视频生成赛道,中国团队和 OpenAI 的差距极小,甚至在某些维度(日活、4K、实时生成)已经实现反超。原因很简单——中国有全球最大的短视频市场、最成熟的创作者生态、最激烈的内容竞争,这些需求端的压力倒逼视频生成技术以更快的速度迭代。

"文字时代的赢家是 OpenAI,图片时代的赢家是 Midjourney,视频时代的赢家可能还没出现。" —— Dawn Vision

成本曲线被砍断:三个行业的地震

AI 视频跨过可用门槛,意味着什么?最直接的冲击,是内容生产成本的数量级下降。

广告行业:一条 30 秒的品牌产品片,传统流程需要策划、脚本、选角、搭景、拍摄、后期,预算 5 万到 50 万不等,周期一到两周。现在用 Sora 或可灵,一个熟练的创作者在 2-3 小时内可以生成 10 条以上的候选视频,从中筛选优质版本做精修,总成本可以压到 500 块以内——成本下降 100 倍,周期从两周缩到半天。

电商行业:主图视频和详情页视频是电商的标配,但过去拍摄一条产品主图视频需要布景、灯光、模特、剪辑,至少 3 天时间、数千元成本。现在上传一张产品图,AI 在 3 分钟内生成环绕展示、场景代入、使用演示等多种版本的视频,成本几乎为零。淘宝和抖音的商家已经开始大规模用 AI 视频替代传统拍摄。

短视频创作:创作者的工作流正在从"拍剪"变成"说剪"。过去一个短视频创作者的典型工作流是:写脚本、准备道具、拍摄、导入剪辑软件、加字幕、配乐、调色、发布——一条高质量短视频需要 4-8 小时。现在越来越多创作者的流程是:写 Prompt、生成视频素材、AI 自动剪辑加字幕、微调后发布,整个过程压缩到 30 分钟以内。创作者的核心竞争力从"拍摄和剪辑技术"转向"讲故事的能力和审美判断力"。

不是所有人都开心

每一次生产力工具的革命,都会同时制造赢家和输家。AI 视频的普及也不例外。

影视后期行业首当其冲。过去需要一整个后期团队花几周完成的特效镜头、场景延伸、素材修补工作,现在一个人用 AI 工具几小时就能搞定。中低端的后期外包业务正在以肉眼可见的速度萎缩。

动画师面临的冲击更为深远。AI 视频生成本质上是在"学习"海量动画和影视素材后重建画面,2D 动画、Motion Graphics、简单的 3D 动画都在被快速替代。独立动画师赖以生存的"一个人做一支短片"的壁垒正在被抹平——因为现在任何人都能"一个人做一支短片"。

模特行业的地震也已经开始。电商模特、平面模特、甚至部分商业广告模特的工作正在被 AI 生成的虚拟模特替代——不需要付酬劳、不需要档期、不会耍大牌、可以随时调整五官和身材,还能精准对应任何目标人群的审美偏好。

视频是比文本更大的市场

最后说一个更大的判断。

过去三年,AI 行业的聚光灯几乎全部打在大模型对话上——ChatGPT、Claude、Gemini、豆包、通义千问,估值和融资都围绕"AI 对话"展开。但一个被忽略的事实是:今天互联网流量的 70% 以上是视频流量,全球数字广告市场的 60% 以上投放在视频形式上,短视频平台的用户时长是纯文本平台的十倍以上。

视频 is a bigger market than text. The commercial value of AI video generation could ultimately exceed that of large-model chat -- because it cuts directly into proven, massive paying markets like advertising, e-commerce, film/TV, education, and gaming, rather than creating an entirely new demand.

OpenAI clearly sees this, which is why Sora wasn't launched as a standalone product but deeply integrated into ChatGPT as part of the workflow. China's Kling, Jimeng, and Vidu see it too, which is why they're moving faster on productization and commercialization than text models ever did.

Summer 2026 may be remembered as the real starting point for AI video. Just as ChatGPT's launch in late 2022 wasn't the end of large models but the beginning, Sora's public release is just the first starting gun of the video generation era. Over the next 12 months, more players, more products, and more use cases will emerge; the barrier to video creation will keep getting pushed down, until "making a video" is as easy as "writing a paragraph."


See you tomorrow.

The shoe that's been dangling for two years has finally dropped.

On June 25, OpenAI officially opened Sora video generation to all ChatGPT Plus users. Not a waitlist, not invite-only beta, not cherry-picked demos -- a real product that any ordinary user paying $20/month can open ChatGPT and use.

The significance of this day rivals the original ChatGPT launch: it marks AI video generation officially crossing the "usability threshold." For the past two years, AI video tools were stuck in the "generate a few seconds of eye-candy clips for social media novelty" phase. Today's Sora is a tool that can support actual production workflows.

What Does Sora Bring?

Every number on the production Sora spec sheet is rewriting industry expectations:

It supports 1080p resolution and up to 60 seconds of video generation -- already covering the vast majority of short-form ad duration needs. More critically is cinematic camera control: users can specify push-in, pull-out, pan, tracking, and other camera movements; AI generates smooth camera motion following cinematic language, not the stiff jitter of "the picture is moving but there's no cinematography" that plagued earlier versions.

Character consistency is another qualitative leap. One of the biggest pain points of AI video was "the same person's face keeps changing within a shot," making it unusable for storytelling. Production Sora supports reference-image character locking; generated characters maintain high consistency in facial features, clothing, and posture throughout the full 60 seconds -- meaning AI can finally shoot "stories with people in them," not just landscapes and B-roll.

Additionally, Sora supports mixed text+image input: you can upload a product photo and have AI generate video around it, or upload storyboard sketches and have AI generate dynamic footage in your visual style. This capability is nuclear-grade for advertising and e-commerce -- because it means "using AI to edit/iterate on footage" becomes possible, not just "generating from scratch."

China's Big Three: Fiercer in Video Than Text

The same week Sora opened up, Chinese AI video teams were also dropping major moves, the timing almost looking coordinated.

Kuaishou's Kling AI crossed 10 million daily active users, becoming the highest-DAU AI video product globally. Kling's edge lies in its understanding of Chinese-language context and Chinese user needs -- the quality of generated Chinese scenes, Asian faces, and localized storylines has surpassed Sora in many benchmarks. More importantly, Kling is deeply embedded in the Kuaishou ecosystem; users can publish generated videos directly to Kuaishou, creating a "generate-distribute-monetize" closed loop.

ByteDance's Jimeng AI announced 4K resolution video generation the same week, making it the world's first product to mass-produce 4K AI video. 4K means AI-generated video reaches professional broadcast standards for the first time -- not just good enough for Douyin, but good enough for TV ads, outdoor billboards, and theatrical pre-roll.

Vidu (Shengshu Technology) took a differentiated route: real-time video generation. Users type text in the chat box and video starts playing almost simultaneously, with latency under 3 seconds. This "real-time interactive" experience transforms video generation from "rendering wait" to "conversational creation" -- imagine telling AI "dim the light," "have the protagonist turn around," "change the background to the beach," and the video changes instantly, like talking to a director.

A noteworthy phenomenon: in the large-model text race, Chinese teams have always been playing catch-up; but in video generation, the gap between Chinese teams and OpenAI is minimal, and in some dimensions (DAU, 4K, real-time generation) they've already pulled ahead. The reason is simple -- China has the world's largest short-video market, the most mature creator ecosystem, and the most intense content competition, and demand-side pressure is forcing video generation technology to iterate faster.

"The winner of the text era was OpenAI, the winner of the image era was Midjourney, and the winner of the video era might not have emerged yet." -- Dawn Vision

The Cost Curve Gets Slashed: Earthquakes in Three Industries

What does AI video crossing the usability threshold mean? The most immediate impact is an order-of-magnitude drop in content production costs.

Advertising: A 30-second brand product spot traditionally required planning, scripting, casting, set building, shooting, and post-production; budgets ranged from 50K to 500K RMB with a one-to-two-week cycle. Now, using Sora or Kling, a skilled creator can generate 10+ candidate videos in 2-3 hours, pick the best versions for polish, and keep total cost under 500 RMB -- a 100x cost reduction, cycle compressed from two weeks to half a day.

E-commerce: Product hero videos and detail-page videos are table stakes in e-commerce, but shooting a product video used to require set design, lighting, models, and editing -- at least three days and thousands of RMB. Now upload a single product photo, and AI generates multiple versions -- 360 showcase, contextual lifestyle, usage demo -- in under 3 minutes at near-zero cost. Taobao and Douyin merchants are already replacing traditional shoots with AI video at scale.

Short-form video creation: Creator workflows are shifting from "shoot-and-edit" to "prompt-and-edit." A short-form creator's typical workflow used to be: write a script, prep props, shoot, import into editing software, add subtitles, score, color grade, publish -- a high-quality short took 4-8 hours. More and more creators' workflow now is: write prompts, generate video footage, AI auto-edits with subtitles, tweak and publish -- the whole process compressed to under 30 minutes. A creator's core competitive edge is shifting from "shooting and editing skill" to "storytelling ability and aesthetic judgment."

Not Everyone Is Happy

Every productivity tool revolution creates winners and losers simultaneously. The mass adoption of AI video is no exception.

Film/TV post-production is taking the first hit. VFX shots, scene extensions, and footage repair that used to take an entire post team weeks can now be done by one person with AI tools in hours. Low-to-mid-tier post-production outsourcing business is visibly shrinking.

Animators face an even deeper impact. AI video generation is essentially reconstructing visuals after "learning" from massive animation and film footage; 2D animation, Motion Graphics, and simple 3D animation are being rapidly displaced. The moat that independent animators relied on -- "one person making a short film" -- is being leveled, because now anyone can "one-person a short film."

The modeling industry is already feeling the earthquake. E-commerce models, print models, and even some commercial ad models are seeing work replaced by AI-generated virtual models -- no fees, no scheduling conflicts, no diva behavior, adjustable facial features and body types on demand, precisely matched to any target demographic's aesthetic preferences.

Video Is a Bigger Market Than Text

One bigger judgment to close.

For the past three years, the AI spotlight has been almost entirely on large-model chat -- ChatGPT, Claude, Gemini, Doubao, Tongyi Qianwen, with valuations and funding all centered around "AI conversation." But an overlooked fact: over 70% of internet traffic today is video, over 60% of global digital ad spend goes to video formats, and user time on short-video platforms is over 10x that of purely text-based platforms.

Video is a bigger market than text. The commercial value of AI video generation could ultimately be larger than large-model chat -- because it cuts directly into proven, massive paying markets (advertising, e-commerce, film/TV, education, gaming) rather than creating a new demand from scratch.

OpenAI clearly sees this, which is why Sora wasn't released as a standalone product but deeply integrated into ChatGPT as part of the workflow. China's Kling, Jimeng, and Vidu see it too, which is why they're pushing productization and commercialization faster than text models did.

Summer 2026 may well be remembered as the real starting line for AI video. Just as ChatGPT's launch in late 2022 wasn't the end of large models but the beginning, Sora's full release is only the first starting gun of the video generation era. Over the next 12 months, more players, more products, and more use cases will surface; the barrier to video creation will keep dropping, until "making a video" is as simple as "writing a sentence."


See you tomorrow.

文字时代的赢家是 OpenAI,图片时代的赢家是 Midjourney,视频时代的赢家可能还没出现。

—— Dawn Vision

The winner of the text era was OpenAI, the winner of the image era was Midjourney, and the winner of the video era might not have emerged yet.

-- Dawn Vision
Sora 正式版能力拆解 · AI 视频生成赛道竞争格局 · 内容生产成本曲线下移 · 视频 vs 文本 AI 商业价值对比
Sora production capabilities breakdown - AI video generation competitive landscape - Content production cost curve decline - Video vs. text AI commercial value comparison
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 10 个源信号自动生成,经编辑部人工审核。素材来源包括:OpenAI Sora 发布公告、可灵/即梦/Vidu 产品更新、AI 视频生成技术评测、广告电商行业调研。

This article was auto-generated by the Dawn Vision cognitive engine processing 10 source signals, with editorial review. Source materials include: OpenAI Sora launch announcement, Kling/Jimeng/Vidu product updates, AI video generation technical reviews, and advertising/e-commerce industry research.