Focus · 焦点

ChatGPT Images 2.5发布
涂鸦即生成,AI图片进入手绘时代

ChatGPT Images 2.5 Launches
Sketch to Generate, AI Images Enter Hand-Drawn Era

在对话框里随手画个草图,AI就能生成专业级图片——OpenAI把图片创作的门槛降到了'会涂鸦就行'。延迟降低50%、新增模板库、支持图片批注修改,全球用户每周已生成超30亿张图片。

Doodle a rough sketch in the chat box, and AI generates professional-grade images — OpenAI has lowered the barrier for image creation to 'can draw stick figures.' 50% latency reduction, new template library, image annotation support, and over 3 billion images generated globally every week.

No.050 2026.09.09 约 10 分钟阅读 ~10 min read

9月8日,OpenAI在官方博客上发布了ChatGPT Images 2.5。一句话总结:你可以在对话框里随手画个草图,AI就能把它变成专业级图片

这个功能叫Sketch。不需要专业绘画技能,不需要复杂的提示词工程,你只需要在ChatGPT对话中输入@Sketch,然后用鼠标或触控板画一个房间布局、一件衣服轮廓、甚至只是一团涂鸦——Images 2.5会把它理解为视觉参考,生成对应的高质量图片。

这不是渐进式更新,这是图片创作范式的根本转变。

从「描述图片」到「画出图片」

过去两年,AI图片生成的主流交互方式是文字描述。你想要一张海报,得写一段详细的提示词:主体是什么、背景是什么、风格是什么、光线从哪个方向打、色调是暖还是冷。描述越精确,生成效果越好;描述模糊,结果就一塌糊涂。

这导致了一个尴尬的现实:AI图片生成的门槛不是技术门槛,是语言门槛。很多人脑子里有清晰的画面,但就是不知道怎么用文字准确描述出来。更讽刺的是,那些擅长写提示词的人,往往是文字工作者而非视觉创作者——AI图片工具在某种程度上成了「文字好的人的视觉玩具」。

Sketch功能正在打破这个壁垒。它把交互方式从「描述」变成了「表达」——你可以用最原始的涂鸦来表达你想要的画面,AI负责理解你的意图并生成专业级结果。这就像从「用文字描述一首歌给作曲家听」变成了「自己哼一段旋律,作曲家帮你编曲」。

"Sketch不是让不会画画的人变成画家,而是让脑子里有画面但不会用文字描述的人,终于能把画面变成现实。"—— 一位创意工作者的评价

技术细节:延迟降低50%,API双模型

除了Sketch功能,Images 2.5还带来了一系列技术升级。

首先是延迟降低50%。对比前代Images 2.0,新版本的图片生成速度大幅提升。在实际体验中,这意味着你画完草图后,AI生成图片的等待时间从十几秒缩短到几秒——体验上从「等待渲染」变成了「即时反馈」。

其次是模板库(Templates)。OpenAI内置了一系列常用创作模板,覆盖海报、商品图、社交媒体封面等场景。你不需要从零开始画草图,可以选择一个模板作为起点,然后在此基础上修改——这进一步降低了使用门槛。

第三是图片批注修改。你可以在已生成的图片上直接添加评论或批注,告诉AI「这里颜色换一下」「这个元素放大」「背景改成海滩」——AI会根据你的批注进行定向修改,而不是重新生成一张全新的图。这解决了AI图片生成长期以来的一个痛点:微调困难。

第四是提示词随图分享。当你生成了一张满意的图片,可以把提示词和图片一起分享给他人。别人看到你的图片后,可以直接复用你的提示词,替换成自己的照片或元素——这在团队协作和创意分享场景中非常实用。

面向开发者,OpenAI推出了两款API模型:GPT-Image-2.5 Flare(快速)和GPT-Image-2.5 Sunburst(高精度)。Flare适合需要快速迭代的场景,Sunburst适合对质量要求极高的商业用途。

30亿张/周:一个被低估的数字

OpenAI在博客中披露了一个数据:全球用户每周通过ChatGPT Images和API生成超过30亿张图片。

30亿张/周是什么概念?折算下来大约每秒5000张。这意味着在你读完这句话的时间里,全球用户已经通过ChatGPT Images生成了上万张图片。这个数字已经超过了大多数传统图片库的月度更新量——Shutterstock每月新增约2000万张图片,Getty Images的存量约8000万张。

AI图片生成正在从「新鲜玩具」变成「基础设施」。当每周有30亿张图片通过AI生成时,它已经不是小众工具了——它是全球图片供应链的重要组成部分。

竞争格局:Midjourney、Stable Diffusion怎么办?

Images 2.5的发布,让AI图片生成赛道的竞争格局再次洗牌。

Midjourney的优势在于审美和艺术性。它的图片风格独特、质量极高,在设计师和艺术家群体中有很强的口碑。但Midjourney的劣势也很明显:交互方式单一(主要是Discord机器人),不支持草图输入,微调能力弱。

Stable Diffusion的优势在于开源和可定制性。开发者可以在本地部署、微调模型、开发插件,自由度极高。但Stable Diffusion的劣势是使用门槛高,需要一定的技术背景,普通用户很难上手。

ChatGPT Images的优势在于集成度和易用性。它直接嵌入ChatGPT对话界面,用户不需要切换工具;Sketch功能大幅降低了输入门槛;图片批注修改解决了微调问题;模板库进一步降低了使用门槛。而且,ChatGPT的用户基数远超Midjourney和Stable Diffusion——当Images 2.5向所有ChatGPT用户开放时,它的潜在用户量是以亿为单位的。

当然,这不意味着Midjourney和Stable Diffusion会被淘汰。它们各有自己的生态位:Midjourney在高端创意市场仍有不可替代的地位,Stable Diffusion在开源社区和定制化场景中仍是首选。但Images 2.5确实把AI图片生成的主流赛道从「专业工具」拉向了「大众产品」。

iPhone时刻还是Nokia时刻?

有人说Images 2.5是AI图片生成的iPhone时刻——它把一个专业工具变成了大众消费品。

这个类比有道理,但也不完全准确。iPhone之所以改变世界,不只是因为易用,更是因为它创造了一个全新的生态(App Store)。Images 2.5目前还是一个工具,它的生态还没有完全建立起来——模板库是初级形态,图片批注修改是初级形态,提示词分享也是初级形态。

但方向是对的。当图片创作的门槛降到「会涂鸦就行」,当全球每周有30亿张图片通过AI生成,当API双模型让开发者可以构建更复杂的图片应用——Images 2.5正在为AI图片生成的「App Store时刻」铺路。

OpenAI这次发布的时间点也值得玩味。9月8日,正好是秋季发布会季节的开始。苹果、Google、微软都在筹备各自的秋季活动。OpenAI选择在这个时间点发布Images 2.5,显然是想在秋季AI产品大战中抢得先机。

AI图片生成的战场,正在从「谁的模型更强」转向「谁的生态更大」。Images 2.5是OpenAI在这个战场上的重要落子。至于它能不能成为真正的iPhone时刻,要看接下来几个月OpenAI能不能围绕它建立起一个繁荣的开发者和创作者生态。

但有一点是确定的:涂鸦即生成的时代,已经来了。

明天见。

On September 8, OpenAI published ChatGPT Images 2.5 on its official blog. One sentence summary: you can doodle a rough sketch in the chat box, and AI turns it into a professional-grade image.

The feature is called Sketch. No professional drawing skills needed, no complex prompt engineering required. Just type @Sketch in a ChatGPT conversation, then use your mouse or trackpad to draw a room layout, a clothing silhouette, or even just a scribble — Images 2.5 interprets it as a visual reference and generates corresponding high-quality images.

This isn't an incremental update. This is a fundamental paradigm shift in image creation.

From "Describing Images" to "Drawing Images"

Over the past two years, the dominant interaction method for AI image generation has been text descriptions. Want a poster? You need to write a detailed prompt: what's the subject, what's the background, what's the style, where does the light come from, warm or cool tones. The more precise the description, the better the result; vague descriptions produce garbage.

This created an awkward reality: the barrier to AI image generation isn't technical — it's linguistic. Many people have crystal-clear mental images but simply can't articulate them in words. More ironically, the people who excel at writing prompts tend to be writers, not visual creators — AI image tools have become, to some extent, "visual toys for people who are good with words."

Sketch is breaking down that barrier. It shifts the interaction from "describing" to "expressing" — you can communicate your desired image through the most primitive doodles, and AI handles the intent interpretation and professional-grade output. It's like going from "describing a song in words to a composer" to "humming a melody and having the composer arrange it."

"Sketch doesn't turn people who can't draw into artists. It lets people who have clear mental images but can't describe them in words finally bring those images to life."— A creative professional's assessment

Technical Details: 50% Latency Drop, Dual API Models

Beyond the Sketch feature, Images 2.5 brings a suite of technical upgrades.

First: 50% latency reduction. Compared to the previous Images 2.0, the new version dramatically improves image generation speed. In practice, this means the wait time after sketching drops from ten-plus seconds to just a few seconds — the experience shifts from "waiting for a render" to "instant feedback."

Second: Templates. OpenAI has built in a library of creation templates covering posters, product images, social media covers, and more. You don't need to start from scratch with a doodle — pick a template as a starting point and modify from there, further lowering the barrier to entry.

Third: image annotation editing. You can add comments or annotations directly on generated images, telling the AI "change the color here," "make this element bigger," "switch the background to a beach" — and the AI makes targeted modifications based on your annotations instead of regenerating entirely new images. This addresses a long-standing pain point in AI image generation: fine-tuning difficulty.

Fourth: prompt sharing with images. When you generate an image you're happy with, you can share the prompt alongside the image. Others can see your image, reuse your prompt, and swap in their own photos or elements — incredibly useful for team collaboration and creative sharing.

For developers, OpenAI has released two API models: GPT-Image-2.5 Flare (fast) and GPT-Image-2.5 Sunburst (high-precision). Flare suits scenarios requiring rapid iteration; Sunburst is built for commercial use cases demanding maximum quality.

3 Billion/Week: An Underestimated Number

OpenAI disclosed a figure in its blog: globally, users generate over 3 billion images per week through ChatGPT Images and the API.

What does 3 billion/week mean? That works out to roughly 5,000 images per second. By the time you finish reading this sentence, global users will have generated tens of thousands more images through ChatGPT Images. This number already exceeds most traditional stock photo libraries' monthly additions — Shutterstock adds about 20 million images per month; Getty Images holds roughly 80 million total.

AI image generation is transitioning from "novelty toy" to "infrastructure." When 3 billion images are generated via AI every week, it's no longer a niche tool — it's a significant component of the global image supply chain.

Competitive Landscape: What Happens to Midjourney and Stable Diffusion?

The Images 2.5 launch reshuffles the AI image generation competitive landscape.

Midjourney's strengths lie in aesthetics and artistry. Its images have distinctive styles and extremely high quality, with strong word-of-mouth among designers and artists. But Midjourney's weaknesses are clear: limited interaction methods (primarily a Discord bot), no sketch input support, and weak fine-tuning capabilities.

Stable Diffusion's strengths lie in open source and customizability. Developers can deploy locally, fine-tune models, build plugins — maximum freedom. But Stable Diffusion's weakness is its high usage barrier, requiring technical background that makes it hard for ordinary users to get started.

ChatGPT Images' strengths are integration and usability. It's embedded directly in the ChatGPT conversation interface — no tool-switching needed. The Sketch feature dramatically lowers the input barrier. Image annotation editing solves the fine-tuning problem. Templates further reduce the learning curve. And ChatGPT's user base dwarfs both Midjourney and Stable Diffusion — when Images 2.5 opens to all ChatGPT users, its potential user base is measured in the hundreds of millions.

Of course, this doesn't mean Midjourney and Stable Diffusion will be eliminated. Each occupies its own ecological niche: Midjourney remains irreplaceable in the high-end creative market, and Stable Diffusion remains the go-to in open-source communities and customized scenarios. But Images 2.5 is definitively shifting the mainstream AI image generation race from "professional tools" toward "mass-market products."

iPhone Moment or Nokia Moment?

Some are calling Images 2.5 the iPhone moment for AI image generation — turning a professional tool into a mass consumer product.

The analogy has merit but isn't entirely accurate. The iPhone changed the world not just because of usability, but because it created an entirely new ecosystem (the App Store). Images 2.5 is still a tool — its ecosystem isn't fully built yet. Templates are embryonic. Image annotation editing is embryonic. Prompt sharing is embryonic.

But the direction is right. When the barrier to image creation drops to "can draw stick figures," when 3 billion images are generated via AI globally every week, when dual API models let developers build increasingly complex image applications — Images 2.5 is paving the way for AI image generation's "App Store moment."

The timing of this release is also worth noting. September 8, right at the start of fall announcement season. Apple, Google, Microsoft — all are preparing their autumn events. OpenAI's choice to release Images 2.5 now is clearly about gaining an early lead in the fall AI product war.

The AI image generation battlefield is shifting from "whose model is stronger" to "whose ecosystem is bigger." Images 2.5 is OpenAI's important move on this battlefield. Whether it becomes a true iPhone moment depends on whether OpenAI can build a thriving developer and creator ecosystem around it in the coming months.

But one thing is certain: the era of sketch-to-generate has arrived.

See you tomorrow.

Sketch不是让不会画画的人变成画家,而是让脑子里有画面但不会用文字描述的人,终于能把画面变成现实。

—— Dawn Vision编辑部

Sketch doesn't turn people who can't draw into artists. It lets people who have clear mental images but can't describe them in words finally bring those images to life.

— The Dawn Vision Editorial Desk
ChatGPT Images 2.5 · Sketch · 涂鸦生图 · OpenAI · AI图片生成 · 延迟降低50% · 模板库 · 30亿张/周 · 创意工具
ChatGPT Images 2.5 · Sketch · doodle to image · OpenAI · AI image generation · 50% latency drop · templates · 3B/week · creative tools
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 15 个源信号生成,经编辑部人工审核。素材来源:OpenAI官方博客、IT之家、太平洋科技、站长之家、ZAKER新闻。

This article was generated by the Dawn Vision cognitive engine processing 15 source signals, with human editorial review. Sources: OpenAI Blog, IT Home, PConline, Chinaz, ZAKER.