AI 视频 · 多模态

Google图片搜索25周年
从索引到生成的视觉革命

Google Images at 25
The Visual Revolution from Indexing to Generation

Google图片搜索迎来25周年,回顾从2001年首个图片索引到今天多模态生成式搜索的演变,视觉内容的探索和创作方式正在被AI彻底重构。

Google Images turns 25, looking back at the evolution from the first image index in 2001 to today's multimodal generative search — AI is fundamentally restructuring how visual content is explored and created.

No.015 2026.07.15 约 5 分钟阅读 ~5 min read

2001年,Google上线了图片搜索。起因是一个很具体的需求:用户搜"Jennifer Lopez绿色范思哲礼服",文字搜索找不到满意的结果——因为那件衣服太火了,全网都在讨论,但没有一个文字结果能让用户"看到"它到底长什么样。

25年后的今天,Google图片搜索每月处理超过120亿次视觉查询,Google Lens每月识别超过100亿张图片。从"搜图片"到"用图片搜"再到"生成图片",视觉搜索走过了四分之一个世纪的演变。

三个阶段:索引→理解→生成

Google把视觉搜索的25年分成了三个阶段,这个划分很有启发性。

第一阶段(2001-2015):索引时代。这个阶段的核心问题是"找到相关的图片"。技术上主要靠文字标签和页面上下文——你搜"猫",找到alt文本或周围文字里有"猫"的图片。这个阶段的视觉搜索本质上还是文字搜索的延伸,图片本身只是被索引的对象,计算机并不真正"理解"图片里有什么。

第二阶段(2015-2022):理解时代。2015年前后深度学习爆发,计算机视觉准确率一年上一个台阶。2017年Google Lens上线,用户可以用相机拍任何东西直接搜索——拍一朵花知道它是什么花,拍一家餐厅看到评价和菜单,拍一道题得到解题步骤。这个阶段的核心是"理解图片内容":计算机不仅知道这张图的标签是"猫",还知道这是一只橘猫、坐在沙发上、旁边有个咖啡杯、窗外是晴天。

第三阶段(2022-至今):生成时代。多模态大模型的出现,把视觉搜索从"找图"变成了"创造"。现在你在Google图片搜索里输入"一个未来主义风格的猫咖啡馆",它不仅能给你找到相关的图片,还能直接生成几张符合描述的全新图片;你拍一张自己客厅的照片,说"把沙发换成中世纪现代风格",它能直接给你渲染出改造后的效果。搜索和创作的边界消失了。

下一个25年:视觉成为主要交互界面

回望25年,最震撼的不是技术本身的进步,而是视觉搜索的演变映射了整个人机交互方式的变迁。

PC时代,我们用键盘输入文字跟计算机交互;移动互联网时代,我们用手指触摸屏幕、用相机拍照片;AI时代,视觉正在成为最自然的交互界面——我们指着一个东西问AI这是什么、给AI看一张图说我要改成这样、让AI把我们想象中的画面直接生成出来。

"文字是人类的发明,但视觉是人类的本能。当AI能理解和生成视觉内容,人机交互的下一个范式就是'所见即所得,所指即所说'。"—— Google搜索团队

下一个25年,视觉搜索会变成什么样?Google给出的方向是"沉浸式、对话式、创造性"——你可以在一个场景里无缝地搜索、理解、修改、生成视觉内容,不需要在多个App之间切换。比如你在装修房子,对着毛坯房拍张照,说"我想要日式简约风格",AI直接给你渲染出装修后的效果,每一件家具都可以直接点击购买;你在计划旅行,对着地图说"我想去一个人少、有海、能冲浪的地方",AI直接给你生成实景预览和行程规划。

25年前,人们需要一张图片来理解"Jennifer Lopez的礼服到底有多美";25年后,AI可以直接为你生成你想象中的任何画面。视觉从"被索引的内容"变成了"交互的界面"和"创造的媒介"——这个转变,才刚刚开始。

明天见。

In 2001, Google launched Image Search. It was triggered by a very concrete need: users searching for "Jennifer Lopez green Versace dress" couldn't find satisfactory results in text search — the dress was so famous the whole internet was talking about it, but no text result let users "see" what it actually looked like.

Twenty-five years later, Google Images processes over 12 billion visual queries per month, and Google Lens recognizes over 10 billion images monthly. From "searching for pictures" to "searching with pictures" to "generating pictures," visual search has traversed a quarter-century of evolution.

Three Eras: Index → Understand → Generate

Google divides visual search's 25 years into three phases — a useful framework.

Phase 1 (2001-2015): Indexing era. The core problem here was "finding relevant pictures." Technically it relied mostly on text tags and page context — search "cat" and you'd find images where alt text or surrounding text contained "cat." Visual search in this era was essentially an extension of text search; images were just indexed objects, and computers didn't truly "understand" what was in them.

Phase 2 (2015-2022): Understanding era. Deep learning exploded around 2015, computer vision accuracy leaped forward every year. Google Lens launched in 2017, letting users point their camera at anything and search directly — photograph a flower to identify it, snap a restaurant to see reviews and menus, capture a math problem to get step-by-step solutions. The core of this phase was "understanding image content": computers didn't just know an image was tagged "cat" — they knew it was an orange tabby, sitting on a couch, next to a coffee mug, sunny outside the window.

Phase 3 (2022-present): Generation era. Multimodal foundation models transformed visual search from "finding images" to "creating them." Now when you type "a futuristic cat cafe" into Google Images, it doesn't just find related pictures — it directly generates brand-new images matching your description. Take a photo of your living room and say "change the sofa to mid-century modern," and it renders the renovated result directly. The boundary between search and creation has disappeared.

The Next 25 Years: Vision Becomes the Primary Interface

Looking back 25 years, what's most striking isn't the technical progress itself — it's that visual search's evolution maps directly onto the broader trajectory of human-computer interaction.

In the PC era, we interacted with computers by typing text on keyboards. In the mobile internet era, we touched screens with our fingers and took photos with cameras. In the AI era, vision is becoming the most natural interaction interface — we point at something and ask AI what it is, show AI a picture and say "I want it like this," have AI directly generate the images we imagine.

"Text was a human invention, but vision is human instinct. When AI can understand and generate visual content, the next paradigm of human-computer interaction becomes 'what you see is what you get, what you point at is what you say.'"— Google Search Team

What will visual search look like in another 25 years? Google's direction is "immersive, conversational, creative" — you'll be able to seamlessly search, understand, modify, and generate visual content within a single scene, no app-switching required. For example, you're renovating: take a photo of the bare apartment, say "I want Japanese minimalist style," and AI directly renders the furnished result, every piece of furniture clickable to purchase. You're planning a trip: point at a map and say "I want somewhere quiet, by the ocean, good for surfing," and AI generates a photorealistic preview and itinerary directly.

Twenty-five years ago, people needed an image to understand "how beautiful Jennifer Lopez's dress actually was." Twenty-five years from now, AI can directly generate any image you can imagine. Vision has gone from "indexed content" to "interaction interface" to "creative medium" — and that transformation is just getting started.

See you tomorrow.

"文字是人类的发明,但视觉是人类的本能。当AI能理解和生成视觉内容,人机交互的下一个范式就是'所见即所得,所指即所说'。"

—— Google搜索团队

"Text was a human invention, but vision is human instinct. When AI can understand and generate visual content, the next paradigm of human-computer interaction becomes 'what you see is what you get, what you point at is what you say.'"

— Google Search Team
Google图片搜索 · 25周年 · 视觉搜索 · 多模态AI · 生成式搜索 · Google Lens · 人机交互 · 视觉交互界面
Google Images · 25th anniversary · visual search · multimodal AI · generative search · Google Lens · HCI · visual interaction interface
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 6 个源信号生成,经编辑部人工审核。素材来源:Google AI Blog。

This article was generated from 6 source signals processed by the Dawn Vision cognitive engine, with editorial review. Source: Google AI Blog.