A photograph has a boundary. Fei-Fei Li is building AI that carries on beyond it.
On September 1, her company World Labs introduced Atlas, a model that works with text, images, video and 3D information. It can generate new viewpoints and reconstruct scenes. Where the input leaves gaps, the company explains, Atlas invents plausible details. More reference images constrain that invention. The model is entering early access with selected partners. World Labs’ Atlas announcement
Before a machine acts on a generated room, someone needs to know which parts were observed and which were invented.
Li’s career makes that question especially interesting. She helped give computer vision a way to measure what it could see. Her next venture is testing what AI can infer from what it cannot.
The woman who helped make vision measurable
Li is a Stanford computer-science professor and World Labs’ co-founder and CEO. Her research spans computer vision, robotic learning and spatial intelligence. ImageNet and its associated challenge are central to that history: a large collection of labeled images and a common task against which researchers could compare progress. Stanford’s profile of Fei-Fei Li
That achievement was collective. The ImageNet challenge paper, co-authored by Li, Olga Russakovsky, Jia Deng and colleagues, documents the work behind assembling annotations and evaluating recognition systems. Organizing a shared test made it easier to see which methods were improving and where they struggled. The ImageNet challenge paper
Her newer ambition stretches beyond naming objects in a picture. In a November 2025 essay, Li describes perception, imagination and action as connected capabilities. People reason about how things fit, where they can move and what a physical interaction might cause, often without putting those judgments into words. She argues that AI needs a richer relationship with space. Li’s essay on spatial intelligence
The ambition is easy to understand. Measuring its success is harder.
When a good guess becomes a physical assumption
The company's camera-control comparison also has a specific scope. Atlas receives explicit camera geometry, while rival models receive text descriptions of camera movements. World Labs acknowledges that different prompting could improve competitors’ results. These are company-published evaluations of particular tasks, rather than independent proof of general physical understanding. Atlas’s evaluation details
That qualification leaves plenty of room for the work to matter. It gives readers a more precise account of what the demonstrations can support.
The route back to a real robot
World Labs has committed resources to that physical ambition. In February it announced $1 billion in new funding, naming investors including AMD, Autodesk, NVIDIA and Fidelity Management & Research Company. In July it announced that robotics and simulation company SceniX was joining it. Funding announcement, SceniX announcement
A week after the acquisition announcement, World Labs described a process that begins with a physical task, builds a simulation, trains a robot’s behavior there and checks its transfer back to hardware. The company reported one-hour autonomous runs involving tasks such as manipulating cables and moving test tubes. Those July demonstrations are earlier work, distinct from September’s Atlas announcement. World Labs’ real-to-sim-to-real demonstrations
Its claim of “zero real-world training data” needs a careful reading. It refers to training the demonstrated robot policies—the systems selecting the robot’s actions. Real observations and recordings still help construct the simulated tasks. The process has a physical starting point, even when a particular policy learns inside a simulation.
The potential attraction is practical. Testing variations in a simulated task could help teams discover weaknesses before committing more hardware time. The value would depend on how well those variations predict what happens when the real robot returns to work.
The next benchmark
World Labs’ own taxonomy distinguishes rendering observations, simulating changes and planning actions. These functions connect, but success in one does not establish success in all three. That is a useful distinction whenever the phrase “world model” starts carrying more meaning than the evidence underneath it. A functional taxonomy of world models
The next evidence should make the work easier for outsiders to assess: how a simulation predicts failures on unfamiliar tasks, how much recording and engineering each new setting requires, and whether independent teams reproduce the gains.
Li’s earlier work helped a research community make progress visible. Her next chapter would become more consequential if spatial intelligence acquired equally clear ways to expose its weaknesses. The convincing world is the beginning. Knowing where it stops describing reality is what will make it useful.
How this article was made
This analysis was prepared using AI-assisted research and drafting from the linked public sources. Company demonstrations and our interpretation are distinguished. Publication was authorized by the publisher. Vastkind has not independently tested Atlas or the robot demonstrations; no original interview or separate human fact-check is claimed.
Photograph
Fei-Fei Li at AI for Good, 8 June 2017. © ITU / R. Farrell (ITU Pictures). CC BY 2.0. Resized and converted to WebP. Image source, license.




