Computer Vision 2.0: Light as the First-Class Citizen

Computation over memory: from collected pixels to simulated light.

With thousands of open-source datasets for almost every task, and millions of dollars spent on annotating the longest of long-tail cases, Computer Vision is now in our Roombas, vehicles, and phones. The enabling advance was that instead of hand-writing logic to detect pedestrians, we specify the program's behavior and leave the logic to a trainable network. In that paradigm, large datasets of pixels are the first-class citizens.

But pixels are a poor substrate: they are a low-dimensional projection of the path light actually takes through the scene. As simulation becomes ubiquitous and its marginal cost falls, a shift is underway: light, not pixels, is becoming the first-class citizen.

Figure 1. Where the pixels come from: collected versus simulated

Two parallel loops. CV 1.0: a camera collects frames of pixels into a frozen dataset, humans label them, the AI trains on them and is deployed; new events in the field mean another round of collection and labelling. CV 2.0: the world lives inside a light simulation whose sensor produces pixels with labels attached, the AI trains on them and is deployed, and the events it meets update the world inside the simulation. CV 1.0 memory:collect and label Collect pixels into a dataset camera pixels dataset, frozen Pixelsno labels labelled by hand AI Deploy the AI acts in the world AI event new events mean another round of collection and labelling CV 2.0 simulate,act,update Light simulation the world inside the simulator light world sensor Pixelswith labels trains AI Deploy the AI acts in the world AI event events in the field update the world inside the simulation
CV 1.0 collects frames of pixels from a camera into a dataset, pays humans to label them, and freezes the result. When the deployed AI meets something new, the answer is another round of collecting and labelling. In CV 2.0 the world lives inside a light simulation: rays leave a source, bounce off the scene, and land on a simulated sensor, so the pixels arrive with their labels attached. Those pixels train the AI, the AI is deployed to act in the real world, and the events it meets there update the world inside the simulation.

Light as a First-Class Citizen

The insight of Computer Vision 2.0 is treating light as the fundamental unit of visual understanding. Instead of learning patterns in pixels, we model how light interacts with the environment. Three developments made this viable:

  1. GPU compute costs have dropped below $0.50 per hour for high-end cards
  2. Physics engines can now simulate complex light interactions in real-time
  3. Neural networks can efficiently learn from simulated data

This solves problems pixel-based systems cannot. A traditional system needs thousands of images to learn how reflections work on different surfaces. A physics-based system understands reflection from first principles and needs only the material properties and the lighting.

Computation over Memory: From Datasets to Simulations

Treating light as a first-class citizen means simulating it and how it interacts with the world. Pixels were the simulation we had before. With access to individual photons, you can simulate whatever you want. The physics of light-matter interaction is well known, and for most scenes the marginal cost per ray is near zero. So instead of collecting data with a specific sensor, we simulate it with physics engines, and more realistically still with AI-based physics engines.

The transition to simulation is the largest shift in computer vision since deep learning. It was not possible a decade ago: the simulation-to-reality gap was too large.

CV 2.0 integrates simulation into the stack itself, rather than using it as a tool to bolster performance in the real world. Instead of storing terabytes of real-world data, we generate scenarios with a simulation engine.

The bitter lesson echoes this change: move from datasets that humans tediously collect, and scenarios that humans think up, to scenarios generated by computation. An accurate simulation is a general-purpose method for dataset collection, because it improves as compute increases. AI still learns from data; the data is now unbounded.

Figure 2. Computation over memory: filling the long tail

A long-tail curve of scenarios: collected datasets cover the frequent head; simulation generates the rare tail. how often it happens scenarios, ordered from common to rare → the head: collected the tail: simulated rain at dusk unprotected left turn in fog once in years
Collected datasets are dense where the world is repetitive and thin where it is not. A simulation engine fills the tail on demand, and its coverage grows with compute rather than with the fleet.

Quantifiable advantages

  1. Storage: a trillion-parameter vision model can generate more unique scenarios than have been collected in history. This is a simplification, but the point stands: we can generate more data than we can collect, and all we have to do is ask.
  2. Edge cases: simulation generates millions of rare scenarios that real-world collection would see once in years
  3. Cost: a million synthetic images of rain cost about $1,000; collecting and annotating real ones costs $10,000 or more

Practical Implementation

CV 2.0 changes how teams are built. The traditional team, split across data collection, sensing, annotation, and ML, is replaced by a leaner, more specialized team at 30 to 40% of the original size.

The New Team Structure

Instead of large data collection and annotation teams, a CV 2.0 organization is built around three groups:

Figure 3. The new org structure for CV robotics

The CV 2.0 team loop: simulation engineering feeds ML engineering, which ships to the real world; operations finds the missing scenarios and sends them back to event simulation. Simulation Engineering Physics accuracy light transport and rendering Event simulation scenarios and edge cases Data engineering validates against real data ML Engineering sim-to-real transfer Real world deployed system Operations finds what is missing simulated data models gaps observed in the field scenario requests
The team is a loop. Simulation engineering replaces data collection and annotation; operations is a small team whose job is to notice what the simulator does not yet contain and send it back to event simulation.

1. Simulation Engineering

This team has three units:

  • Physics accuracy: n-th order light simulation, rendering models, and computational efficiency. This is the team building the simulation engine: physically based, but increasingly ML-driven.
  • Event simulation: test scenarios and edge cases, plus the validation frameworks around them. This is high-level work (an unprotected left turn in fog at Harvard Square) rather than light-matter interaction (how fog scatters light), so LLMs can drive much of it.
  • Data engineering: validates the simulator against real-world data, owns its quality metrics, and is the interface to ML engineering.

2. Operations Team

A small team that identifies the real-world scenarios the simulator is missing. It replaces the data collection team: someone still has to go out and find the event the simulator does not yet contain.

3. ML Engineering

These engineers work across simulation and the real world, and own the transfer between the two. This group looks most like today's CV teams. Its focus is accurate real-time ML systems.

Future Implications

CV 2.0 is not an efficiency improvement. It is a different way to build visual systems. Organizations that go simulation-first will build more reliable systems at a fraction of today's cost and time. The question is not whether this happens, but who leads it.

If you build computer vision systems today, start integrating simulation into how you build your stack. The difference is launching in 2 months instead of 12.