Computer Vision 2.0: Light as the First-Class Citizen
Computation over memory: from collected pixels to simulated light.
With thousands of open-source datasets for almost every task, and millions of dollars spent on annotating the longest of long-tail cases, Computer Vision is now in our Roombas, vehicles, and phones. The enabling advance was that instead of hand-writing logic to detect pedestrians, we specify the program's behavior and leave the logic to a trainable network. In that paradigm, large datasets of pixels are the first-class citizens.
But pixels are a poor substrate: they are a low-dimensional projection of the path light actually takes through the scene. As simulation becomes ubiquitous and its marginal cost falls, a shift is underway: light, not pixels, is becoming the first-class citizen.
Figure 1. Where the pixels come from: collected versus simulated
Light as a First-Class Citizen
The insight of Computer Vision 2.0 is treating light as the fundamental unit of visual understanding. Instead of learning patterns in pixels, we model how light interacts with the environment. Three developments made this viable:
- GPU compute costs have dropped below $0.50 per hour for high-end cards
- Physics engines can now simulate complex light interactions in real-time
- Neural networks can efficiently learn from simulated data
This solves problems pixel-based systems cannot. A traditional system needs thousands of images to learn how reflections work on different surfaces. A physics-based system understands reflection from first principles and needs only the material properties and the lighting.
Computation over Memory: From Datasets to Simulations
Treating light as a first-class citizen means simulating it and how it interacts with the world. Pixels were the simulation we had before. With access to individual photons, you can simulate whatever you want. The physics of light-matter interaction is well known, and for most scenes the marginal cost per ray is near zero. So instead of collecting data with a specific sensor, we simulate it with physics engines, and more realistically still with AI-based physics engines.
The transition to simulation is the largest shift in computer vision since deep learning. It was not possible a decade ago: the simulation-to-reality gap was too large.
CV 2.0 integrates simulation into the stack itself, rather than using it as a tool to bolster performance in the real world. Instead of storing terabytes of real-world data, we generate scenarios with a simulation engine.
The bitter lesson echoes this change: move from datasets that humans tediously collect, and scenarios that humans think up, to scenarios generated by computation. An accurate simulation is a general-purpose method for dataset collection, because it improves as compute increases. AI still learns from data; the data is now unbounded.
Figure 2. Computation over memory: filling the long tail
Quantifiable advantages
- Storage: a trillion-parameter vision model can generate more unique scenarios than have been collected in history. This is a simplification, but the point stands: we can generate more data than we can collect, and all we have to do is ask.
- Edge cases: simulation generates millions of rare scenarios that real-world collection would see once in years
- Cost: a million synthetic images of rain cost about $1,000; collecting and annotating real ones costs $10,000 or more
Practical Implementation
CV 2.0 changes how teams are built. The traditional team, split across data collection, sensing, annotation, and ML, is replaced by a leaner, more specialized team at 30 to 40% of the original size.
The New Team Structure
Instead of large data collection and annotation teams, a CV 2.0 organization is built around three groups:
Figure 3. The new org structure for CV robotics
1. Simulation Engineering
This team has three units:
- Physics accuracy: n-th order light simulation, rendering models, and computational efficiency. This is the team building the simulation engine: physically based, but increasingly ML-driven.
- Event simulation: test scenarios and edge cases, plus the validation frameworks around them. This is high-level work (an unprotected left turn in fog at Harvard Square) rather than light-matter interaction (how fog scatters light), so LLMs can drive much of it.
- Data engineering: validates the simulator against real-world data, owns its quality metrics, and is the interface to ML engineering.
2. Operations Team
A small team that identifies the real-world scenarios the simulator is missing. It replaces the data collection team: someone still has to go out and find the event the simulator does not yet contain.
3. ML Engineering
These engineers work across simulation and the real world, and own the transfer between the two. This group looks most like today's CV teams. Its focus is accurate real-time ML systems.
Future Implications
CV 2.0 is not an efficiency improvement. It is a different way to build visual systems. Organizations that go simulation-first will build more reliable systems at a fraction of today's cost and time. The question is not whether this happens, but who leads it.
If you build computer vision systems today, start integrating simulation into how you build your stack. The difference is launching in 2 months instead of 12.