Who is betting on what in AI, and against whom.
Predict the world, not text. Learn the dynamics of the environment from video and sensors so an agent can imagine outcomes and plan.
A language model predicts the next word. A world model predicts the next moment: what the scene will look like, where the objects will be, what happens if you push. Train that on video and sensor streams and you get something an agent can plan inside, by imagining outcomes before acting.
Two sub-camps disagree about what to predict. LeCun's JEPA line predicts abstract latent features and argues that generating pixels wastes capacity on irrelevant detail. The video camp (Genie, Cosmos, Luma, Runway, Decart) generates frames directly and points out that its models exist, ship, and produce training data for robots.
The money arrived before the results. AMI Labs raised a $1.03B seed with no product; World Labs raised $1B; Wayve, Waabi, Luma, Runway, Decart, Odyssey and General Intuition each raised hundreds of millions. Genie 3 can generate a playable world in real time, but no latent world model has yet planned a multi-step physical task from raw video better than a vision-language-action model does.
Unfamiliar terms are in the glossary.
The money arrived before the results. Generative video world models (Genie 3, Cosmos 3) exist and are useful as simulators and data engines. JEPA-style latent models have no headline demonstration yet; AMI Labs has $1.03B and no product. LeCun argues pixel reconstruction wastes capacity; the video camp keeps shipping. World-model startups raised more than $6B in the first half of 2026.
What would prove them right. A latent world model that plans a multi-step physical task from raw video with far less data than a vision-language-action model, or one that beats LLM plus tools on long-horizon planning.
Researchers have held a research or faculty role; scientist-founders count. Founders, executives and investors are listed separately so nobody mistakes a boardroom for a lab. Where someone argues for a different camp than the one they work in, it says so.
Universities and institutes with people in this camp, from the roles recorded here. Not a ranking, and not complete: a place is listed when someone in the atlas works there.
Senior authors in this camp's core literature, found through the alphaXiv research index and listed on the strength of one paper each. Being on this list means they publish in the field, not that they have taken a side. The full list is on the People page.



NASDAQ:GOOGLGenie 3 (August 2025) generates playable worlds in real time; Project Genie opened it to subscribers in January 2026.


NASDAQ:NVDACosmos world foundation models as a data engine for robots; Cosmos 3 announced at GTC 2026.Also active here: Genesis AI (Genesis physics simulator as the data engine.); Meta (FAIR still does V-JEPA under Rob Fergus.); OpenAI (Sora positioned as a world simulator.).
Joint-Embedding Predictive Architecture. Encode two views of the world, predict one embedding from the other, never decode to pixels. LeCun's 2022 position paper laid out the whole program.
Instead of a probability over outputs, learn a scalar energy that is low for plausible states. Lets a model say "this is compatible" without enumerating every pixel.
The JEPA camp's central argument: the world has too much unpredictable detail to model at the pixel level, so predict what matters and ignore the rest.
Learn a compact world model from experience, then train the policy inside the model's imagination. Hafner's Dreamer line took this from Atari to real robots.
Generate the next frame conditioned on the player's action, fast enough to play. Genie 3 does it at 24 frames per second.
Fei-Fei Li's term for models that perceive, generate and reason about 3D space rather than text. World Labs' Marble is the product version.
A generated world that stays consistent when you turn around and come back, rather than re-imagining it each frame.
Ha and Schmidhuber: an agent that learns a compressed model of its environment and trains inside its own dream.
LeCun's position paper. JEPA, hierarchical planning and why predicting pixels is the wrong target.
I-JEPA: prediction in representation space, the first working piece of the plan.
DreamerV3. One learned world model, many tasks, no per-task tuning.
Action-conditioned worlds learned from unlabeled video; the generative branch of the camp.
Video world model that plans on a real robot with no robot data in pretraining.
| Lab | Access | Note |
|---|---|---|
| Google DeepMind | NASDAQ:GOOGL | Alphabet (GOOGL): Genie 3, Veo. |
| NVIDIA | NASDAQ:NVDA | Nvidia (NVDA): Cosmos world models plus the GPUs. |
| Meta | NASDAQ:META | Meta (META): V-JEPA continues at FAIR, but the company moved its bet to camp 1. |
| AMI Labs | private | $1.03B seed at $3.5B pre; Bezos, Nvidia, Samsung, Temasek. |
| World Labs | private | $1B round with $200M from Autodesk. |
| Wayve | private | $8.6B; Microsoft, Nvidia, Uber and SoftBank on the cap table. |
| Runway | private | $5.3B. |
| Luma AI | private | $4B+; Saudi HUMAIN is the lead. |
| Decart | private | $4B; reported Anthropic acquisition talks. |
| Waabi | private | $3B; Uber and Volvo. |
Not investment advice. Private valuations are what the last round implied; "reported" means the press, not the company, gave the figure. Full table on the Capital page.