
World models help AI predict how environments change and what actions might cause. This article explains the concept, its relationship with language models, and developments across Meta, DeepMind, Waymo and World Labs-while examining the gap between convincing simulations and reliable predictions.
An AI system can describe a busy warehouse in convincing detail. Ask it to predict what happens when a robot moves a box, blocks an aisle or loses sight of an approaching forklift, and the challenge changes. The system needs to track an environment, anticipate how it might evolve and account for the consequences of an action.
That challenge sits at the centre of world-model research. As AI moves into robotics, autonomous driving and interactive simulation, researchers are developing models that learn how environments behave, rather than relying only on descriptions of them.
The idea has a long history, but recent work has brought it renewed attention. Meta’s V-JEPA family, Google DeepMind’s Genie 3, World Labs’ Atlas and smaller research models such as LeWorldModel illustrate several ways of approaching the problem. They share an interest in prediction, yet differ considerably in what they learn, what they produce and how they are evaluated.
Understanding those differences matters. A world model can be useful without recreating reality in full, while a striking demonstration can still fall short of reliable prediction.
What Is a World Model in Artificial Intelligence?
A world model is a learned representation of an environment that helps an AI system predict its state or how it will change. For planning and control, the model often predicts what could happen after particular actions.
The “world” may be a room, a road, a game or another bounded environment. The term does not imply that the model understands everything about reality.
Consider a robot pushing a cup across a table. A useful model might predict whether the cup will slide, tip over or reach a target position. It may make that prediction through compact internal features, without generating a detailed video of the movement or explicitly calculating every physical force.
This is also where the language of “understanding” requires care. A model can learn patterns that support useful predictions without discovering physical laws in the scientific sense. Performance outside the conditions it has encountered remains an empirical question.
A Research Idea with Deep Roots
World models are established ideas in reinforcement learning and robotics. Agents can use a learned model to explore possible outcomes before choosing an action, reducing their dependence on repeated real-world trials.
David Ha and Jürgen Schmidhuber’s 2018 World Models research offered an influential example, combining compressed visual representations with a model of environmental dynamics. Today’s work builds on this broader history while exploring larger datasets, richer environments and new training methods.
World Models vs Large Language Models: What Is the Difference?
Large language models are typically trained to predict tokens in sequences. Many world models instead learn to predict environmental states, observations or compact representations of them, sometimes conditioned on actions.
The distinction concerns the training objective and intended use. It is not a clean division between models that “generate” and models that “understand.”
Language models can participate in planning, and sequence prediction can produce internal representations of an environment. Research on Othello-playing sequence models, for example, found representations of the board state. That evidence comes from a constrained game setting; it does not establish general physical understanding. It does show why describing language models as mere memorisation systems is too simplistic.
Conversely, a world model can be generative. DeepMind’s Genie 3 produces interactive visual environments, while JEPA-style approaches make predictions in a learned representation space.
Three Broad Approaches to World Modeling
These approaches overlap, but offer a useful way to organise the field:
- Predictive representation models learn compact features of observations and predict missing information or future features. JEPA-based systems belong here.
- Generative environment models produce observations or interactive worlds, allowing users or agents to explore simulated outcomes. Genie 3 illustrates this direction.
- Models for learning control help an agent improve its behaviour using imagined trajectories. DreamerV3 is a prominent example.
The useful question is what each model enables: better perception, better predictions, more effective planning or some combination of these capabilities.
How JEPA Supports World-Model Research
Joint Embedding Predictive Architecture, or JEPA, is closely associated with Yann LeCun’s proposals for learning predictive representations of the world.
In simple terms, a JEPA system encodes available information and learns to predict the representation of related information that is missing or unavailable. The prediction happens in an embedding space rather than through reconstruction of every pixel.
Meta’s I-JEPA, introduced in 2023, applied this principle to images. Subsequent V-JEPA research extended feature prediction to video. These representation-learning tasks are related to world modeling, but are not automatically complete systems for predicting action consequences.
Why Predict Features Instead of Every Pixel?
A robot deciding where to move may need to preserve information about object position and motion while tolerating some variation in background appearance. Predictive representations aim to retain information useful for a task without spending equal effort on every visual detail.
That selectivity is an ambition, not a guarantee. Training must prevent representations from becoming uninformative, and evaluation must establish whether the features preserve what downstream tasks require.
For planning, a system also needs a way to predict the effects of candidate actions and choose between them. JEPA supplies a learning framework; its usefulness depends on the model and the surrounding system.
Key World-Model Developments in 2025 and 2026
The following developments provide a view across predictive representations, generative simulation and control. They are selected milestones, rather than an exhaustive ranking of the field.
Meta V-JEPA 2: From Video Learning to Robot Planning
Meta’s V-JEPA 2 paper, published in June 2025, described pretraining using a dataset containing more than one million hours of internet video, alongside images.
For robot planning, the researchers developed an action-conditioned version, V-JEPA 2-AC, through post-training using less than 62 hours of robot interaction video from the DROID dataset. They then demonstrated tasks such as picking and placing objects with Franka robot arms in two lab environments.
The zero-shot result is meaningful but specific: the demonstrations did not require data collection from the deployment environments or task-specific training. The system had already learned from robot interaction data. This was evidence of transfer for bounded manipulation tasks, rather than a robot acquiring unrestricted capabilities from observation alone.
V-JEPA 2.1: More Detailed Visual Representations
Meta’s V-JEPA 2.1 research appeared in March 2026, with an updated family designed to learn more detailed, spatially structured and temporally consistent features from images and video.
The researchers reported improvements across tasks including action anticipation, depth estimation and robot grasping. Its relevance is practical: recognising a scene at a high level is different from preserving the local information needed to locate an object or track it through movement.
These results strengthen the case for predictive visual representations, while remaining tied to the tasks and evaluation settings studied.
LeWorldModel: A Smaller Model for Predictive Control
LeWorldModel, or LeWM, introduced in March 2026, examined whether a compact JEPA-based world model could be trained stably from raw pixels without a pretrained encoder.
The authors reported a model with approximately 15 million parameters and a training objective containing two loss terms. Compared with the end-to-end alternative examined in the paper, the approach reduced tunable loss hyperparameters from six to one.
They also reported planning speeds up to 48 times faster than foundation-model-based world-model baselines on the evaluated tasks. The contribution is encouraging for smaller research teams, but the speed figure is benchmark-specific. It should not be treated as a general claim that compact models outperform larger models in every environment.
DreamerV3: Learning Behaviour Through Imagined Experience
The DreamerV3 Nature paper, published in 2025 following an initial 2023 preprint, described an agent that learns an environment model and improves its behaviour through imagined trajectories.
It offers another perspective on world models: their value can lie in helping an agent learn and act more effectively, without making realistic visual generation the central goal.
Google DeepMind Genie 3: Interactive Generated Worlds
DeepMind introduced Genie 3 in August 2025. The company described a model capable of generating environments that users can navigate in real time at 24 frames per second and 720p resolution, with consistency lasting a few minutes.
Here, prediction becomes an interactive experience: the environment changes in response to navigation and other inputs.
DeepMind also identified limits, including a constrained action space, challenges in modeling interactions between independent agents and limited interaction duration. These qualifications matter when assessing whether generated worlds can support training or evaluation beyond a demonstration.
Waymo World Model: A Concrete Driving-Simulation Use Case
In February 2026, Waymo introduced the Waymo World Model, built on Genie 3 and adapted for driving simulation.
According to Waymo, the system generates camera and lidar outputs and allows engineers to vary driving inputs, scene layouts and language-controlled conditions. The company showcased rare scenarios that would be difficult to collect at scale in real traffic.
This gives world-model research a concrete operational purpose: expanding the situations an autonomous-driving system can encounter during testing. The announcement illustrates the application; it does not, by itself, establish the simulator’s accuracy in every scenario or quantify its effect on road safety.
World Labs Atlas: Spatial Generation and Reconstruction
World Labs announced Atlas in September 2026 as a model for spatial intelligence, combining text, images, video and 3D information.
The company describes capabilities spanning camera-controlled generation, spatial reconstruction and space-time simulation. Its approach highlights another part of the world-model problem: maintaining spatial relationships as a viewpoint changes and reconstructing environments from limited observations.
World Labs announced early access with selected partners. Its reported capabilities and benchmarks should therefore be attributed to the company. In particular, filling unseen regions with plausible content is useful for creation, but a reconstruction intended for measurement requires evidence that those regions match reality.
AMI: Investment in Predictive World Models
AMI, co-founded by Yann LeCun, has attracted significant investment for research into world models. A March 2026 investor announcement described its focus on self-supervised learning and JEPA, alongside plans for industrial collaboration.
The investment signals interest in this research direction. Whether it produces dependable systems for varied real-world environments will depend on the evidence emerging from that work.
Where World Models Could Make a Practical Difference
Robotics and autonomous driving provide direct reasons to model action consequences. A manipulation robot must anticipate how an object will respond; a driving simulator must generate useful variations of traffic and environmental conditions.
Interactive world generation has a different set of applications, including creative exploration and potential training environments. The criteria differ too. A creator may accept a plausible invented space, while an engineer testing a physical system needs validated behaviour within known limits.
Scientific and industrial applications remain promising directions, but the results above should not be treated as proof of general capabilities in materials discovery, climate prediction or biological simulation. Those domains require their own models, data and validation.
World Models and Digital Twins: Related, but Different
A digital twin represents a particular asset, process or system and is maintained through relevant data connections. A learned world model may contribute predictive capabilities to such a system, but it is not automatically a digital twin. For a warehouse, vehicle or production line, a useful twin also depends on correct configuration, timely data and validation against the actual operation. Generating a plausible environment addresses only part of that work.
How to Judge Whether a World Model Is Useful
Visual quality is easy to notice. Predictive reliability is harder to establish. For models intended to support action, three questions provide a practical starting point:
- Does the model predict the consequences of different actions? A useful prediction must respond appropriately when the proposed action changes.
- Do predictions remain useful over longer sequences? Small errors can accumulate as the model forecasts further ahead.
- Does using the model improve task performance? Evaluation should connect prediction to outcomes such as successful control, rather than relying only on appealing outputs.
Uncertainty matters as well. A prediction based on incomplete observations should not be presented as equally dependable as one made under familiar, well-observed conditions.
The Main Challenges Facing World Models
Long-horizon prediction remains difficult. Models must also handle partial observations, changes in viewpoint and unfamiliar conditions without losing important information about the environment.
There is a further challenge in separating appearance from behaviour. A scene can look coherent while responding incorrectly to an action. Conversely, a compact latent model can produce useful control predictions without generating an attractive scene at all.
Language integration offers opportunities, as the V-JEPA 2 work illustrates through video question answering. However, connecting language, perception and action does not remove the need to evaluate each capability and the system as a whole.
What Comes Next for World Models in AI?
World-model research is expanding the ways AI systems learn, predict and interact. The field includes compact control models, large visual representation models and generative environments, each addressing different parts of the problem.
The next test is practical: can these models make dependable predictions under the conditions where people need to use them? Progress will be measured through better task performance and evidence that the models transfer beyond their original evaluations. For readers following the field, that is a more useful measure than the realism of a demonstration alone.
Resources
- World Models — David Ha and Jürgen Schmidhuber (2018)
- Emergent Linear Representations in World Models of Self-Supervised Sequence Models — Nanda, Lee and Wattenberg (2023)
- I-JEPA: The First AI Model Based on Yann LeCun’s Vision for More Human-like AI — Meta (2023)
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Research Paper (2025)
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning — Research Paper (2026)
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels — Research Paper (2026)
- Mastering Diverse Control Tasks Through World Models — DreamerV3, Nature (2025)
- Genie 3: A New Frontier for World Models — Google DeepMind (2025)
- The Waymo World Model: A New Frontier for Autonomous Driving Simulation — Waymo (2026)
- Atlas: A World Model for Spatial Intelligence — World Labs (2026)
Discover more from Poniak Times
Subscribe to get the latest posts sent to your email.





