World Models

A world model is a learned model of how an environment evolves in response to actions, letting an agent predict consequences and plan without acting in the real world. The term goes back to Ha and Schmidhuber’s “World Models” (2018), which learned a compact model of a game environment and trained a controller inside it. Genie (2024) learned interactive, action-controllable virtual worlds from unlabelled internet video alone, at 11B parameters. Google DeepMind announced Genie 3 on 2025-08-05 as a general purpose world model that generates navigable worlds from a text prompt in real time at 24 frames per second and 720p, consistent for a few minutes (DeepMind blog).

2026 context

The label is now used for several overlapping things: action-conditioned video generators that simulate interactive scenes, latent predictive models for planning, and simulators used to train robots. Interest is driven by physical AI, where data from the real world is scarce and costly, and by games and agent-training environments (digital-twin-universe). Some researchers argue LLMs alone lack a grounded model of the world; this is a debate, not a settled result.

Related: diffusion-models, multimodal-models, artificial-general-intelligence, autonomous-vehicles, ai-research-labs.

Sources

Open items

  • Genie 3 availability after the 2025-08-05 announcement and other 2025-2026 world-model products were not verified. The “LLMs lack a grounded world model” debate is stated without a cited source.