OTHER

DeepMind Claims Its New Genie 3 World Model is a Crucial Advancement Toward AGI

Google DeepMind has launched Genie 3, its latest foundational world model intended for training general-purpose AI agents. This advancement is deemed crucial by the AI lab in their quest for “artificial general intelligence,” which is similar to human-like cognitive abilities.

“Genie 3 is the first real-time interactive general-purpose world model,” remarked Shlomi Fruchter, a research director at DeepMind, during a press conference. “It goes beyond earlier narrow models, as it’s not confined to a single environment; it can generate both photo-realistic and imaginative worlds, and everything in between.”

Currently in research preview and not available to the public, Genie 3 expands on its predecessor, Genie 2 (which can create new environments for agents), and DeepMind’s latest video generation model, Veo 3 (known for its thorough understanding of physics).

Image Credits:Google DeepMind

Using a basic text prompt, Genie 3 can generate several minutes of interactive 3D environments at 720p resolution and 24 frames per second — a notable enhancement from the 10 to 20 seconds that Genie 2 could produce. Furthermore, the model incorporates “promptable world events,” enabling users to adjust the generated environment via prompts.

Importantly, Genie 3 ensures physical consistency over time in its simulations, as it can remember previously generated elements — a characteristic that DeepMind claims was not explicitly programmed into the model.

Fruchter highlighted that while Genie 3 could be valuable in sectors like education, gaming, and creative prototyping, its true significance lies in training agents for general tasks, which he believes is critical for achieving AGI.

“World models are central to AGI, especially for embodied agents, where simulating real-world scenarios presents significant challenges,” noted Jack Parker-Holder, a research scientist on DeepMind’s open-endedness team, during the briefing.

Techcrunch event

San Francisco
|
October 27-29, 2025

Image Credits:Google DeepMind

Genie 3 aims to resolve this bottleneck. Like Veo, it does not rely on a fixed physics engine; rather, DeepMind states that the model learns how the world functions — how objects move, fall, and interact — by remembering prior outputs and reasoning over extended durations.

“The model is auto-regressive, meaning it generates one frame at a time,” Fruchter explained to TechCrunch in an interview. “It refers back to previously generated content to decide what occurs next, which is a vital aspect of the architecture.”

This memory function contributes to the consistency in Genie 3’s simulated environments, allowing it to understand physics similarly to how humans recognize that a glass at the edge of a table is about to fall or when to duck from a descending object.

Significantly, DeepMind asserts that the model has the potential to fully engage AI agents — encouraging them to learn from experiences in ways that resemble human learning in real-world contexts.

As an example, DeepMind showcased testing Genie 3 with an updated version of its generalist Scalable Instructable Multiworld Agent (SIMA), instructing it to meet a series of objectives. In a warehouse, the agent was directed to perform tasks such as “approach the bright green trash compactor” or “walk to the packed red forklift.”

“In all instances, the SIMA agent successfully accomplished its objective,” Parker-Holder stated. “It receives commands from the agent, takes the goal, surveys the simulated world, and executes actions accordingly. Genie 3 simulates forward, and its success depends on its consistency.”

Image Credits:Google DeepMind

However, Genie 3 has its limitations. For instance, while researchers claim it understands physics, a demonstration involving a skier on a mountain did not accurately depict snow’s interaction with the skier.

Additionally, the selection of actions available to an agent is somewhat constrained. Although promptable world events allow for various environmental changes, these are not necessarily actions performed by the agent itself. Furthermore, accurately modeling intricate interactions among multiple autonomous agents in a shared environment continues to be a challenge.

Genie 3 is also limited to a few minutes of continuous interaction, while hours would be needed for effective training.

Despite these drawbacks, the model signifies a major leap in enabling agents to go beyond mere reactions to stimuli, potentially allowing them to plan, explore, seek uncertainties, and learn through trial and error — all crucial elements of self-directed learning deemed necessary for moving toward general intelligence.

“We haven’t yet observed a Move 37 moment for embodied agents, where they can execute unprecedented actions in the real world,” Parker-Holder remarked, alluding to the iconic moment during the 2016 game of Go where DeepMind’s AlphaGo made a groundbreaking move against world champion Lee Sedol, showcasing AI’s ability to discover new strategies beyond human understanding.

“But now, we may be on the brink of a new era,” he concluded.