Google DeepMind's Genie 3 generates interactive worlds in real time
Genie 3 is a general-purpose world model that generates navigable, interactive 3D environments from text prompts in real time at 720p and 24 fps, staying consistent for a few minutes.
Key facts
- Announced 5 August 2025
- Real-time generation at 24 frames per second, 720p
- Environments remain consistent for a few minutes, with visual memory of about a minute
- Supports 'promptable world events' that alter the scene via text
- Released as a limited research preview
What happened
DeepMind showed a model that renders explorable worlds frame-by-frame in response to user actions.
Why it matters
World models are seen as a path to training embodied agents and robots in unlimited simulated environments.
Changelog
- 2026-09-29: created
Videos (2)
Google Just Turned Street View Into a Video Game
Bilawal Sidhu · 2026-05-19 · ai-madeDescription by Gemini, which watched the video:
Summary
In this video, creator and former Google Maps product lead Bilawal Sidhu reviews Google DeepMind’s Project Genie (Genie 3) integration with Google Maps Street View imagery, announced around Google I/O. He demonstrates how interactive real-time world-generation models can turn 360-degree Street View panoramas into playable, editable 3D-like simulation environments.
What is shown
- [00:00 - 00:44] Introduction to grounding Genie 3 experiences using Google Street View panoramic imagery, showing early demo clips (raccoon on a scooter, Formula 1 car, runner in Austin).
- [00:45 - 00:52] The Project Genie interface showing prompt fields for Environment ("Choose a location from Google Maps") and Character, with a third-person camera toggle.
- [00:53 - 01:29] Driving simulation of a Google Maps-themed Formula 1 car navigating the Las Vegas Strip, complete with an AI-generated speedometer HUD, race checkpoints, and Parisian landmarks.
- [01:30 - 02:02] Third-person simulation of a raccoon and a fox riding scooters around and through the Palace of Fine Arts in San Francisco.
- [02:03 - 02:23] Simulation featuring Google Maps mascot Pegman running past the Ferry Building in San Francisco.
- [02:24 - 03:12] An avatar running along the Ann and Roy Butler Hike-and-Bike Trail over Lady Bird Lake in Austin, Texas, jumping over a railing into the water, and switching to a boat simulation under railway bridges.
- [03:13 - 03:20] Indoor walkthrough of the White House generated from indoor Street View "special collects."
- [03:21 - 03:49] Conceptual transformations, including underwater scuba diving beneath the Golden Gate Bridge, snowstorms on city streets, and historical black-and-white aerial imagery.
- [03:50 - 04:15] Discussion of world models illustrated by a Spider-Man pointing meme representing competing approaches (JEPA, LLM, SLAM, Video-Gen, 3DGS, Google Maps).
- [04:16 - 05:40] Breakdown of retrieval-augmented generation (RAG) for world models using the "Seoul World Model" academic paper as an architectural comparison.
- [06:31 - 06:45] A TechCrunch quote from Jack Parker-Holder noting real-time models lag offline video models by roughly 6 to 12 months in quality.
Claims & numbers
- The presenter states that Genie 3 is Google's real-time interactive world model that autoregressively generates the next video frame based on user controls and inputs.
- The presenter notes that the current version of Project Genie relies only on Street View panoramic photography rather than aerial imagery.
- The presenter quotes Jack Parker-Holder (from a TechCrunch article) stating that this kind of interactive world model is "maybe six to 12 months behind video in terms of the accuracy and quality."
Notable quotes
- [00:19] "What that means is you can reference actual Street View photography of a physical area and use that as a basis for your generation."
- [01:13] "And this is particularly cool because this is just referencing the panoramic imagery. They're not even feeding in the aerial imagery into it yet."
- [06:34] "'I think for this kind of model, it's maybe six to 12 months behind video in terms of the accuracy and quality, so I think it's something we will solve,' Parker-Holder said."
Assessment
This is a creator review and demonstration video examining early access to Google DeepMind's Project Genie Street View integration. The interactive gameplay sequences are actual prototype screen recordings from Genie 3, highlighting both impressive dynamic generation and noticeable visual hallucination artifacts when deviating far from original camera angles.
Lyrics & themes
This video is spoken commentary and demonstration rather than a song. The narration revolves around turning physical mapping data into real-time interactive virtual simulations:
- Real-world holodeck: "How do you take the complexity of reality and put it inside a simulation so you can do anything inside it?" [00:03]
- Interactive generation: "This model is autoregressively predicting the next frame... it can just generate everything on the fly for you." [01:50]
- World simulation editing: "So kind of by bringing reality into latent space, you can now edit it and do things that would have been otherwise very hard or tedious to do in traditional tools." [05:03]
- The future of game engines: "Is this what you imagine GTA 7 is actually going to look like?" [07:33]
Lore & references
- Pegman: The yellow human-shaped icon from Google Maps, animated here as a playable 3D character exploring San Francisco.
- World Models Meme: A classic multi-Spider-Man meme highlighting the rivalry between different paradigms for digital reality representation: Meta's JEPA, LLMs, robotics SLAM, generative video models, 3D Gaussian Splatting (3DGS), and geospatial datasets like Google Maps.
- Seoul World Model (SWM): Reference to a research paper on retrieval-augmented generation (RAG) conditioning video diffusion models on city-scale Street View databases.
- GTA 7: A running gaming culture reference speculating that neural world models will eventually replace traditional polygon-based game engines in future open-world titles.
Visual style & craft
The video blends standard creator video essay production—a lighted webcam talking-head shot and screen recordings of web articles and X (Twitter) threads—with direct gameplay captures of Google’s Genie 3 neural world simulator. The generated simulations exhibit characteristic neural video artifacts, including edge warping, object morphing when pivoting cameras, and dreamlike background hallucinations, contrasting with the static, crisp 2D UI overlays and web interfaces.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.
People are Creating INSANE Worlds with Genie 3
RandomAI · 2026-01-30 · ai-madeDescription by Gemini, which watched the video:
Summary
This video is an overview presented by an AI-voiced narrator on the channel RandomAI, showcasing user creations and interactive gameplay demos generated with Google DeepMind’s Genie 3 world model. The presenter highlights how users across social media are simulating existing games, photorealistic environments, and historical events, while analyzing the current capabilities and constraints of the model.
What is shown
- [00:04] Montage of Genie 3 generated clips (paper airplane over waterfalls, jet ski on tropical ocean, San Francisco superhero flight).
- [00:36] A simulation posted by Riley Goodside of a discarded cigarette pack sliding across a New York subway platform controlled via WASD keys.
- [01:08] A physics demo by Shlomi Fruchter featuring a reflective silver sphere navigating amongst yellow spheres.
- [01:40] A daytime trailer-park bodycam simulator holding a taser, posted by Chris First.
- [02:01] A third-person recreation of Fortnite gameplay running near Tomato Town, noting HUD text distortion.
- [02:47] A low-poly stylized wooden roller coaster simulation winding around castle towers.
- [03:15] A helicopter flight simulator over an urban skyline, followed by a flying winged cat simulation over city skyscrapers [03:48].
- [04:12] A Grand Theft Auto VI-style third-person walking simulation down an Ocean Drive-inspired avenue with sports cars and walking pedestrians.
- [04:54] A sports car driving through a Minecraft cherry blossom biome.
- [05:22] A The Last of Us third-person urban survival clip of a character traversing an overgrown, ruined city street.
- [05:39] A downhill skier navigating a snowy slope with cabins and trees.
- [05:54] A historical recreation of the Crucifixion at Golgotha, depicting crowds, Roman soldiers, and the three crosses.
- [06:30] A recreation of The Legend of Zelda: Breath of the Wild featuring Link gliding with a paraglider and sprinting through open hills.
- [07:26] Discussion of Genie 3 limitations, including a 1-minute real-time exploration cap, paywalling under Google's Ultra subscription, and US region locking.
Claims & numbers
- The presenter claims Genie 3 was announced by Google in 2025 as a foundational world model.
- The presenter claims it will take only "six to seven months" until world models like Genie 3 can generate a fully playable AAA game from a single text prompt.
- The presenter notes the current demo is capped at up to "one minute" of real-time interactive exploration.
- An on-screen graphic claims the model is locked behind Google’s AI Ultra tier priced at "$250/month".
- The presenter claims the prototype is region-locked to the United States.
Notable quotes
- [00:00] "Google just made the best world-building AI model out there. Genie 3 public for everyone to use, and people are already using this to create some of the most diabolical and insane worlds."
- [02:32] "I think that it has only like six to seven months left till Genie 3 or the world-building models are able to generate a completely good, playable AAA game using just a single prompt."
- [07:34] "The interactivity is there, but you can only look around a specific world for a bit, like for only a minute, so that is a problem."
Assessment
This is an AI-generated reaction/curation video compiling viral Genie 3 demonstration clips shared on X. The footage originates from real Genie 3 research prototype demos shared by prominent AI researchers and testers (such as DeepMind's Shlomi Fruchter and prompt engineer Riley Goodside), though the presenter's timeline claim of full AAA game generation within 6–7 months is speculative hype.
Lyrics & themes
- The video is non-musical and consists of an AI-narrated script structured into distinct sections: an introduction, interactive physics demos, game recreations (Fortnite, GTA 6, Zelda, Minecraft), serious/educational use cases, and limitations.
- Theme quote 1 [00:27]: "Will this AI model completely destroy and revolutionize the gaming and VR industry as we know them?"
- Theme quote 2 [01:19]: "Now that is the good thing about Genie 3, that you can become anything in the world. So you can play as a ball, or in a first-person mode, or even in third-person mode..."
- Theme quote 3 [06:17]: "So this could mean a lot for educational videos and learning history by directly looking at it from a first-person view..."
Lore & references
- Shlomi Fruchter: Genie research co-lead at Google DeepMind; his post demonstrating physics and reflection rendering is directly reviewed.
- Riley Goodside: Well-known prompt engineer; featured for his unconventional prompt making a cigarette pack the playable character.
- Gaming Franchises: References to Grand Theft Auto VI, Fortnite, The Legend of Zelda: Breath of the Wild, Minecraft, and The Last of Us to benchmark the fidelity of real-time neural world rendering against commercial game engines.
- Project Genie / AI Ultra: Mentions Google's restricted rollout mechanism for interactive world models.
Visual style & craft
The video combines automated screen captures and embedded social media video posts from X with canned graphic assets (paper textures, animated icons, clean 2D vector text overlays). The narration is synthesized using an AI text-to-speech voice with standard conversational inflections, and the video editing follows an automated script-to-video workflow common to aggregator channels.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.
Related events
- Google's Veo 3 generates video with native audio ★★★★
- OpenAI previews Sora, a text-to-video 'world simulator' ★★★★
- General Intuition and Kyutai release MIRA, a real-time multiplayer world model of Rocket League ★★
Sources (3)
- officialGenie 3: A new frontier for world models (Google DeepMind)
- officialGenie (Google DeepMind models page)
- discussionWikipedia: Genie (world model)
id: 2025-08-05-genie-3 · updated 2026-09-29 · open in the interactive timeline