If you work in the AEC (Architecture, Engineering, and Construction) industry, your LinkedIn feed is likely split into two realities.
In Reality A, you have the day-to-day grind of BIM: managing complex Revit families, solving Solibri clashes, and obsessing over metadata. It is precise, laborious, and rigid.
In Reality B, you have the AI explosion generating breathtaking architectural hallucinations that are useless for construction because they lack depth, scale, and physics.
For the last two years, we’ve been waiting for the bridge between these two worlds. We’ve been waiting for an AI that understands that a building isn’t just a picture—it’s a system.
Google’s Project Genie could be just that.
On the surface, Genie (Generative Interactive Environments) is marketed as a fun research project that turns images into playable 2D platformer games. But if you look closer at the underlying technology, you’ll see the blueprint for the next generation of Building Information Modelling.
Here is why a “World Model” like Genie matters more to designers than to gamers.
To understand why Genie is a big deal for BIM, you have to understand the difference between generative video and a world model.
If you ask Sora or Runway to generate a video of a “modern hospital lobby,” it predicts pixels based on visual patterns. It creates a movie. You are a passive observer.
Genie, however, learns a world model. It is trained on internet videos to understand action and consequence. It learns that if you press “right,” the camera pans. If you jump, gravity pulls you down. It isn’t just hallucinating pixels; it is simulating a rudimentary physics engine.
Why does this matter for BIM? Because BIM is, at its core, a simulation of reality. Currently, we build that simulation manually, brick by digital brick. Genie suggests a future where the simulation is inferred.
Right now, the “Text-to-BIM” workflow is non-existent. You can do “Text-to-Image” (Midjourney) or “Text-to-3D-Mesh” (Rodin/Shap-E), but those are static objects.
Imagine applying Genie’s architecture to 3D space (which is the inevitable next step).
The prompt: “A university atrium, timber structure, heavy foot traffic, afternoon light.”
The output: Not a render, but a navigable environment.
Designers could generate a “playable” version of a massing model in seconds. You could hand a controller (or an iPad) to a client and say, “Walk around.”
The client can turn corners. They can look up. If they walk into a wall, they stop. This moves the design phase from looking at pictures of spaces to experiencing the logic of spaces.
BIM is great at hard physics (structural loads, thermal bridging). It is less evolved at “soft” physics (how humans move, crowd dynamics, intuitive flow).
Because Genie is trained on video data of agents moving through worlds, a sufficiently advanced version could be trained on CCTV footage of building interiors.
It could predict where bottlenecks happen in a lobby during rush hour.
It could simulate egress paths during a fire alarm—not by following a coded rule set, but by “imagining” how humans behave in that geometry.
We move from static BIM (geometry + data) to behavioural BIM (geometry + data + predicted usage).
This is the boring part that makes the most money.
To train autonomous construction robots (like Boston Dynamics’ Spot or masonry drones), you need millions of hours of training data in messy, chaotic construction sites. You can’t get that data easily in the real world because it’s dangerous and slow.
Genie can generate infinite, interactive variations of “messy construction sites” that follow the laws of physics. It can create the synthetic data required to train the robots that will eventually build the designs we model in tools like Revit, ArchiCAD and others.
Before we get too excited, we have to address the “object” in the room. Genie generates visual consistency, not semantic data.
Genie sees: A grey pixelated barrier that stops movement.
BIM needs: A
Basic Wall: Generic - 200mm, with a (locally certified) Fire Rating of 1 hour and a specific U-Value.
A World Model can dream a building, but it cannot engineer it. Yet.
The likely future workflow isn’t that Genie replaces your BIM authoring tool. It’s that Genie becomes the “Sketchpad.” The AI generates the interactive massing and the spatial logic, and a secondary “Translator AI” maps that geometry to BIM objects, properties and classifications.
We are moving away from the era where we tell computers explicitly where to put every line (CAD/BIM) and toward an era where we tell computers the rules of the world and ask them to generate the solution.
Google’s Genie is currently focussed on visualisation, but the engine driving it (the ability to learn physics, navigation, and object permanence from video) is the engine that will eventually power the “Concept Design” tab in your BIM software of the future.
The dream of walking through a drawing is as old as architecture itself. We just got a lot closer to waking up inside it.
While Google’s Project Genie is still in the research lab, the following startups and tools are building the practical “translation” engines you can use or test today.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.