One-Sentence Answer
After teaching machines to read (LLMs) and to see and draw (image models), the industry is now racing to teach AI space — and the result is a new class of "world models" that generate explorable 3D reality from a sentence, a photo, or a video.
Why "Spatial Intelligence" Is the Next Big Thing
Fei-Fei Li — the researcher who helped kick off the deep-learning era with ImageNet — has a sharp framing: AI is incomplete without spatial intelligence.
Her argument is simple. Language is powerful, but a huge part of human intelligence is spatial: understanding that objects have volume, that you can walk around them, that the world persists when you turn your head. A model that only manipulates text or flat pixels will never fully understand or act in the physical world. Robots, AR glasses, autonomous machines, immersive games — all of them need AI that reasons in 3D.
That is the bet behind the current wave of world models. And in 2025–2026, that bet turned from research demos into shipping products.
The Three Camps of Spatial AI
The field is moving fast, but it roughly splits into three approaches. Understanding the difference helps you pick the right tool.
| Player | Flagship | Output type | Best for | |--------|----------|-------------|----------| | World Labs (Fei-Fei Li) | Marble + World API | Persistent, editable 3D worlds | Creators, developers building spatial apps | | Tencent Hunyuan | Hunyuan World 2.0 | Exportable 3DGS / Mesh scenes | Game devs, virtual sets, digital twins | | Google DeepMind | Genie 3 + SIMA 2 | Real-time interactive video worlds | Research, agent training, embodied AI |
1. World Labs — "3D as Code"
World Labs, co-founded by Fei-Fei Li and Justin Johnson, is the purest expression of the spatial-intelligence thesis. Their tagline says it all: "Text became the universal interface for software; 3D is becoming the universal interface for space."
Their key releases:
- Marble — a frontier multimodal world model, now publicly available. Give it text, an image, or a video and it generates an explorable 3D world.
- World API (January 2026) — brings Marble's world-generation into your own applications, so developers can build spatial experiences programmatically.
- Spark 2.0 — a streamable, level-of-detail system for 3D Gaussian Splatting, so these worlds can run smoothly on the web.
- RTFM (Real-Time Frame Model) — generates video in real time as you interact with it.
World Labs' framing of "3D as code" is the most developer-forward vision in the space: worlds you can generate, edit, simulate, and share the way we currently do with software.
2. Tencent Hunyuan World — The Practical Workhorse
If World Labs is the visionary, Hunyuan World 2.0 is the tool you can use in a production pipeline today. (I wrote a dedicated deep-dive on it — read it here.)
The short version: give it text, an image, or a video and it produces a high-fidelity, walkable 3D scene (in 3D Gaussian Splatting or Mesh format) that you can export into Unity or Unreal Engine and keep editing. It covers world generation, real-world reconstruction (digital twins), and a "character mode" that turns a scene into a playable experience.
Crucially, Tencent has pushed parts of the Hunyuan World family (and FlashWorld) toward open source — which matters enormously for small creators and indie developers who cannot pay enterprise prices.
3. Google DeepMind Genie 3 — Real-Time Interactive Worlds
Genie 3 takes the most cinematic approach. Given a text prompt, it generates a dynamic world you can navigate in real time at 24 frames per second, holding consistency for a few minutes at 720p.
Genie 3 models physical properties — water, lighting, weather — and lets you steer through the world as it's generated. DeepMind frames it explicitly as a stepping stone toward AGI: world models let you train AI agents (like their SIMA 2 agent) in an unlimited curriculum of simulated environments.
The trade-off: Genie 3 worlds are interactive video rather than exportable 3D assets, and they're still research-stage — limited duration, limited consistency over long horizons. But as a glimpse of where this goes, it's stunning.
What Actually Changes for You
This isn't just a research story. Spatial AI collapses the cost and time of anything that involves 3D:
For Content Creators
Virtual sets and backdrops from a prompt. Immersive story worlds. 360° environments for video without a location scout or a render farm.
For Game Developers
Rapid level prototyping. Instead of grey-boxing a level by hand, describe it and iterate. Export to Unity/UE and refine.
For Small Businesses & Marketers
Product visualisations, virtual showrooms, and immersive brand experiences that used to require an agency and a five-figure budget.
For Robotics & Simulation
Training environments generated on demand. Robots and autonomous systems need to learn in 3D, and world models can produce endless varied scenarios.
For Everyday People
The same democratisation we saw with text (ChatGPT) and images (Midjourney) is coming to space. You'll be able to turn an idea, a memory, or a photo into a small world you can walk through — no 3D modelling skills required.
The Honest Limitations
Spatial AI is early. A few things to keep in mind:
- Consistency over time is hard. Real-time models drift; walk too far and the world can lose coherence.
- Fine editing is still clunky. Generating a world is easier than precisely changing one thing in it.
- Compute is heavy. These models are demanding, though streaming and level-of-detail tricks are improving fast.
- Standards are unsettled. 3DGS, mesh, and various export formats coexist; the pipeline from "generated world" to "polished production asset" still needs work.
None of this is a reason to ignore it. It's the same messy, fast-moving early phase we saw with image and video generation — right before they became everyday tools.
How to Start Exploring Today
- Try Marble at World Labs to feel what a generated, explorable world is like.
- Try Hunyuan World at 3d-models.hunyuan.tencent.com/world — the most practical for exporting real assets.
- Watch Genie 3 demos from Google DeepMind to see the real-time frontier.
- Think in use cases, not tech. What in your work involves 3D, space, or environments? Start there.
Final Thoughts
We spent 2023–2025 learning to prompt for words and pictures. The next few years will be about prompting for worlds. Spatial AI won't replace 3D artists any more than image models replaced illustrators — but it will dramatically lower the barrier to entry, and it will create entirely new formats we haven't imagined yet.
Fei-Fei Li calls spatial intelligence the missing piece that will let AI "understand the real world." Whether or not it's the road to AGI, it's clearly the road to a new creative medium. If you make things — games, videos, products, experiences — this is worth watching closely, and worth experimenting with now while it's still early.
Which of these have you tried — Marble, Hunyuan World, or Genie 3? Reach out on social media and tell me what you built. I'm collecting real use cases from creators in our region.
Resources
- World Labs — Marble, World API, and the spatial-intelligence manifesto
- Fei-Fei Li: "From Words to Worlds" — the manifesto
- Tencent Hunyuan World — try it, export to Unity/UE
- Google DeepMind Genie 3 — real-time world model
- My deep-dive on Hunyuan World 2.0




