Zendoric
← Back to the day · September 2, 2026

World Labs unveils Atlas, a multimodal world model that generates, reconstructs and simulates 3D spaces

🕒 Published on Zendoric: September 2, 2026 · 08:27

✨ AI-generated · how it's made

World Labs, the company focused on so-called "spatial intelligence," has unveiled Atlas, its new next-generation world model. According to the company itself, it is an "omni" model trained from scratch to operate natively on text, images, video and 3D data.

World Labs, the company focused on so-called "spatial intelligence", has unveiled Atlas, its new next-generation world model. According to the company itself, it is an "omni" model trained from scratch to operate natively on text, images, video and 3D data. The core idea is that all those modalities are combined into a single shared "spatial context", from which the model generates what comes next, maintaining three-dimensional coherence with everything it has already seen and imagining whatever lies off-camera.

The underlying promise is a system capable of covering three fronts at once: generating imagined worlds for creative uses, reconstructing real scenes with high fidelity, and simulating environments so that robots can plan actions. According to the article, Atlas will power future versions of Marble and other company products.

On architecture, the company describes Atlas as a "multimodal autoregressive diffusion transformer": a transformer that generates each element of a multimodal sequence one at a time (autoregressively, like an LLM), but that produces those elements through diffusion (specifically, as a "rectified flow" model that gradually removes noise, something typical of modern image and video generators). The novelty World Labs stresses is that every image or depth map entering or leaving the model is anchored to an explicit 3D position via a camera pose, which makes that context spatial rather than merely temporal or textual. This, they claim, allows far more precise camera control than the text instructions customary in video models, and opens up possibilities such as placing two unrelated reference images in 3D space and asking the model to generate a world that connects them coherently (corridors, doors, intermediate corners, etc.).

On specific capabilities, the article details four major blocks. The first is generation with camera control: from one or several reference images, Atlas generates video with precisely defined camera positions and angles, reaching up to one minute of video at 1440p resolution, and allowing complete camera trajectories to be designed to produce "controllable long videos" in which, according to the company, the user acts as the scene's director instead of merely firing off random generations.

The second block is spatial reconstruction: Atlas can reconstruct real scenes from between one and dozens of images, with no need for specialized capture equipment. The more images it receives, the less it has to "imagine" and the more faithful the reconstruction; the company claims that with just two or three images it already matches or beats models specialized in 3D reconstruction, and that it can make use of more than a hundred input images to recreate real environments with fidelity. The article illustrates this with examples such as the progressive reconstruction of a garden and a house by adding images one by one, or the recreation of Stanford's Main Quad from between two and twenty-five street-level photos, even generating aerial views of the campus. In addition to 2D images and video, Atlas can produce explicit 3D outputs —point clouds and "3D Gaussian splats"—, the same representation Marble already uses, which eases its integration into the company's other products.

The third block is spatio-temporal simulation, where Atlas acts as a world simulator, understanding both spatial structure and its evolution over time. Two applications stand out here: "bullet time"-style video reframing from a handful of ordinary cameras (the article notes that the demo recordings were made with mobile phones and action cameras, without professional equipment), and robotics simulation through "Real-to-Sim" workflows: from casual videos shot on a phone, Atlas helps reconstruct an environment while also generating the RGB and depth data a robot moving through that space would "see", for both navigation and object manipulation, allowing objects, positions, lighting and background to be varied to generate diverse training data.

The fourth block, presented as secondary to the model's main goal, is image generation: Atlas can follow complex instructions, render text and generate 360-degree panoramas from text or images, with a wide variety of visual styles.

On performance evaluation, World Labs presents two comparisons. In generation with camera control, it pits Atlas against several video models (identified in the article as MiniMax H3, Gemini Omni Flash, Happy Horse 1.1, FLUX 3 and Seedance 2.5), using external human evaluators who judge which model better follows a given camera trajectory. According to those figures, Atlas was preferred by between 75% and 94% of voters depending on the rival model, with a margin that, they claim, grows the more complex the camera trajectory. It is worth noting that the comparison uses native camera control for Atlas versus text descriptions for the other models, since these do not accept cameras as native input; the company itself acknowledges that better "prompt engineering" could improve its rivals' results, although it opted for text as the more common modality. In 3D reconstruction from sparse views, Atlas is compared with specialized open-source models (Pi3X, π³, VGGT-Ω 1B, Depth Anything 3 and MapAnything) across several academic benchmarks (DTU, ETH3D, KITTI, NRGBD, 7-Scenes, T&T and ScanNet), showing a lower mean relative point error than all the models compared, according to figures reproduced by World Labs' own team under a common evaluation protocol.

On scaling, the company says it trained a series of models of increasing size and compute during development, and that each new level of compute unlocked new capabilities, which leads it to expect the trend to hold in future models.

As for the launch, Atlas enters an "early access" phase with selected partners; anyone wanting to try it must request access through World Labs' own site. The company also says it is hiring research and engineering staff. The article additionally links, as related content, an interview by Martin Casado (a16z) with Fei-Fei Li and Yunzhu Li on the acquisition of SceniX and the role of simulation in robotics, as well as a piece on the "real-to-sim-to-real" approach to training robot policies, both published by World Labs in late July 2026.

These figures should be read with the caution usual in product announcements: all the benchmarks, comparisons and claims of superiority come from World Labs' own team, and the article mentions no independent evaluation beyond the external "human raters" used for the preference test. Nor does it detail model sizes, training data volume or computational cost, and the model is not yet publicly available, only through an early-access request. That said, the technical approach —an autoregressive transformer with diffusion that anchors each image to a camera pose within a shared spatial context— is consistent with the industry's broader trend toward models that unify image generation, video and 3D reconstruction under a single architecture, and fits the growing interest in using these systems as simulation engines for robotics.

🔗 Related on Zendoric

Sources & references