AWSM: Agentic World Simulation and Mapping
Pronounced “awesome”. From real spaces to worlds phygital agents can use (phygital = physical + digital).
📖 Read the full interactive article: phygital-ai.github.io/agentic-world-simulation-and-mapping (中文版: zh.html)
AWSM uses large-model agents to turn visual observations into editable 3D scenes, grounded by camera-pose and depth estimates, including IMU-informed visual–inertial constraints that tie geometry to physical scale. This repository holds the four frozen reconstructions (Blender .blend + .glb) from the study, their manifests, and the full metric tables.
| Headline result (M1 → M4) | Before | After | Relative change |
|---|---|---|---|
| Bidirectional surface error | 0.376 m | 0.072 m | ≈ −81% |
| Model-depth AbsRel (180 views) | 16.44% | 7.60% | ≈ −54% |
These are complete-pipeline comparisons in one controlled NVIDIA Isaac Sim scene (World Lobby), not an isolated IMU ablation.
A scene can look plausible and still be spatially wrong
A reconstruction agent can assemble a convincing room while getting a corner, a distance, or a passage wrong. For an embodied agent those errors change where a destination lies and which route a body can follow. AWSM asks how geometric evidence (pose, depth, metric constraints) can constrain agentic modeling so the result is a usable spatial reference, not just a good-looking render.
One agentic workflow, four reconstruction routes
The reconstruction agent runs a shared observe → build → verify → freeze → evaluate loop: it inspects the evidence, writes Blender Python to build named objects, materials, cameras and collision proxies, renders review views, revises, and freezes the scene before any ground truth is shown.
| Route | Inputs | Geometry source |
|---|---|---|
| M1 | 180 sampled RGB frames | Visual inference only, no metric scale |
| M2 | Full RGB video | ViPE RGB-only poses → pose-conditioned Depth Anything 3 |
| M3 | Video + IMU + camera–IMU calibration | ORB-SLAM3 monocular-inertial poses → pose-conditioned DA3 |
| M4 | Video + GT camera poses (diagnostic) | GT poses → pose-conditioned DA3 |
Results
1. Better pose does not mechanically imply better native depth
| Route | ATE (m) | Rotation (°) | Native depth AbsRel | Model depth AbsRel |
|---|---|---|---|---|
| M1 | — | — | — | 16.44% † |
| M2 | 0.167 | 0.31 | 19.06% | 9.29% |
| M3 | 0.121 | 0.19 | 20.38% | 8.37% |
| M4 | GT | GT | 18.87% | 7.60% |
† M1 uses a GT-assisted Sim(3) alignment and is diagnostic only.
M3 has a better trajectory than M2, yet slightly worse native depth. After object-level modeling the order flips: M3's final scene is closer to GT.
2. Geometric evidence turns a plausible layout into a more faithful space
| Route | Model → GT (m) | GT → model (m) | Bidirectional mean (m) |
|---|---|---|---|
| M1 | 0.345 | 0.407 | 0.376 |
| M2 | 0.132 | 0.103 | 0.118 |
| M3 | 0.069 | 0.095 | 0.082 |
| M4 | 0.069 | 0.074 | 0.072 |
3. Inspect the scenes, not just the scores
The interactive article lets you orbit and zoom each frozen scene side by side with GT.
Beyond the lobby: a phone-captured Office Café
A qualitative real-capture case study: an object-centric scene, a solved camera trajectory, and physics replays in one artifact. Every object is interactable; the shake replay shows the physical response under strong shaking.
Map-based embodied demo
In NVIDIA Isaac Sim, the reception desk is marked on the reconstructed map, routes are set, and a drone, humanoid, quadruped and wheeled robot navigate there and line up. Routes are predefined and localization uses simulator pose. The demo illustrates downstream use, not a navigation benchmark.
▶ Watch the demo video in the article.
Files
scenes/M1..M4/scene.blend frozen Blender scene (editable objects, materials, cameras)
scenes/M1..M4/scene.glb the same scene as GLB
scenes/M1..M4/manifest.json route manifest with hashes and registration
tables_1_7.json all seven frozen metric tables and evaluation protocols
awsm.bib citation
Download one route:
hf download AWSM01/AWSM --include "scenes/M3/*" --local-dir awsm
Limitations
- One synthetic scene and one engineering run per route; no claim of general superiority or statistical significance.
- The four routes are complete systems, not a single-variable ablation; M2 vs. M3 cannot be attributed to IMU alone.
- M4 uses GT camera poses and is diagnostic, not deployable.
- The 100-view depth and appearance sets contain modeling RGB inputs; they are not a held-out novel-view benchmark.
- The robot demo uses predefined routes and simulator pose.
Citation
@misc{awsm_2026,
title = {{AWSM}: Agentic World Simulation and Mapping},
author = {{Phygital AI}},
year = {2026},
howpublished = {Interactive research article},
url = {https://phygital-ai.github.io/agentic-world-simulation-and-mapping/},
note = {Reconstruction benchmarks, editable scene assets, and map-based embodied demonstration}
}









