AWSM: Agentic World Simulation and Mapping

Pronounced “awesome”. From real spaces to worlds phygital agents can use (phygital = physical + digital).

📖 Read the full interactive article: phygital-ai.github.io/agentic-world-simulation-and-mapping (中文版: zh.html)

AWSM overview

AWSM uses large-model agents to turn visual observations into editable 3D scenes, grounded by camera-pose and depth estimates, including IMU-informed visual–inertial constraints that tie geometry to physical scale. This repository holds the four frozen reconstructions (Blender .blend + .glb) from the study, their manifests, and the full metric tables.

Headline result (M1 → M4) Before After Relative change
Bidirectional surface error 0.376 m 0.072 m ≈ −81%
Model-depth AbsRel (180 views) 16.44% 7.60% ≈ −54%

These are complete-pipeline comparisons in one controlled NVIDIA Isaac Sim scene (World Lobby), not an isolated IMU ablation.

A scene can look plausible and still be spatially wrong

A reconstruction agent can assemble a convincing room while getting a corner, a distance, or a passage wrong. For an embodied agent those errors change where a destination lies and which route a body can follow. AWSM asks how geometric evidence (pose, depth, metric constraints) can constrain agentic modeling so the result is a usable spatial reference, not just a good-looking render.

One agentic workflow, four reconstruction routes

The reconstruction agent runs a shared observe → build → verify → freeze → evaluate loop: it inspects the evidence, writes Blender Python to build named objects, materials, cameras and collision proxies, renders review views, revises, and freezes the scene before any ground truth is shown.

Route Inputs Geometry source
M1 180 sampled RGB frames Visual inference only, no metric scale
M2 Full RGB video ViPE RGB-only poses → pose-conditioned Depth Anything 3
M3 Video + IMU + camera–IMU calibration ORB-SLAM3 monocular-inertial poses → pose-conditioned DA3
M4 Video + GT camera poses (diagnostic) GT poses → pose-conditioned DA3

Results

1. Better pose does not mechanically imply better native depth

Route ATE (m) Rotation (°) Native depth AbsRel Model depth AbsRel
M1 — — — 16.44% †
M2 0.167 0.31 19.06% 9.29%
M3 0.121 0.19 20.38% 8.37%
M4 GT GT 18.87% 7.60%

† M1 uses a GT-assisted Sim(3) alignment and is diagnostic only.

M3 has a better trajectory than M2, yet slightly worse native depth. After object-level modeling the order flips: M3's final scene is closer to GT.

Figure 1. Trajectory, native depth, and model depth

2. Geometric evidence turns a plausible layout into a more faithful space

Route Model → GT (m) GT → model (m) Bidirectional mean (m)
M1 0.345 0.407 0.376
M2 0.132 0.103 0.118
M3 0.069 0.095 0.082
M4 0.069 0.074 0.072

Figure 2. M1–M4 and input GT across five fixed views

3. Inspect the scenes, not just the scores

M4 frozen model (view 72) Input GT RGB (view 72)
M4 view 72 GT view 72

Figure 5. Model-depth error at five fixed views

The interactive article lets you orbit and zoom each frozen scene side by side with GT.

Beyond the lobby: a phone-captured Office Café

A qualitative real-capture case study: an object-centric scene, a solved camera trajectory, and physics replays in one artifact. Every object is interactable; the shake replay shows the physical response under strong shaking.

Procedural model Strong-shake replay
Office Café model Office Café shake

Map-based embodied demo

In NVIDIA Isaac Sim, the reception desk is marked on the reconstructed map, routes are set, and a drone, humanoid, quadruped and wheeled robot navigate there and line up. Routes are predefined and localization uses simulator pose. The demo illustrates downstream use, not a navigation benchmark.

Planned routes Navigation
Planned routes Navigation demo

▶ Watch the demo video in the article.

Files

scenes/M1..M4/scene.blend   frozen Blender scene (editable objects, materials, cameras)
scenes/M1..M4/scene.glb     the same scene as GLB
scenes/M1..M4/manifest.json route manifest with hashes and registration
tables_1_7.json             all seven frozen metric tables and evaluation protocols
awsm.bib                    citation

Download one route:

hf download AWSM01/AWSM --include "scenes/M3/*" --local-dir awsm

Limitations

  • One synthetic scene and one engineering run per route; no claim of general superiority or statistical significance.
  • The four routes are complete systems, not a single-variable ablation; M2 vs. M3 cannot be attributed to IMU alone.
  • M4 uses GT camera poses and is diagnostic, not deployable.
  • The 100-view depth and appearance sets contain modeling RGB inputs; they are not a held-out novel-view benchmark.
  • The robot demo uses predefined routes and simulator pose.

Citation

@misc{awsm_2026,
  title        = {{AWSM}: Agentic World Simulation and Mapping},
  author       = {{Phygital AI}},
  year         = {2026},
  howpublished = {Interactive research article},
  url          = {https://phygital-ai.github.io/agentic-world-simulation-and-mapping/},
  note         = {Reconstruction benchmarks, editable scene assets, and map-based embodied demonstration}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support