University of Maryland and AWS evaluate GPT-6 Astra for 3D scene coding, hit 53% indoor accuracy
The agent writes, runs, and revises code iteratively, producing a program that captures objects, geometry, layout, and camera position. To measure performance they built LEGO‑Bench, a benchmark of 208 images from 104 indoor and outdoor scenes rendered from simulator ground truth.
Key points
- LEGO‑Bench benchmark uses 208 rendered images with hidden ground‑truth geometry for automated scoring.
- GPT‑6 Astra scored 53.4% indoor and 39.6% outdoor accuracy; weaker GPT configs stayed near 15%.
- LEGO‑Plugin replaces self‑assessment with concrete measurements, improving weaker agents by up to 62.7%.
Six GPT configurations were tested. All delivered a usable scene, but geometric accuracy varied widely. GPT‑6 Astra, the strongest model, achieved 53.4 % accuracy on indoor scenes and 39.6 % on outdoor scenes, while weaker setups hovered around 15 %. Raising the reasoning budget boosted Astra’s office subset score from 32.3 % to 61.8 %. The authors found agents struggled with self‑assessment, often mis‑ranking revisions. Their LEGO‑Plugin, which replaces self‑judgment with concrete measurements, lifted weaker models by up to 62.7 % and gave the top model a modest two‑point gain.
The generated scenes can be queried for vision tasks, yielding roughly half the object‑detection performance of the specialized DINO model, but larger gaps for segmentation and depth compared with SAM 3 and Depth Anything 3. Researchers see GPT‑6 Astra’s spatial understanding as a notable step, while industry players like Unity are already releasing plugins for agents such as Claude Code and Codex.
Model page: GPT-6 Astra →
The story so far
2 episodes →- University of Maryland and AWS evaluate GPT-6 Astra for 3D scene coding, hit 53% indoor accuracythis story
AI agents build 3D scenes from photos but have no idea if they got it right
The Decoder · 3 October 2026
Loading the full article…
This text was published by The Decoder and written by Jonathan Kemper. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- OpenAI releases GPT-6.1 Sol for Codex users, 20% cheaper with Astra-like performance · 5 src
- Meta, OpenAI and Uber launch proactive AI agents that interrupt users · 1 src
- Anthropic introduces mods for Claude Code · 2 src
- Prime Intellect launches Prime Inference for serving open models · 1 src
- Microsoft releases Agent 365 to manage AI agents list · 1 src
Comments
via GitHub Discussions