THE CRUNCH

Researchers at the University of Maryland and AWS have built LEGO-Anything, a system in which a coding agent turns a single photo into an editable 3D scene by writing and refining a Blender program step by step. The catch, revealed by their new LEGO-Bench benchmark, is that the agents are poor judges of their own work: when asked which of two scene versions better matches the original, their geometric judgments land near or below chance level.

The team built LEGO-Bench to measure this properly. Real photos offer no exact 3D answer key, so the benchmark renders 208 images from 104 professionally built simulator scenes, keeping the true geometry hidden for automated scoring. All six tested GPT configurations delivered a usable scene almost every time, but accuracy varied wildly: the best model, GPT-6 Astra, scored 53.4 percent on indoor scenes and 39.6 percent outdoors, while weaker configurations managed around 15 percent. Accuracy drops as scenes get more complex.

The most telling finding concerns self-assessment. When models had to judge which of two versions better matched the original, their geometric judgments landed near or below chance level, even when evaluating their own work. The researchers traced common failures to poor initial attempts, revisions that undid earlier progress, and unreliable self-judgment, and concluded that refinement should rely on concrete measurements rather than the agent's own opinion.

That insight led to LEGO-Plugin, a training-free extension that anchors the starting scene in the reference image, replaces self-judgment with concrete measurements, and protects correct progress from regressive edits. It improved all six models, with weaker agents gaining up to 62.7 percent and the top model only about two percentage points. Scenes reconstructed without the plugin produced usable but unremarkable results when used for standard vision tasks such as object detection, segmentation and depth estimation.