🤖 ByteDance Unveils Lucida: Video Transforms into an Editable 3D Scene

ByteDance Seed, in collaboration with Peking University and Zhejiang University, introduced Lucida (arXiv:2608.30821): from video or a set of frames, it reconstructs a room as a set of individual editable 3D objects, not a single static model. In the Place stage, a VLM in a render → edit → re-render loop moves objects in a 3D editor until the scene matches the frames.

🌍 According to the authors' measurements, the scene F-Score on R2S-Scene increased from 0.794 for SAM 3D to 0.924, and object pose accuracy on CA-1M — from 57.8% for RecGen to 83.4%. The goal is simulation-ready individual mesh assets for robotics and digital twins; this is currently confirmed only by the team's own tests.

👤 A demo on PlayCanvas is already open: you can spin around an apartment reconstructed from real footage. There is no open source, API, or public access yet.

Source 1: https://lucida-r2s.github.io/ Source 2: https://arxiv.org/abs/2608.30821