🤖 AI Summary
Current robotic manipulation methods suffer from limited generalization due to the scarcity of large-scale, high-fidelity, and physically plausible scene data. This work proposes TableVerse, a novel approach that deterministically reconstructs physically consistent tabletop layouts from unstructured web images and introduces a Real2Sim pipeline to automatically convert real-world images into high-fidelity, simulation-ready environments. Building upon this foundation, the authors integrate task-conditioned trajectory generation, physics-based stability validation, and large-scale synthesis to construct TableVerse-100K—a dataset comprising 100,000 unique scenes paired with collision-free interaction trajectories. This dataset substantially strengthens the data foundation for generalizable robotic manipulation.
📝 Abstract
The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified procedural generation, they frequently suffer from physical implausibility and fail to capture the complex, dense clutter of actual human environments. In this paper, we introduce TableVerse, a fully automated Real2Sim pipeline that shifts the paradigm from imaginative layout generation to deterministic reconstruction from unstructured, in-the-wild image data. Our framework seamlessly processes unscripted internet media into high-fidelity, simulation-ready tabletop environments with accurate metric scales, authentic topologies, and verified mechanical stability. Furthermore, an automated task-conditioned trajectory generation framework is integrated to synthesize high-quality, collision-free pick-and-place demonstrations. Leveraging this complete pipeline, we construct the TableVerse-100K Dataset, a large-scale corpus comprising 100,000 unique, physically consistent environments paired with interactive manipulation trajectories. By capturing diverse asset compositions, realistic spatial distributions, and high-quality demonstrations, TableVerse-100K establishes a highly scalable and high-fidelity data foundation, providing significant value to facilitate future research in generalizable robotic manipulation tasks.