GALA: Geometry-Aware Language Model for Controllable Object Arrangement

Chucheng Xiang1,2   Runze Wang1,2   Zhi Deng2,†   Ruchao Bao1,2   Yuanwei Zhang3,2   Yuxuan Xie2   Hanliu Wang2   Liangzhen Fei1,2   Wenzheng Wu1,2   Cheng Wan4,2   Peifeng Li1,2   Zhongyuan Liu2   Ligang Liu1,†

1University of Science and Technology of China   2Tencent   3Tsinghua University   4Hong Kong University of Science and Technology
Corresponding authors

SIGGRAPH Asia 2026

GALA scene layout and geometric constraint results

Abstract

Reasoning about fine-grained placement constraints, such as support and containment, is essential for realistic 3D scene creation. It remains a significant challenge for existing methods that rely on simplified object representations like oriented bounding boxes (OBBs), which are blind to detailed geometry. We introduce GALA, a geometry-aware language model with explicit point-cloud perception for controllable object placement. GALA autoregressively predicts the 6D pose of one target object at a time, conditioned on the scene context and a placement instruction. A two-stage training strategy first aligns geometric and textual modalities, then fine-tunes the model for placement. We also curate a multi-source scene dataset covering diverse room types and fine-grained furniture arrangements. Experiments show that GALA significantly outperforms baselines without explicit geometric reasoning, and a multi-stage plan–place–verify agent pipeline uses the model for text-driven scene generation.

BibTeX

@inproceedings{Xiang2026GALA,
  title = {GALA: Geometry-Aware Language Model for Controllable Object Arrangement},
  author = {Xiang, Chucheng and Wang, Runze and Deng, Zhi and Bao, Ruchao and Zhang, Yuanwei and Xie, Yuxuan and Wang, Hanliu and Fei, Liangzhen and Wu, Wenzheng and Wan, Cheng and Li, Peifeng and Liu, Zhongyuan and Liu, Ligang},
  booktitle = {SIGGRAPH Asia 2026 Conference Papers},
  year = {2026},
  publisher = {ACM}
}