AI summaryⓘ
The authors created UrbanGround, a realistic 3D model of Hong Kong, to test how well multimodal large language models (MLLMs) can understand and navigate a city from a first-person perspective. They examined if these models can recognize local scenes, use that knowledge to find destinations farther away, and adapt when routes or pedestrian movement change. They found that while MLLMs can handle simple visual tasks and short-range navigation, they struggle with orientation and adapting to dynamic conditions over longer explorations. Errors tend to pile up because these skills don’t combine smoothly into ongoing, goal-directed movement. The authors see UrbanGround as a tool to study and improve MLLM agents in real-world urban settings.
Multimodal Large Language ModelsUrban Navigation3D Geospatial DataFirst-Person ViewSpatial ReasoningClosed-Loop InteractionPedestrian MotionOrientationGoal-Directed BehaviorSimulated Urban Environment
Authors
Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
Abstract
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.