Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System

Jan 1, 2026·
Haokun Liu
Haokun Liu
,
Zhaoqi Ma
,
Yunong Li
,
Junichiro Sugihara
,
Yicheng Chen
,
Jinjie Li
,
Moju Zhao
· 1 min read
Abstract
Heterogeneous multi-robot systems show great potential in complex tasks requiring coordinated hybrid cooperation, yet traditional approaches relying on static models often struggle with task diversity and dynamic environments. We propose a hierarchical framework integrating a prompted large language model (LLM) and a GridMask-enhanced fine-tuned vision language model (VLM): the LLM performs task decomposition and global semantic map construction, while the VLM extracts task-specified semantic labels and 2D spatial information from aerial images to support local planning. The aerial robot follows a globally optimized semantic path and continuously provides bird-view images that guide the ground robot’s local semantic navigation and manipulation, including target-absent scenarios where implicit alignment is maintained. To our knowledge, this is the first demonstration of an aerial-ground heterogeneous system integrating VLM-based perception with LLM-driven task reasoning and motion planning.
Type
Publication
Advanced Intelligent Systems
publications

Video

Haokun Liu
Authors
PhD Student
I am a PhD student at the DRAGON Lab, The University of Tokyo, advised by Junior Assoc. Prof. Moju Zhao. My research focuses on vision-language-action (VLA) models and language-conditioned navigation for aerial and ground robots.