(香港技术院时空智能系统工程研究中心 新闻)

2025年12月23日,香港技术院时空智能系统工程研究中心鄂超、张斌等核心团队成员发表研究成果《面向时空认知的探索式无人机自主视觉导航》,该论文聚焦极端环境下卫星定位信号失效、无人机易迷航撞毁,以及现有视觉-语言-导航技术遇未定义目标点易导航失败的难题,提出无人机探索式导航系统UAV E-Nav。

该无人机探索式导航系统UAV E-Nav,构建了大型预训练无人机导航语言、视觉和行动模型,将高级领航员的时空认知与安全飞行经验融入模型调优过程,实现语言、视觉与知识的跨模态融合适配。该系统由大语言模型(Large Language models,LLM)、导航定位模型 NPM(Navigation Positioning model,NPM)和导航转向模型NSM(Navigation Steering model,NSM)构成,可从自然语言中提取地标名称与行动副词,透过图像语言模型完成现实世界映射,经导航转向模型完成下一飞行方向决策,并同步存储于“心理地图”中;系统无需任何微调或语言标注数据,对真实世界场景具备良好的泛化能力。基于该系统完成了真实场景无人机导航实例化验证,在关闭卫星定位功能的条件下,透过自然语言指令成功实现了复杂户外环境中的远程探索式导航。本研究为无人机探索式导航提供了新的思路与范式,提升了无人机人机交互的工程实用性。

图1 探索式视觉语言导航

图2 UAV E-Nav:我们的系统将实时观察图像和自由形式的文本指令作为输入,并使用三个预训练模型进行导航:用于提取地标的导航大语言模型(Navigation Language Model,NLM)、用于执行地标点和图像匹配的导航定位模型(Navigation Positioning model,NPM)和用于导航路径规划的导航转向模型(Navigation Steering model,NSM)。这使得UAV E-Nav 能够遵循复杂环境中的指令来进行视觉观察导航。

Exploratory UAV Autonomous Visual Navigation with Spatio-Temporal Cognition

On December 23, 2025, core team members including Chao E and Bin Zhang from the Research Center for Spatio-Temporal Intelligent Systems Engineering of the Hong Kong Academy of Technology published the research paper Exploratory UAV Autonomous Visual Navigation with Spatio-Temporal Cognition. Focusing on the challenges of satellite positioning signal failure in extreme environments (which causes UAVs to easily get lost and crash) and the common navigation failure of existing vision-language-navigation technologies when encountering undefined target points, the paper proposes the UAV exploratory navigation system UAV E-Nav.

The UAV exploratory navigation system UAV E-Nav constructs a large pre-trained UAV navigation language-vision-action model, integrating the spatio-temporal cognition and safe flight experience of senior navigators into the model tuning process to achieve cross-modal fusion and adaptation of language, vision, and knowledge. The system comprises three core components: the Large Language Model (LLM), the Navigation Positioning Model (NPM), and the Navigation Steering Model (NSM). It extracts landmark names and action adverbs from natural language instructions, completes real-world mapping via a vision-language model, determines the next flight direction through the navigation steering model, and synchronously stores information in a “cognitive map”. Requiring no additional fine-tuning or language annotation data, the system demonstrates strong generalization capability in real-world scenarios.

Instance verification of UAV navigation based on this system has been completed. With satellite positioning fully disabled, the system successfully realized long-distance exploratory navigation in complex outdoor environments using only natural language instructions. This research provides new ideas and paradigms for UAV exploratory navigation, and improves the engineering practicality of human-UAV interaction.

Fig. 1. Exploratory Visual Language Navigation of UAV E-Nav

Fig. 2. UAV E-Nav: Our system takes real-time observation images and free-form text instructions as inputs, and uses three pre-trained models for navigation: Navigation Language Model (NLM) for extracting landmarks, Navigation Positioning Model (NPM) for performing landmark and image matching, and Navigation Steering Model (NSM) for navigation path planning. This enables UAV E-Nav to follow instructions in complex unknown environments for visual observation and navigation.