One Agent to Guide Them All: Empowering MLLMs for Vision-and-Language Navigation via Explicit World Representation
一个引导所有代理:通过显式世界表示增强多模态大语言模型用于视觉-语言导航
机构 * Australian Institute for Machine Learning, Adelaide University(澳大利亚机器学习研究所,阿德莱德大学) ; The University of Manchester(曼彻斯特大学) ; Zhejiang University(浙江大学) ; Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR))
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract)
AI总结 通过显式世界表示增强多模态大语言模型,实现视觉-语言导航的解耦框架,提升导航性能与现实应用能力。