Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation
Fly0: 解耦语义定位与几何规划以实现零样本空中导航
Zhenxing Xu, Yihong Lu, Weidong Bao, Zhengqiu Zhu, Jingxuan Zhou, Zhichuang Wang, Ji Wang, Lihua Liu, Wei He
机构
*
National Key Laboratory of Big Data and Decision(大数据与决策国家重点实验室)
;
National University of Defense Technology(国防科技大学)
;
State Key Laboratory of Digital Intelligent Modeling and Simulation(数字智能建模与仿真国家重点实验室)
;
Information Support Force Engineering University(信息支援力量工程大学)
CommentsPreprint. Accepted at NeurIPS 2025 Workshops on SPACE in Vision, Language, and Embodied AI (SpaVLE) as Oral, Embodied World Models for Decision Making (EWM), Aligning Reinforcement Learning Experimentalists and Theorists (ARLET), and Scaling Environments for Agents (SEA)
Ying Shen, Zhiyang Xu, Jiuhai Chen, Shizhe Diao, Jiaxin Zhang, Yuguang Yao, Joy Rimchala, Ismini Lourentzou, Lifu Huang
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Maryland(马里兰大学)
;
Nvidia(英伟达)
;
Salesforce AI Research(Salesforce AI研究)
;
Intuit AI Research(Intuit AI研究)
专题命中
图文多模态
:multimodal(abstract,comments);multimodal foundation model(abstract);分类 cs.CV
Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
用于3D理解的参数高效CLIP适配:通过统一分词实现
Guofeng Mei, Qinfeng Xiao, Bin Ren, Luigi Riz, Juan Liu, Xiaoshui Huang, Xu Zheng, Nicu Sebe, Ming-Hsuan Yang, Fabio Poiesi
机构
*
Fondazione Bruno Kessler(布鲁诺·科塞拉基金会)
;
University of Trento(特伦托大学)
;
University of Pisa(比萨大学)
;
Beijing Forestry University(北京林业大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Hong Kong University of Science and Technology (GZ)(香港科技大学)
;
Shandong University(山东大学)
;
University of California, Merced(加州大学默塞德分校)
机构
*
Nanjing University of Science and Technology(南京理工大学)
;
National University of Singapore(新加坡国立大学)
;
Beihang University(北京航空航天大学)
;
Nanjing Forestry University(南京林业大学)
Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?
贫困地图绘制的柏拉图式表示:统一的视觉语言代码还是智能体诱导的新颖性?
Satiyabooshan Murugaboopathy, Connor T. Jerzak, Adel Daoud
机构
*
AI and Global Development Lab(人工智能与全球发展实验室)
;
Institute for Analytical Sociology(分析社会学研究所)
;
Institute of Computer Science(计算机科学研究所)
;
Department of Government(政府系)
机构
*
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences (UCAS)(中国科学院大学先进交叉学科学院)
;
School of Electronic, Electrical and Communication Engineering, UCAS(中国科学院大学电子、电气与通信工程学院)
;
School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)
Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
计算机视觉中的推理:分类、模型、任务与方法论
Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris
机构
*
Department of Computer Science and Engineering, Birbhum Institute of Engineering and Technology(计算机科学与工程系,比罗尔理工学院)
;
College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)
;
Faculty of Computer Science and Information Technology, Universiti Malaya(计算机科学与信息技术学院,马来亚大学)
Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT
问题真的重要吗?视觉-语言SFT的无训练数据选择
Peng Sun, Yi Yang, Huawen Shen, Yi Ban, Tianfan Fu, Yanbo Wang, Yuqiang Li
机构
*
Nanjing University(南京大学)
;
Institute of Information Engineering(信息工程研究所)
;
North University of China(中国北方大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning
CityRiSE:基于强化学习的大视觉语言模型城市社会经济地位推理框架
Tianhui Liu, Hetian Pang, Xin Zhang, Jie Feng, Pan Hui, Yong Li
机构
*
Information Hub, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心)
;
Department of Electronic Engineering, BNRist, Tsinghua University(清华大学电子工程系)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang
机构
*
Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国)
;
School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国)
;
State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)