机构
*
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学交叉科学学院)
;
Tsinghua University(清华大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation
Fly0: 解耦语义定位与几何规划以实现零样本空中导航
Zhenxing Xu, Yihong Lu, Weidong Bao, Zhengqiu Zhu, Jingxuan Zhou, Zhichuang Wang, Ji Wang, Lihua Liu, Wei He
机构
*
National Key Laboratory of Big Data and Decision(大数据与决策国家重点实验室)
;
National University of Defense Technology(国防科技大学)
;
State Key Laboratory of Digital Intelligent Modeling and Simulation(数字智能建模与仿真国家重点实验室)
;
Information Support Force Engineering University(信息支援力量工程大学)
专题命中
GUI与屏幕智能体
:MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.AI
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
SleepWalk:一种三层压力测试基准,用于指导的视觉-语言导航
Niyati Rawal, Sushant Ravva, Shah Alam Abir, Saksham Jain, Aman Chadha, Vinija Jain, Suranjana Trivedy, Amitava Das
机构
*
Indian AI Research Organization (IAIRO)(印度人工智能研究组织)
;
ĸragya Lab, BITS Pilani Goa(BITS Pilani Goa 的 ĸragya 实验室)
;
University of Dhaka(达卡大学)
;
Delhi Technological University(德里技术大学)
;
Apple(苹果公司)
;
Meta
Grounding Computer Use Agents on Human Demonstrations
基于人类演示的计算机使用智能体基础构建
Aarash Feizi, Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Kaixin Li, Rabiul Awal, Xing Han Lù, Johan Obando-Ceron, Juan A. Rodriguez, Nicolas Chapados, David Vazquez, Adriana Romero-Soriano, Reihaneh Rabbany, Perouz Taslakian, Christopher Pal, Spandana Gella, Sai Rajeswar
机构
*
Mila - Quebec AI Institute(魁北克AI研究所)
;
McGill University(麦吉尔大学)
;
Université de Montréal(蒙特利尔大学)
;
ServiceNow Research(ServiceNow研究)
;
University of Waterloo(滑铁卢大学)
;
University of Oxford(牛津大学)
;
National University of Singapore(新加坡国立大学)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
École de Technologie Supérieure(高级技术学院)
;
CIFAR AI Chair(CIFAR人工智能主席)
机构
*
AnnLab(安实验室)
;
Institute of Semiconductors, Chinese Academy of Sciences(中国科学院半导体研究所)
;
Zhongguancun Academy(中关村学院)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
State Key Laboratory of High Performance Ceramics(高性能陶瓷国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Electronic, Electrical and Communication Engineering(电子电气与通信工程学院)
;
University of ChineseAcademy of Sciences(中国科学院大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(title,abstract);分类 cs.AI
机构
*
Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室,数字工程中心,斯克尔科沃科学与技术研究所)
机构
*
Tsinghua University(清华大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司人工智能研究院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Peking University(北京大学)
专题命中
GUI与屏幕智能体
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract)
SCLARO: A Dataset for Grounded Scenario-Level Scene Understanding and ScenarioCLIP for Benchmarking
SCLARO:面向场景级场景理解的数据集及用于基准测试的ScenarioCLIP
Advik Sinha, Saurabh Atreya, Aashutosh A, Sk Aziz Ali, Abhijit Das
机构
*
Machine Intelligence Group, Department of CS&IS, Birla Institute of Technology and Science, Pilani – Hyderabad Campus(人工智能组,计算机科学与信息系,比拉理工学院和科学学院,比拉-海得拉巴校区)
专题命中
GUI与屏幕智能体
:visual language model(title);分类 cs.CV
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系)
;
Communication University of China(中国传媒大学)
;
Imperial College London(伦敦帝国学院)
;
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区网络科学与技术学院)
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
GUIDE:通过实时网络视频检索和即插即用标注解决GUI代理的领域偏见
Rui Xie, Zhi Gao, Chenrui Shi, Zirui Shang, Lu Chen, Qing Li
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
State Key Laboratory for General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,北京通用人工智能研究院)
;
Beijing Institute of Technology(北京理工大学)
CommentsAccepted to ECCV 2026. 30 pages: 15-page main paper followed by supplementary material as an appendix (Sections A-F). Project page: https://sharryXR.github.io/GUIDE/
NormAct: Benchmarking Embodied Agents' Proactive Compliance with Unspoken Social Norms
NormAct:具身规划中隐藏社会规范遵守的基准
Shiyun Zhao, Xinwei Song, Tianyu Guo, Xiaomeng Gao, Mingyuan Liu, Xu Han, Yuanyuan Zhang, Zhenliang Zhang, Xue Feng, Bo Dai
机构
*
State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI)(通用人工智能国家重点实验室,北京通用人工智能研究院)
;
China Academy of Information and Communications Technology(中国信息通信研究院)
;
ShanghaiTech University(上海科技大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI