机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
;
College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
Agent-X:评估视觉中心智能体任务中的深度多模态推理
Tajamul Ashraf, Amal Saqib, Hanan Ghani, Muhra AlMahri, Yuhao Li, Noor Ahsan, Umair Nawaz, Jean Lahoud, Hisham Cholakkal, Mubarak Shah, Philip Torr, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学)
;
University of Central Florida(中央佛罗里达大学)
;
University of Oxford(牛津大学)
G-DRAGON: Geospatial Reasoning and Dynamic Planning for Retrieval-Augmented Outdoor Navigation
G-DRAGON:面向检索增强的户外导航的地理空间推理与动态规划
Dongzhihan Wang, Yi Du, Jianan Sun, Yuan Xue, Yingchen Zhang, Bing Xiao, Chen Wang, Liang Xu
机构
*
Spatial AI & Robotics Lab(空间人工智能与机器人实验室)
;
University at Buffalo(布法罗大学)
;
School of Future Technology(未来技术学院)
;
Shanghai University(上海大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
GeoMathCode: 理解几何问题求解中交织的数学-代码推理
Yingji Zhang, Yong Dai, André Freitas
机构
*
Idiap Research Institute(Idiap研究 institute)
;
X-Humanoid
;
Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系)
;
Cancer Biomarker Centre, CRUK Manchester Institute(癌症生物标志物中心,CRUK曼彻斯特研究所)
专题命中
视觉推理
:multimodal large language model(abstract)
机构
*
School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
;
National Institute of Health Data Science, Peking University(北京大学健康数据科学国家研究院)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Tianjin Institute of Cardiology, the Second Hospital of Tianjin Medical University(天津医科大学第二医院心内科)
;
National University of Singapore(新加坡国立大学)
;
Jarvis Lab, Tencent(腾讯 Jarvis实验室)
;
HeartVoice Medical Technology(HeartVoice医疗科技)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract)
Language Movement Primitives: Grounding Language Models in Robot Motion
语言运动基元:将语言模型锚定在机器人运动中
Yinlong Dai, Benjamin A. Christie, Daniel J. Evans, Dylan P. Losey, Simon Stepputtis
机构
*
Collab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(合作组,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
;
TEA Lab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(TEA实验室,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos
你可以比看见更早定位:一种用于压缩视频中时序句子定位的高效流程
Xiang Fang, Daizong Liu, Pan Zhou, Guoshun Nan
机构
*
The Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,网络安全科学与工程学院,华中科技大学)
;
Peking University(北京大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
层次化局部-全局Transformer用于时间语句定位
Xiang Fang, Daizong Liu, Pan Zhou, Zichuan Xu, Ruixuan Li
机构
*
Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,华中科技大学网络安全科学与工程学院)
;
Wangxuan Institute of Computer fTechnology, Peking University(王宣计算机技术研究院,北京大学)
;
School of software, Dalian University of Technology(软件学院,大连理工大学)
;
School of Computer Science, and Technology, Huazhong University of Science, and Technology(计算机科学与技术学院,华中科技大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Third Affiliated Hospital of Sun Yat-Sen University(中山大学第三附属医院)
;
Hong Kong Metropolitan University(香港 Metropolitan 大学)