机构
*
SKLCCSE Lab Beihang University Beijing China(SKLCCSE实验室 北航)
;
Department of Data Science City University of Hong Kong Hong Kong China(数据科学系 香港城市大学)
;
Beijing Advanced Innovation Center Beihang University Beijing China(北京先进创新中心 北航)
;
Beihang University(北航)
;
City University of Hong Kong(香港城市大学)
机构
*
Tsinghua University(清华大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司人工智能研究院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Peking University(北京大学)
专题命中
GUI与屏幕智能体
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract)
机构
*
College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
;
Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院智能信息处理上海市重点实验室)
;
TeleAI, China Telecom(中国电信天翼人工智能公司)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang
机构
*
Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国)
;
School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国)
;
State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
机构
*
School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间安全学院)
;
College of Computer Science, Chongqing University(重庆大学计算机科学学院)
;
School of Software and engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)
专题命中
幻觉与鲁棒性
:vision-language model(title);vision language model(abstract);分类 cs.CV
Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
在域转移下对用于乳腺钼靶成像的基础模型的稳健性进行基准测试
Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker
机构
*
College of Engineering and Computer Science, VinUniversity(工程与计算机科学学院,文大大学)
;
Imperial College London(伦敦帝国理工学院)
;
Radiology Department, Vietnam National Cancer Hospital(越南国家癌症医院放射科)
;
VinUni-Illinois Smart Health Center, VinUniversity(文大大学 - 伊利诺伊智能健康中心,文大大学)
;
The Computer Vision and Medical AI Lab, VinUniversity(计算机视觉与医学人工智能实验室,文大大学)
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
SpaceDrive: 在基于视觉语言模型的自动驾驶中引入空间感知
Peizheng Li, Zhenghao Zhang, David Holtz, Hang Yu, Yutong Yang, Yuzhi Lai, Rui Song, Andreas Geiger, Andreas Zell
机构
*
Mercedes-Benz AG(梅赛德斯-奔驰集团)
;
University of Tübingen(图宾根大学)
;
Tübingen AI Center(图宾根人工智能中心)
;
TU Munich(慕尼黑工业大学)
;
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
University of Stuttgart(斯图加特大学)
;
UCLA(加州大学洛杉矶分校)
专题命中
VLM训练与架构
:VLM(title,summary_cn);vision language model(abstract);分类 cs.CV