机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院)
;
Pengcheng Laboratory(鹏城实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
专题命中
视觉推理
:multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV
机构
*
University of California, San Diego(加州大学圣地亚哥分校)
;
ByteDance(字节跳动)
;
University of California, Merced(加州大学默塞德分校)
;
University of Southern California(南加州大学)
;
University at Buffalo(布法罗大学)
;
The University of Queensland(昆士兰大学)
专题命中
视觉推理
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
CommentsPublished in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-industry.103
Journal refProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) ACL 2025 1457-1465
机构
*
School of Information Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)信息科学与技术学院)
;
Guangdong Provincial Key Laboratory of Space-Aerial Networking and Intelligent Sensing(广东省空间-航空网络与智能感知重点实验室)
;
Information Systems Technology and Design, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计)
;
School of System Design and Intelligent Manufacturing, Southern University of Science and Technology(南方科技大学系统设计与智能制造学院)
专题命中
视觉推理
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.LG
CommentsThe paper has been submitted to IEEE Internet of Things Magazine
Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
Fengze Yang, Bo Yu, Yang Zhou, Xuewen Luo, Zhengzhong Tu, Chenxi Liu
机构
*
Department of Civil & Environmental Engineering University of Utah(土木与环境工程系 犹他大学)
;
Zachry Department of Civil and Environmental Engineering Texas A&M University(扎克里系 土木与环境工程系 德克萨斯农工大学)
;
Department of Computer Science & Engineering Texas A&M University(计算机科学与工程系 德克萨斯农工大学)
专题命中
视觉推理
:vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI
机构
*
School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China(信息科学与工程学院,东华大学,上海,中国)
;
School of Computer Science, Fudan University, Shanghai, China(计算机科学学院,复旦大学,上海,中国)
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
Comments13 pages, 7 figures, published to ACL 2025
Journal refIn Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 620-632, 2025, Vienna, Austria
机构
*
University of Science and Technology of China(科学技术大学)
;
University of Adelaide(阿德莱德大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
专题命中
视觉推理
:visual language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI
Comments8 pages, 3 figures, this paper has been accepted by ACM MM 2025
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院)
;
Meta Reality Lab(Meta现实实验室)
;
Xi’an Jiao Tong University(西安交通大学)
;
National University of Singapore(新加坡国立大学)
;
The Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学系统枢纽部门(广州))
Tactical Decision for Multi-UGV Confrontation with a Vision-Language Model-Based Commander
Li Wang, Qizhen Wu, Lei Chen
机构
*
Yangtze Delta Region Academy(扬子江地区学院)
;
Beijing Institute of Technology(北京理工大学)
;
School of Automation(自动化学院)
;
Beihang University(北航大学)
;
Advanced Research Institute of Multidisciplinary Sciences(多学科科学高级研究机构)
专题命中
视觉推理
:vision-language model(title,abstract);vision language model(abstract);分类 cs.AI
机构
*
Technical Aspects of Multimodal Systems (TAMS), Department of Informatics, University of Hamburg(汉堡大学信息学院多模态系统技术部门)
;
Hon Hai Research Institute (HHRI)(鸿海研究机构)
;
Department of Electrical Engineering, National Taiwan University(台湾大学电子工程系)
;
Department of Computer Science and Technology, National Tsinghua University(清华大学计算机科学与技术系)
专题命中
视觉推理
:vision-language model(title,abstract);visual language model(abstract);分类 cs.CV
机构
*
Rikkyo University(立命馆大学)
;
University of Tokyo(东京大学)
;
ATR(ATR研究所)
;
Institute of Science Tokyo(东京科学研究院)
;
National Institute of Informatics(国家信息研究所)
;
NII LLMC(日本信息处理学会LLMC)
;
Sony Semiconductor Solutions(索尼半导体解决方案)
机构
*
Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
;
ARC Lab, Tencent PCG(腾讯PCG ARC实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
The University of Hong Kong(香港大学)
专题命中
视觉推理
:vision language model(title,abstract);multimodal large language model(abstract);分类 cs.AI