ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
ENC-Bench:用于评估多模态大语言模型在电子航海图理解中的基准
Ao Cheng, Xingming Li, Xuanyu Ji, Xixiang He, Qiyao Sun, Chunping Qiu, Runke Huang, Qingyong Hu
机构
*
National University of Defense Technology(国防科技大学)
;
Intelligent Game and Decision Lab(智能游戏与决策实验室)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);InternVL(abstract);grounding(abstract);分类 cs.CV
EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
EagleVision:基于BEV的双阶段框架用于空间智能
Jiaxu Wan, Xu Wang, Mengwei Xie, Hang Zhang, Mu Xu, Yang Han, Hong Zhang, Ding Yuan, Yifan Yang
机构
*
Atlas Lab(Atlas实验室)
;
School of Aerospace, BUAA(北京航空航天大学航空学院)
;
School of Software, BUAA(北京航空航天大学软件学院)
;
State Key Laboratory of HERATT(HERATT国家重点实验室)
;
Key Laboratory of SDODS (MOE) Project(SDODS重点实验室(教育部))
专题命中
视觉定位与Grounding
:grounding(title,abstract);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院)
;
School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)
;
CSIRO(澳大利亚联邦科学与工业研究组织)
;
Northwestern Polytechnical University(西北工业大学)
;
University of Macau(澳门大学)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
基于多模态大语言模型的零样本人-物交互检测
Shiyu Xuan, Dongkai Wang, Zechao Li, Jinhui Tang
机构
*
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics(西南财经大学计算机与人工智能学院)
;
Nanjing Forestry University(南京林业大学)
机构
*
School of Artificial Intelligence, Tianjin University(天津大学人工智能学院)
;
Low-Altitude Intelligence Lab, Xiong’an National Innovation Center(雄安国家创新中心低空智能实验室)
;
Xiong’an Guochuang Lantian Technology Co., Ltd.(雄安国创莲田科技有限公司)
;
College of Electronic Science and Technology, National University of Defense Technology(国防科技大学电子科学学院)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构
*
Harbin Institute of Technology(哈尔滨工业大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Xidian University(西安电子科技大学)
;
East China Normal University(东华大学)
;
Wuhan University(武汉大学)
;
Southeast University(东南大学)
;
National University of Defense Technology(国防科技大学)
;
Chinese University of Hong Kong(香港中文大学)
专题命中
视觉定位与Grounding
:vision language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV
机构
*
ByteDance Intelligent Creation(字节跳动智能创作)
;
Tsinghua University(清华大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AGO: Adaptive Grounding for Open World 3D Occupancy Prediction
Peizheng Li, Shuxiao Ding, You Zhou, Qingwen Zhang, Onat Inak, Larissa Triess, Niklas Hanselmann, Marius Cordts, Andreas Zell
机构
*
Mercedes-Benz AG(梅赛德斯-奔驰集团)
;
University of Tübingen(图宾根大学)
;
Tübingen AI Center(图宾根人工智能中心)
;
University of Bonn(波恩大学)
;
RPL
;
KTH Royal Institute of Technology(皇家理工学院)
;
TU Berlin(柏林技术大学)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
Fanxiao Li, Jiaying Wu, Canyuan He, Wei Zhou
机构
*
School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院)
;
National University of Singapore(新加坡国立大学)
;
Engineering Research Center of Cyberspace, Yunnan University(云南大学网络空间研究院)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
Bastian Pätzold, Jan Nogga, Sven Behnke
机构
*
Autonomous Intelligent Systems, University of Bonn(博恩大学自主智能系统中心)
;
Lamarr Institute for Machine Learning and AI(拉马尔人工智能与机器学习研究所)
;
Center for Robotics, University of Bonn(博恩大学机器人中心)
Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
Lihua Zhou, Mao Ye, Shuaifeng Li, Nianxin Li, Jinlin Wu, Xiatian Zhu, Lei Deng, Hongbin Liu, Jiebo Luo, Zhen Lei
机构
*
Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong, China(人工智能与机器人研究中心,香港科学与创新研究院,中国科学院,香港,中国)
;
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科学与技术大学计算机科学与工程学院)
;
Surrey Institute for People-Centred Artificial Intelligence, CVSSP, University of Surrey(以人为中心的人工智能研究院,CVSSP, Surrey大学)
;
School of Electronics and Information Engineering, Shenzhen University(电子与信息工程学院,深圳大学)
;
University of Rochester(罗切斯特大学)
机构
*
University of California, San Diego(加州大学圣地亚哥分校)
;
ByteDance(字节跳动)
;
University of California, Merced(加州大学默塞德分校)
;
University of Southern California(南加州大学)
;
University at Buffalo(布法罗大学)
;
The University of Queensland(昆士兰大学)
Uncertainty-aware Medical Diagnostic Phrase Identification and Grounding
Ke Zou, Yang Bai, Bo Liu, Yidi Chen, Zhihao Chen, Yang Zhou, Xuedong Yuan, Meng Wang, Xiaojing Shen, Xiaochun Cao, Yih Chung Tham, Huazhu Fu
机构
*
College of Computer Science, Sichuan University(四川大学计算机学院)
;
Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所)
;
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)
;
Department of Radiology, West China Hospital, Sichuan University(四川大学华西医院放射科)
;
College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院)
;
Department of Mathematics, Sichuan University(四川大学数学系)
;
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院)
;
Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore and the Singapore Eye Research Institute, Singapore National Eye Centre(新加坡国立大学 Yong Loo Lin 医学院眼科系及新加坡眼科研所、新加坡国家眼科中心)
专题命中
视觉定位与Grounding
:grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV