机构
*
Southern University of Science and Technology(南方科技大学)
;
City University of Hong Kong(香港城市大学)
;
Hong Kong University of Science and Technology(香港科技大学)
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
视觉语言模型(VLMs)在失明或被误导时的表现如何?针对科学图表的VLMs行为评估
Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
机构
*
University of Aberdeen(阿伯丁大学)
;
International Institute of Information Technology Hyderabad(海德拉巴国际信息技术学院)
;
University of Technology Nuremberg(纽伦堡工业大学)
;
University of Southern California(南加利福尼亚大学)
机构
*
Remote Sensing Lab, National Technical University of Athens(遥感实验室,国家技术大学雅典分校)
;
Remote Sensing Lab, Department of Engineering and Sciences, National Technical University of Athens(遥感实验室,工程与科学系,国家技术大学雅典分校)
;
Universitas Mercatorum(默卡托大学)
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
手术内窥镜视频的时间视觉语言模型的鲁棒性研究
Darakshan Rashid, Raza Imam, Ufaq Khan, Muhammad Bilal, Shazad Ashraf, Dwarikanath Mahapatra, Mohammad Yaqub, Muhammad Haris Khan, Imran Razzak, Brejesh Lall, Lena Maier-Hein, Yutong Xie
机构
*
Indian Institute of Technology Delhi(印度理工学院德里分校)
;
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Birmingham City University(伯明翰城市大学)
;
University Hospitals Birmingham(伯明翰大学医院)
;
Khalifa University(哈利法大学)
;
German Cancer Research Center (DKFZ)(德国癌症研究中心)
Test-Time Hallucination Control in Large Vision-Language Models
大型视觉语言模型的测试时幻觉控制
Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian, Mohammadreza Alimoradijazi, Aijun An, Hossein Rahmani
机构
*
Australian National University(澳大利亚国立大学)
;
The University of New South Wales(新南威尔士大学)
;
University of Technology Sydney(悉尼科技大学)
;
York University(约克大学)
;
Lancaster University(兰卡斯特大学)
Local Margin Restoration for Test-Time Adaptation of Vision-Language Models
用于视觉-语言模型测试时适应的局部间隔恢复
Yan Huang, Guowei Wang, Xu Wang, Kangjun Liu, Xin Lin
机构
*
Guangzhou University(广州大学)
;
The Second Affiliated Hospital of Guangzhou University of Chinese Medicine(广州中医药大学第二附属医院)
;
Jinan University(暨南大学)
;
Pengcheng Laboratory(鹏城实验室)
Unifying Adversarially Robust Model Experts in Vision-Language Models
统一视觉-语言模型中的对抗鲁棒模型专家
Nguyen Duc Thai, Junhao Dong, Sua Qi Rong, Hua Yu, Yew-Soon Ong
机构
*
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
Center for Frontier AI Research, Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局前沿人工智能研究中心)
机构
*
School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间安全学院)
;
College of Computer Science, Chongqing University(重庆大学计算机科学学院)
;
School of Software and engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)
专题命中
幻觉与鲁棒性
:vision-language model(title);vision language model(abstract);分类 cs.CV
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling
面向多模态大语言模型的长视频快速有效理解:自适应准高斯采样
Kun Zhang, Chenxin Fang, Tao Chen, Baiyang Song, Yunhang Shen, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
专题命中
幻觉与鲁棒性
:multimodal large language model(title,abstract);分类 cs.CV