机构
*
Mohamed Bin Zayed University of AI, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)
;
Khalifa University, UAE(卡比拉大学,阿联酋)
;
Australian National University, Australia(澳大利亚国立大学,澳大利亚)
专题命中
幻觉与鲁棒性
:MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
TSHA:用于可信安全危害评估场景的视觉语言模型基准
Qiucheng Yu, Ruijie Xu, Mingang Chen, Jianfeng Dong, Xin Tan
机构
*
City University of Hong Kong(香港城市大学)
;
East China Normal University(华东师范大学)
;
University of Western Australia(西澳大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai Development Center of Computer Software Technology(上海计算机软件技术开发中心)
;
Zhejiang Gongshang University(浙江工商大学)
专题命中
幻觉与鲁棒性
:visual language model(title);vision-language model(abstract);分类 cs.CV、cs.AI
Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo
机构
*
Cognitive Science Program, University of California, Berkeley(加州大学伯克利分校认知科学项目)
;
Department of Electrical and Computer Engineering, University of California San Diego(加州大学圣地亚哥分校电气与计算机工程系)
;
School of Computer Science, Georgia Institute of Technology & Emory University(佐治亚理工学院计算机科学学院及埃默里大学)
;
Department of Computer Science, Johns Hopkins University(约翰霍普金斯大学计算机科学系)
;
Department of Cognitive Science, University of California San Diego(加州大学圣地亚哥分校认知科学系)
;
Equal Advising Department of Computer Science & Wilmer Eye Institute, Johns Hopkins University(约翰霍普金斯大学计算机科学系及威尔默眼科研究所)
;
Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)
;
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
;
Weinberg Institute for Cognitive Science, University of Michigan(密歇根大学韦恩伯格认知科学研究所)
机构
*
Mohamed Bin Zayed University of AI, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)
;
Khalifa University, UAE(卡布斯大学,阿联酋)
;
Australian National University, Australia(澳大利亚国立大学,澳大利亚)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
高质量文本,稳健视觉:语言在增强视觉语言模型视觉稳健性中的作用
Futa Waseda, Saku Sugawara, Isao Echizen
机构
*
The University of Tokyo(东京大学)
;
National Institute of Informatics(日本信息处理研究所)
;
The University of Tokyo, National Institute of Informatics(东京大学、日本信息处理研究所)
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
SAVER:通过风格感知视觉早期修正减轻大型视觉语言模型中的幻觉
Zhaoxu Li, Chenqi Kong, Yi Yu, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Kot, Xudong Jiang
机构
*
ROSE Lab, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(南洋理工大学罗思实验室,跨学科研究生项目,新加坡)
;
ROSE Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学罗思实验室,电子与电气工程学院,新加坡)
;
City University of Hong Kong, Hong Kong SAR(香港城市大学,香港特别行政区)
;
Shanghai Jiao Tong University, China(上海交通大学,中国)
;
Singapore University of Technology and Design, Singapore(新加坡科技设计大学,新加坡)
;
VinUniversity, Hanoi, Vietnam(越南文大学,河内,越南)
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
MELLA:弥合低资源语言多模态大语言模型的语言能力与文化根基
Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen, Guohang Yan, Yunshi Lan, Botian Shi
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
East China Normal University(东华大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Institute of High Performance Computing, A*STAR(高性能计算研究所,A*STAR)
专题命中
幻觉与鲁棒性
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
Department of Electrical and Computer Engineering, George Washington University(电气与计算机工程系,乔治华盛顿大学)
;
Department of Mechanical and Aerospace Engineering, George Washington University(机械与航空航天工程系,乔治华盛顿大学)
;
Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学)