Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
高质量文本,稳健视觉:语言在增强视觉语言模型视觉稳健性中的作用
Futa Waseda, Saku Sugawara, Isao Echizen
机构
*
The University of Tokyo(东京大学)
;
National Institute of Informatics(日本信息处理研究所)
;
The University of Tokyo, National Institute of Informatics(东京大学、日本信息处理研究所)
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
SAVER:通过风格感知视觉早期修正减轻大型视觉语言模型中的幻觉
Zhaoxu Li, Chenqi Kong, Yi Yu, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Kot, Xudong Jiang
机构
*
ROSE Lab, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(南洋理工大学罗思实验室,跨学科研究生项目,新加坡)
;
ROSE Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学罗思实验室,电子与电气工程学院,新加坡)
;
City University of Hong Kong, Hong Kong SAR(香港城市大学,香港特别行政区)
;
Shanghai Jiao Tong University, China(上海交通大学,中国)
;
Singapore University of Technology and Design, Singapore(新加坡科技设计大学,新加坡)
;
VinUniversity, Hanoi, Vietnam(越南文大学,河内,越南)
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
MELLA:弥合低资源语言多模态大语言模型的语言能力与文化根基
Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen, Guohang Yan, Yunshi Lan, Botian Shi
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
East China Normal University(东华大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Institute of High Performance Computing, A*STAR(高性能计算研究所,A*STAR)
专题命中
幻觉与鲁棒性
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
Department of Electrical and Computer Engineering, George Washington University(电气与计算机工程系,乔治华盛顿大学)
;
Department of Mechanical and Aerospace Engineering, George Washington University(机械与航空航天工程系,乔治华盛顿大学)
;
Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学)
机构
*
The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳))
;
School of Data Science, School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(数据科学学院、人工智能学院、香港中文大学(深圳))
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI
机构
*
Mohamed Bin Zayed University of AI, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)
;
Khalifa University, UAE(哈利法大学,阿联酋)
;
Australian National University, Australia(澳大利亚国立大学,澳大利亚)
An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
聚焦女性安全分析:多模态数据集能否增强VAD模型?
Sangeeta ., Maddikuntla Sai Prajwal, Debi Prosad Dogra, Kamalakar Vijay Thakare, Hyungjoo Jung, Ig-Jae Kim, Heeseung Choi
机构
*
Indian Institute of Technology Bhubaneswar(印度理工学院巴特那分校)
;
Artificial Intelligence and Robotics Institute, Korea Institute of Science and Technology(人工智能与机器人研究所,韩国科学技术院)
;
Yonsei-KIST Convergence Research Institute, Yonsei University(延世大学KIST融合研究中心)
机构
*
GigaAI(字节跳动人工智能实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
University of Leeds(利兹大学)
;
Cornell University(康奈尔大学)
;
Tsinghua University(清华大学)
;
Nanjing University of Science and Technology(南京理工大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
FAWTD(一汽技术开发部)