Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
跨模态注意力校准用于LVLM幻觉缓解
Jiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma, Guanbin Li
机构
*
Sun Yat-sen University(中山大学)
;
The University of Hong Kong(香港大学)
;
Meituan(美团)
;
Inspur Database Technology(Inspur数据库技术)
;
Guilin University of Electronic Technology(桂林电子科技大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
TRACER: Persistent Regularization for Robust Multimodal Finetuning
TRACER: 用于鲁棒多模态微调的持久正则化
Hesam Asadollahzadeh, Feng Liu, Christopher Leckie, Sarah M. Erfani
机构
*
School of Computing and Information Systems (CIS), Faculty of Engineering and IT (FEIT), University of Melbourne, Australia(墨尔本大学计算机科学与信息系统学院(CIS)、工程与信息技术学院(FEIT))
Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers
受神经科学启发的多模态Transformer中视觉趣味性分析
Mathis Immertreu, Fitim Abdullahu, Thomas Kinfe, Helmut Grabner, Patrick Krauss, Achim Schilling
机构
*
Cognitive Computational Neuroscience Group, Patter Recognition Lab, University Erlangen-Nürnberg(认知计算神经科学组、模式识别实验室、埃尔兰根-纽伦堡大学)
;
IDS Institut für Data Science, ZHAW School of Engineering, Winterthur, Switzerland(IDS数据科学研究所、ZHAW工程学院、温特图尔,瑞士)
;
Mannheim Center for Neuromodulation and Neuroprosthetics, University Hospital Mannheim, Heidelberg University(曼海姆神经调制与神经假体中心、曼海姆大学医院、海德堡大学)
;
BGU Ludwigshafen, Germany(BGU路易斯港,德国)
;
Physics and Cognition Group, MCNN, University Hospital Mannheim, Heidelberg University(物理与认知组、MCNN、曼海姆大学医院、海德堡大学)
;
NeuroAI and BCI Group, MCNN, University Hospital Mannheim, Heidelberg University(神经AI与BCI组、MCNN、曼海姆大学医院、海德堡大学)
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
AdaMMS: 为异构多模态大语言模型设计的模型融合方法
Yiyang Du, Xiaochen Wang, Chi Chen, Jiabo Ye, Yiru Wang, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Zhifang Sui, Maosong Sun, Yang Liu
机构
*
Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学)
;
Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学)
;
State Key Laboratory of Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)
;
School of Software Microelectronics, Peking University(软件微电子学院,北京大学)
;
Institute of Intelligent Computing, Alibaba Group(智能计算研究院,阿里巴巴集团)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)
;
Jiangsu Collaborative Innovation Center for Language Competence, Jiangsu, China(江苏省语言能力协同创新中心,江苏,中国)
;
ModelTC Open Source Organization, Beijing, China(ModelTC开源组织,北京,中国)
Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models
重新审视遥感中的变化视觉问答问题:结构化与原生多模态Qwen模型
Yakoub Bazi, Mohamad M. Al Rahhal, Mansour Zuair, Faroun Mohamed
机构
*
Computer Engineering Department, College of Computer and Information Sciences, King Saud University(计算机工程系,计算机与信息科学学院,沙特国王大学)
;
Applied Computer Science Department, College of Applied Computer Science, King Saud University(应用计算机科学系,应用计算机科学学院,沙特国王大学)
Joycelyn Teo, Rui Cao, Zhenyun Deng, Zifeng Ding, Michael Sejr Schlichtkrull, Andreas Vlachos
机构
*
Defence Science and Technology Agency, Singapore(新加坡国防科学与技术局)
;
University of Cambridge, UK(英国剑桥大学)
;
Queen Mary University of London, UK(英国伦敦大学学院)
Automatic Modeling of Social Concepts Evoked by Art Images as Multimodal Frames
艺术图像中引发的社会概念的自动建模作为多模态框架
Delfina Sol Martinez Pandiani, Valentina Presutti
机构
*
Department of Computer Science and Engineering (DISI), University of Bologna, Italy(博洛尼亚大学计算机科学与工程系(DISI))
;
Department of Modern Languages, Literatures and Cultures (LILEC), University of Bologna, Italy(博洛尼亚大学现代语言、文学与文化系(LILEC))
Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
见仁见智:在标签噪声下鲁棒的视觉引导跨模态提示学习
Zibin Geng, Xuefeng Jiang, Jia Li, Zheng Li, Tian Wen, Lvhua Wu, Sheng Sun, Yuwei Wang, Min Liu
机构
*
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
PCALab, VCIP, College of Computer Science, Nankai University(南开大学计算机学院PCALab, VCIP)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
面向视觉语言模型的层次化通用多模态攻击
Peng-Fei Zhang, Zi Huang
机构
*
School of Electrical Engineering and Computer Science, the University of Queensland(电气工程与计算机科学学院,昆士兰大学)
;
Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学)