Mohammad Anas Azeez, Ankan Deria, Zohaib Hasan Siddiqui, Adinath Madhavrao Dukre, Rafiq Ali, Sara Atito, Yutong Xie, Imran Razzak
机构
*
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE(穆罕默德·本·扎耶德人工智能大学,阿布扎比,阿联酋)
;
King Fahd University of Petroleum and Minerals, Dhahran, Saudi Arabia(法赫德国王石油矿产大学,达兰,沙特阿拉伯)
;
Macquarie University, Sydney, Australia(麦考瑞大学,悉尼,澳大利亚)
;
University of Surrey, Guildford, United Kingdom(萨里大学,吉尔福德,英国)
;
University of New South Wales, Sydney, Australia(新南威尔士大学,悉尼,澳大利亚)
专题命中
幻觉与鲁棒性
:grounding(title);multimodal large language model(abstract);分类 cs.CV
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
让 MLLMs 失明:MLLM 内容审核中的对抗走私攻击
Zhiheng Li, Zongyang Ma, Yuntong Pan, Ziqi Zhang, Xiaolei Lv, Bo Li, Jun Gao, Jianing Zhang, Chunfeng Yuan, Bing Li, Weiming Hu
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(北京市多模态信息超智能安全重点实验室)
;
Hellogroup
;
University of Washington(华盛顿大学)
;
Jilin University(吉林大学)
;
ShanghaiTech University(上海科技大学)
专题命中
幻觉与鲁棒性
:MLLM(title);multimodal large language model(abstract);分类 cs.CV
A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models
一种可证明的能量引导测试时间防御提升大视觉-语言模型的对抗鲁棒性
Mujtaba Hussain Mirza, Antonio D'Orazio, Odelia Melamed, Iacopo Masi
机构
*
OmnAI Lab, Computer Science Department, Sapienza University of Rome, Italy(罗马大学计算机科学系OmnAI实验室,意大利)
;
Weizmann Institute of Science, Israel(魏茨曼科学研究所,以色列)
Overthinking Causes Hallucination: Tracing Confounder Propagation in Vision Language Models
过度思考导致幻觉:追踪视觉语言模型中的混杂因素传播
Abin Shoby, Ta Duc Huy, Tuan Dung Nguyen, Minh Khoi Ho, Qi Chen, Anton van den Hengel, Phi Le Nguyen, Johan W. Verjans, Vu Minh Hieu Phan
机构
*
Australian Institute for Machine Learning, University of Adelaide(阿德莱德大学澳大利亚机器学习研究所)
;
Hanoi University of Science and Technology(河内科技大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中
幻觉与鲁棒性
:vision language model(title,abstract);分类 cs.CV
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
SECOND:通过选择性和对比性解码缓解视觉-语言模型中的感知幻觉
Woohyeon Park, Woojin Kim, Jaeik Kim, Jaeyoung Do
机构
*
Department of Electrical and Computer Engineering, Seoul National University(首尔大学电气与计算机工程系)
;
Interdisciplinary Program in Artificial Intelligence, Seoul National University(首尔大学人工智能跨学科项目)
机构
*
School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
专题命中
幻觉与鲁棒性
:multimodal large language model(title,abstract);分类 cs.CV
机构
*
Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)
;
AIM Intelligence
;
National Institute of Technology Agartala(国立阿加尔塔拉理工学院)
;
Fordham University(福特汉姆大学)
;
University of Fukui(福井大学)
;
Independent Researcher(独立研究员)
机构
*
MoE Key Lab of BIPC, University of Science and Technology of China(北京信息科技大学MoE关键实验室,中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
机构
*
Tongji University(同济大学)
;
Computer Network Information Center, CAS(计算机网络信息中心,中国科学院)
;
HIAS, University of Chinese Academy of Sciences(高等研究所,中国科学院大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
Learning To Guide Human Decision Makers With Vision-Language Models
通过视觉-语言模型引导人类决策者
Debodeep Banerjee, Stefano Teso, Burcu Sayin, Andrea Passerini
机构
*
DI, University of Pisa(比萨大学DI)
;
DISI, University of Trento(特伦托大学DISI)
;
CIMEC, University of Trento(特伦托大学CIMEC)
;
Rajib Pal Bankura Sammilani Medical College(拉吉布·帕尔班克尔医学学院)
LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
LMOD+: 一个全面的多模态数据集和基准,用于开发和评估眼科中的多模态大语言模型
Zhenyue Qin, Yang Liu, Yu Yin, Jinyu Ding, Haoran Zhang, Anran Li, Dylan Campbell, Xuansheng Wu, Ke Zou, Tiarnan D. L. Keenan, Emily Y. Chew, Zhiyong Lu, Yih Chung Tham, Ninghao Liu, Xiuzhen Zhang, Qingyu Chen
机构
*
School of Medicine, Yale University(耶鲁大学医学院)
;
School of Computing, Australian National University(澳大利亚国立大学计算机学院)
;
School of Engineering, Imperial College London(伦敦帝国理工学院工程学院)
;
School of Computing, University of Georgia(佐治亚大学计算机学院)
;
Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨秀隆医学学院)
;
National Eye Institute, National Institutes of Health(美国国立卫生研究院眼科研究所)
;
National Library of Medicine, National Institutes of Health(美国国立卫生研究院国家医学图书馆)
;
School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算机技术学院)
专题命中
幻觉与鲁棒性
:multimodal large language model(title,abstract);分类 cs.CV
LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
LLaVAShield: 保障视觉语言模型中的多模态多轮对话安全
Guolei Huang, Qinzhi Peng, Gan Xu, Yao Huang, Yuxuan Lu, Yongjun Shen
机构
*
Southeast University(东南大学)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Zhejiang University of Technology(浙江工业大学)
;
Tsinghua University(清华大学)
;
RealAI
机构
*
Dalian Maritime University(大连海事大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Tsinghua University(清华大学)
;
Nanyang Technological University(南洋理工大学)
;
Xi'an Jiaotong University(西安交通大学)
;
Renmin University of China(中国人民大学)
;
Wuhan University(武汉大学)
专题命中
幻觉与鲁棒性
:MLLM(title);multimodal large language model(abstract);分类 cs.AI
机构
*
MoE Key Lab of BIPC, University of Science and Technology of China(摩埃关键实验室,中国科学技术大学)
;
Nanyang Technological University(南洋理工大学)
;
National University of Singapore(新加坡国立大学)
;
Tianjin University(天津大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(title);vision-language model(abstract);分类 cs.CV
Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification
通过证据不确定性量化检测大视觉-语言模型的误行
Tao Huang, Rui Wang, Xiaofei Liu, Yi Qin, Li Duan, Liping Jing
机构
*
State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室)
;
Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室)
;
School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)
;
School of Automation and Intelligence, Beijing Jiaotong University(北京交通大学自动化与智能学院)
;
Beijing Key Laboratory of Security and Privacy in Intelligent Transportation(北京智能交通安全与隐私重点实验室)