arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 2251 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 2251 篇

2508.08644 2025-08-13 cs.CV 57%

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation

Guiming Cao, Yuming Ou

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07989 2025-08-12 cs.CV cs.HC 57%

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

Xiantao Zhang

机构 * Beihang University(北航大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

Comments 9 pages, 3 figures, 2 tables. Accepted at CV4A11y, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07818 2025-08-12 cs.CV 57%

Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models

Chenyue Song, Chen Hui, Haiqi Zhu, Feng Jiang, Yachun Mi, Wei Zhang, Shaohui Liu

专题命中 幻觉与鲁棒性 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06623 2025-08-12 cs.CV 57%

ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification

Sihan Ma, Qiming Wu, Ruotong Jiang, Frank Burns

机构 * Inner Mongolia University of Science & Technology(内蒙古科技大学) Federal University of Rio de Janeiro(里约热内卢联邦大学)

专题命中 幻觉与鲁棒性 :LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05527 2025-08-08 cs.CV 57%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04567 2025-08-07 cs.CV cs.CL 57%

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

Yifan Li, Kun Zhou, Wayne Xin Zhao, Lei Fang, Ji-Rong Wen

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12591 2025-08-06 cs.CV cs.CL 57%

CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base

Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu, Shuai Zhao, Thong Nguyen, Anh Tuan Luu

机构 * Nanyang Technological University, Singapore(南洋理工大学) National University of Singapore, Singapore(国立新加坡大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02419 2025-08-05 cs.CV cs.CL 57%

Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens

Haohan Zheng, Zhenguo Zhang

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06593 2025-08-05 cs.CV 57%

SAGI: Semantically Aligned and Uncertainty Guided AI Image Inpainting

Paschalis Giakoumoglou, Dimitrios Karageorgiou, Symeon Papadopoulos, Panagiotis C. Petrantonakis

机构 * Department of Electrical and Computer Engineering, Aristotle University of Thessaloniki(阿尔伯塔大学电气与计算机工程系) Information Technologies Institute, CERTH(信息科技研究所)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23202 2025-08-01 cs.CV 57%

Adversarial-Guided Diffusion for Multimodal LLM Attacks

Chengwei Xia, Fan Ma, Ruijie Quan, Kun Zhan, Yi Yang

机构 * School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22744 2025-07-31 cs.CL cs.AI 57%

Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index

Praveenkumar Katwe, Rakesh Chandra, Balabantaray Kali, Prasad Vittala

机构 * International Institute of Information Technology(国际信息研究所) Informatica Business Solutions(Informatica商务解决方案)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

Comments 8

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 57%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21489 2025-07-30 cs.CV 57%

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

Zhichuan Wang, Yang Zhou, Zhe Liu, Rui Yu, Song Bai, Yulong Wang, Xinwei He, Xiang Bai

机构 * Huazhong Agricultural University(华中农业大学) Shenzhen University(深圳大学) The University of Hong Kong(香港大学) University of Louisville(路易斯安那大学) ByteDance(字节跳动) Huazhong University of Science and Technology(华中科技大学)

专题命中 幻觉与鲁棒性 :MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14348 2025-07-29 cs.CV 57%

Manipulating Multimodal Agents via Cross-Modal Prompt Injection

Le Wang, Zonghao Ying, Tianyuan Zhang, Siyuan Liang, Shengshan Hu, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北洋大学) National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) Henan University of Science and Technology(河南科技大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04757 2025-07-25 cs.CV cs.CL 57%

ELITE: Enhanced Language-Image Toxicity Evaluation for Safety

Wonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu, Ashkan Yousefpour, Haon Park, Bumsub Ham, Suhyun Kim

机构 * Yonsei University(延世大学) Kyung Hee University(庆熙大学) Korea Institute of Science(韩国科学研究院) Seoul National University(首尔国立大学) Sookmyung Women's University(_sookmyung女子大学)

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.CV

Comments ICML 2025. Project page at https://velpegor.github.io/ELITE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17118 2025-07-24 cs.AI 57%

HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study

Mandar Pitale, Jelena Frtunikj, Abhinaw Priyadershi, Vasu Singh, Maria Spence

机构 * Nvidia Corporation(英伟达公司) Nvidia GmbH(英伟达德国公司)

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16033 2025-07-23 cs.HC cs.AI 57%

"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives

Ding Wang, Mark Díaz, Charvi Rastogi, Aida Davani, Vinodkumar Prabhakaran, Pushkar Mishra, Roma Patel, Alicia Parrish, Zoe Ashwood, Michela Paganini, Tian Huey Teh, Verena Rieser, Lora Aroyo

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

Comments Accepted to AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society 2025 (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11968 2025-07-17 cs.CV 57%

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

Sahid Hossain Mustakim, S M Jishanul Islam, Ummay Maria Muna, Montasir Chowdhury, Mohammed Jawwadul Islam, Sadia Ahmmed, Tashfia Sikder, Syed Tasdid Azam Dhrubo, Swakkhar Shatabda

机构 * United International University(国际大学) BRAC University(BRAC大学) University of British Columbia(不列颠哥伦比亚大学) Bangladesh University of Professionals(孟加拉专业大学) University of Alberta(阿尔伯塔大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted as long paper, SVU Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10106 2025-07-15 cs.AI 57%

BlueGlass: A Framework for Composite AI Safety

Harshal Nandigramwar, Syed Qutub, Kay-Ulrich Scholl

机构 * Intel Labs, Munich, Germany(英特尔慕尼黑实验室) University of Stuttgart, Stuttgart, Germany(斯图加特大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI

Comments Accepted at ICML 2025 [Actionable Interpretability Workshop]

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00718 2025-07-11 cs.LG cs.SD eess.AS 57%

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models

Isha Gupta, David Khachaturov, Robert Mullins

机构 * ETH Zürich(苏黎世联邦理工学院) University of Cambridge(剑桥大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06814 2025-07-10 cs.CV 57%

HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement

Qingsen Yan, Kangbiao Shi, Yixu Feng, Tao Hu, Peng Wu, Guansong Pang, Yanning Zhang

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03782 2025-07-04 cs.CV 57%

Assessing the Uncertainty and Robustness of the Laptop Refurbishing Software

Chengjie Lu, Jiahui Wu, Shaukat Ali, Mikkel Labori Olsen

机构 * Simula Research Laboratory and University of Oslo(Simula研究实验室和奥斯陆大学) Danish Technological Institute(丹麦技术研究所)

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.CV

Comments 17 pages, 6 figures, 4 tables

Journal ref 2025 IEEE Conference on Software Testing, Verification and Validation (ICST)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01643 2025-07-03 cs.CV 57%

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

Weijie Yin, Dingkang Yang, Hongyuan Dong, Zijian Kang, Jiacong Wang, Xiao Liang, Chao Feng, Jiao Ran

机构 * ByteDance Inc.(字节跳动公司) College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV

Comments We release SAILViT, a series of versatile vision foundation models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00006 2025-07-02 cs.GR cs.LG eess.IV 57%

MVGBench: Comprehensive Benchmark for Multi-view Generation Models

Xianghui Xie, Chuhang Zou, Meher Gitika Karumuri, Jan Eric Lenssen, Gerard Pons-Moll

机构 * University of Tübingen(图宾根大学) Tübingen AI Center(图宾根人工智能中心) Max Planck Institute for Informatics(马克斯·普朗克信息研究所) Independent researcher(独立研究者)

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.LG

Comments 17 pages, 11 figures, 9 tables, project page: https://virtualhumans.mpi-inf.mpg.de/MVGBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24007 2025-06-30 cs.CV 57%

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model

Nokimul Hasan Arif, Shadman Rabby, Md Hefzul Hossain Papon, Sabbir Ahmed

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.CV

Comments Submitted for review in NCAA Springer, 21 pages, 4 figures, 4 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20531 2025-06-26 cs.AI cs.CY 57%

Case-based Reasoning Augmented Large Language Model Framework for Decision Making in Realistic Safety-Critical Driving Scenarios

Wenbin Gan, Minh-Son Dao, Koji Zettsu

机构 * Big Data Integration Research Center, National Institute of Information and Communications Technology (NICT), Tokyo, Japan(大数据整合研究室,日本信息通信技术研究所(NICT),东京)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI

Comments 12 pages, 10 figures, under-review conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17503 2025-06-24 cs.CV 57%

Trustworthy Few-Shot Transfer of Medical VLMs through Split Conformal Prediction

Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz

机构 * ÉTS Montréal(蒙特利尔ÉTS学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments MICCAI 2025. Code: https://github.com/jusiro/SCA-T

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14735 2025-06-18 cs.CL cs.AI 57%

Unleashing the potential of prompt engineering for large language models

Banghao Chen, Zhaofeng Zhang, Nicolas Langrené, Shengxin Zhu

机构 * Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science(interdisciplinary Research and Application for Data Science省重点实验室) BNU-HKBU United International College(北京师范大学-香港浸会大学联合国际学院) Research Center for Mathematics, Beijing Normal University(北京师范大学数学研究中心)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI

Comments v6 - Metadata updated (title, journal ref, DOI). PDF identical to v5 (original submission). Please cite the peer-reviewed Version of Record in "Patterns" (DOI: 10.1016/j.patter.2025.101260)

Journal ref Patterns 6(6) 101260 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12724 2025-06-17 cs.CV 57%

Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency

Hiroshi Tanaka, Anika Rao, Hana Satou, Michael Johnson, Sofia García

专题命中 幻觉与鲁棒性 :LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11521 2025-06-16 cs.CR cs.AI cs.MM 57%

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

Jinming Wen, Xinyi Wu, Shuai Zhao, Yanhao Jia, Yuwen Li

机构 * Jilin University(吉林大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Northeastern University(东北大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏