arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 2255 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 2255 篇

2512.02263 2025-12-03 cs.HC cs.GR 50%

DepthScape: Authoring 2.5D Designs via Depth Estimation, Semantic Understanding, and Geometry Extraction

DepthScape: 通过深度估计、语义理解和几何提取进行2.5D设计创作

Xia Su, Cuong Nguyen, Matheus A. Gadelha, Jon E. Froehlich

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

AI总结 DepthScape通过深度估计、语义理解和几何提取,实现2.5D设计创作,提升视觉真实感与动态效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00928 2025-12-02 cs.MM 50%

Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation

增强多模态大语言模型的模态内理解以实现鲁棒的多模态关键词生成

Jiajun Cao, Qinggang Zhang, Yunbo Tang, Zhishang Xiang, Chang Yang, Jinsong Su

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)

AI总结 AimKP通过增强模态内语义学习和跨模态对齐,提升多模态关键词生成的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19914 2025-11-26 cs.RO 50%

CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model

CoC-VLA: 深入研究对抗域转移以实现可解释的自动驾驶 via 链式因果视觉语言动作模型

Dapeng Zhang, Fei Shen, Rui Zhao, Yinda Chen, Peng Zhi, Chenyang Li, Rui Zhou, Qingguo Zhou

机构 * Lanzhou University(兰州大学) National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学)

专题命中 幻觉与鲁棒性 :VLM(abstract)

AI总结 CoC-VLA通过链式因果视觉语言动作模型实现对抗域转移,提升自动驾驶的可解释性和长尾场景处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15478 2025-11-25 cs.CL 50%

Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models

针对多模态语言模型的红队测试:评估不同提示模态和模型的有害性

Madison Van Doren, Casey Ford

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)

AI总结 本研究通过红队测试评估多模态语言模型在不同提示模态下的安全性,发现Pixtral 12B的有害响应率最高,而Claude Sonnet 3.5最安全,凸显了建立多模态安全基准的必要性。

Journal ref AAAI 2026 AIGOV Workshop and EurIPS 2025 Workshop on Unifying Perspectives on Learning Biases

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16816 2025-11-24 stat.AP stat.CO 50%

Trust-Aware Multimodal Data Fusion for Yield Estimation: A Case Study of the 2020 Beirut Explosion

具有信任意识的多模态数据融合用于产量估计:贝鲁特2020爆炸的案例研究

Lekha Patel, Craig Ulmer, Stephen J. Verzi, Daniel J. Krofcheck, Indu Manickam, Asmeret Naugle, Jaideep Ray

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

AI总结 本文提出一种基于贝叶斯分数后验框架的多模态数据融合方法,用于估计爆炸产量,通过信任权重校准不同观测数据,提升不确定性量化和抗偏差能力。

Comments 19 pages, 4 figures, supplementary material, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10075 2025-11-14 cs.CL 50%

Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts

Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04835 2025-11-10 cs.RO 50%

Conformalized Non-uniform Sampling Strategies for Accelerated Sampling-based Motion Planning

Shubham Natraj, Bruno Sinopoli, Yiannis Kantaros

机构 * Department of Electrical and Systems Engineering, Washington University in St. Louis(电气与系统工程系,华盛顿大学圣路易斯分校)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11350 2025-11-10 cs.RO 50%

Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild

Derek Ming Siang Tan, Shailesh, Boyang Liu, Alok Raj, Qi Xuan Ang, Weiheng Dai, Tanishq Duhan, Jimmy Chiun, Yuhong Cao, Florian Shkurti, Guillaume Sartoretti

专题命中 幻觉与鲁棒性 :vision language model(abstract)

Comments Accepted for presentation at CORL 2025. Code, models, and data are available at https://search-tta.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00776 2025-11-04 cs.SE 50%

A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI

Cuiyun Gao, Guodong Fan, Chun Yong Chong, Shizhan Chen, Chao Liu, David Lo, Zibin Zheng, Qing Liao

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20223 2025-10-24 cs.CR cs.MM 50%

Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations

Divyanshu Kumar, Shreyas Jena, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19886 2025-10-24 cs.CL 50%

An Expert-grounded benchmark of General Purpose LLMs in LCA

Artur Donaldson, Bharathan Balaji, Cajetan Oriekezie, Manish Kumar, Laure Patouillard

机构 * PRé Sustainability B.V.(PRé可持续性公司) Amazon(亚马逊)

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16198 2025-10-21 cs.CL 50%

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad, Hesham Abdelgawad, Mohamed Aref, Ahmed Fares

机构 * Department of Electrical Engineering, Faculty of Engineering at Shoubra, Benha University, Cairo 11629, Egypt(电气工程系,谢布拉工程学院,本海大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13190 2025-10-16 cs.CL 50%

SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16012 2025-10-14 cs.RO 50%

DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning

Boyu Li, Siyuan He, Hang Xu, Haoqi Yuan, Yu Zang, Liwei Hu, Junpeng Yue, Zhenxiong Jiang, Pengbo Hu, Börje F. Karlsson, Yehui Tang, Zongqing Lu

专题命中 幻觉与鲁棒性 :vision language model(abstract)

Comments The experiments in the paper need to be further supplemented, and more methods should be considered for expansion

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01929 2025-10-03 cs.CL 50%

Inverse Language Modeling towards Robust and Grounded LLMs

Davide Gabrielli, Simone Sestito, Iacopo Masi

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) OmnAI Lab(OmnAI实验室)

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22486 2025-09-29 cs.IR cs.CR 50%

Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks

Gaurav Bagwe, Saket S. Chaturvedi, Xiaolong Ma, Xiaoyong Yuan, Kuang-Ching Wang, Lan Zhang

专题命中 幻觉与鲁棒性 :grounding(abstract)

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14446 2025-09-19 q-bio.NC 50%

Mouse vs. AI: A Neuroethological Benchmark for Visual Robustness and Neural Alignment

Marius Schneider, Joe Canzano, Jing Peng, Yuchen Hou, Spencer LaVere Smith, Michael Beyeler

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13127 2025-09-17 cs.CL 50%

Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning

Sijia Cui, Shuai Xu, Aiyao He, Yanna Wang, Bo Xu

机构 * 1 Institute of Automation, Chinese Academy of Sciences, Beijing, China 2 School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China 3 Nanjing Artificial Intelligence Research of IA, Nanjing, China 4 University of Chinese Academy of Sciences,Nanjing, Nanjing, China 5 Nanjing University of Information Science \& Technology, Nanjing, China

专题命中 幻觉与鲁棒性 :grounding(abstract)

Comments Accepted to IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05838 2025-09-09 cs.CY 50%

Towards an Automated Framework to Audit Youth Safety on TikTok

Linda Xue, Francesco Corso, Nicolo' Fontana, Geng Liu, Stefano Ceri, Francesco Pierri

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

Comments 7 pages, 3 figures, submitted to EMNLP 2025 and ECAT Research Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04214 2025-09-05 cs.CR 50%

An Automated, Scalable Machine Learning Model Inversion Assessment Pipeline

Tyler Shumaker, Jessica Carpenter, David Saranchak, Nathaniel D. Bastian

专题命中 幻觉与鲁棒性 :vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03787 2025-09-05 cs.IR cs.CL 50%

Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain

Shakiba Amirshahi, Amin Bigdeli, Charles L. A. Clarke, Amira Ghenai

机构 * University of Waterloo(滑铁卢大学) Toronto Metropolitan University(多伦多 Metropolitan 大学)

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13053 2025-08-07 cs.CL 50%

Evaluating the Robustness of Multimodal Agents Against Active Environmental Injection Attacks

Yurun Chen, Xavier Hu, Keting Yin, Juncheng Li, Shengyu Zhang

机构 * School of Software Technology Zhejiang University Hangzhou Zhejiang China(软件技术学院浙江大学杭州浙江中国) Zhejiang University(浙江大学)

专题命中 幻觉与鲁棒性 :MLLM(abstract)

Comments Accepted at ACM MM 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04980 2025-07-16 cs.RO cs.SY eess.SY 50%

LVLM-MPC Collaboration for Autonomous Driving: A Safety-Aware and Task-Scalable Control Architecture

Kazuki Atsuta, Kohei Honda, Hiroyuki Okuda, Tatsuya Suzuki

机构 * Department of Mechanical Systems Engineering(机械系统工程系) Nagoya University(名古屋大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08203 2025-07-14 cs.CL 50%

TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs

Duygu Nur Yaldiz, Yavuz Faruk Bakman, Sungmin Kang, Alperen Öziş, Hayrettin Eren Yildiz, Mitash Ashish Shah, Zhiqi Huang, Anoop Kumar, Alfy Samuel, Daben Liu, Sai Praneeth Karimireddy, Salman Avestimehr

机构 * University of Southern California(南加州大学) Bogazici University(博鲁斯大学) Capital One(Capital One公司)

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00740 2025-07-02 cs.CR cs.CL cs.DC 50%

Safe Low Bandwidth SPV: A Formal Treatment of Simplified Payment Verification Protocols and Security Bounds

Craig S Wright

机构 * University of Exeter Business School(埃克塞特大学商学院)

专题命中 幻觉与鲁棒性 :grounding(abstract)

Comments 56 pages 5 images

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20944 2025-06-27 cs.MM cs.CR 50%

E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs

Van-Hoang Phan, Long-Khanh Pham, Dang Vu, Anh-Duy Tran, Minh-Son Dao

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

Comments Accepted to AsiaCCS 2025 @ SCID

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15677 2025-06-27 cs.RO 50%

Learning Efficient and Robust Language-conditioned Manipulation using Textual-Visual Relevancy and Equivariant Language Mapping

Mingxi Jia, Haojie Huang, Zhewen Zhang, Chenghao Wang, Linfeng Zhao, Dian Wang, Jason Xinyu Liu, Robin Walters, Robert Platt, Stefanie Tellex

机构 * Brown University(布朗大学) Northeastern University(东北大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13198 2025-06-18 cs.RO 50%

LBAP: Improved Uncertainty Alignment of LLM Planners using Bayesian Inference

James F. Mullen, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00054 2025-06-03 cs.IR cs.CL 50%

Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers

Chaitanya Sharma

机构 * Independent Researcher(独立研究者)

专题命中 幻觉与鲁棒性 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02391 2025-05-27 cs.CL 50%

Attacking Vision-Language Computer Agents via Pop-ups

Yanzhe Zhang, Tao Yu, Diyi Yang

机构 * Georgia Tech(佐治亚理工学院) The University of Hong Kong(香港大学) Stanford University(斯坦福大学)

专题命中 幻觉与鲁棒性 :VLM(abstract)

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏