arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1738 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1738 篇

2510.06541 2025-10-09 cs.CV cs.LG 57%

Cluster Paths: Navigating Interpretability in Neural Networks

Nicholas M. Kroeger, Vincent Bindschaedler

机构 * Department of Computer Science University of Florida(计算机科学系佛罗里达大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06125 2025-10-08 cs.LG 57%

Downsized and Compromised?: Assessing the Faithfulness of Model Compression

Moumita Kamal, Douglas A. Talbert

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments Submitted to and under review at Springer Machine Learning Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04439 2025-10-07 cs.CL 57%

On the Role of Unobserved Sequences on Sample-based Uncertainty Quantification for LLMs

Lucie Kunitomo-Jacquin, Edison Marrese-Taylor, Ken Fukuda

机构 * National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted to UncertaiNLP workshop of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04357 2025-10-07 cs.LG q-fin.CP 57%

From News to Returns: A Granger-Causal Hypergraph Transformer on the Sphere

Anoushka Harit, Zhongtian Sun, Jongmin Yu

机构 * University of Cambridge(剑桥大学) University of Kent(肯特大学) Department of Computer Science, University of Cambridge(剑桥大学计算机科学系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments 6th ACM International Conference on AI in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10246 2025-10-07 cs.LG 57%

Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions

Hazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr, Yarin Gal

机构 * University of Oxford(牛津大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments Accepted to EMNLP(main)2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00015 2025-10-03 cs.CL 57%

Design and Application of Multimodal Large Language Model Based System for End to End Automation of Accident Dataset Generation

MD Thamed Bin Zaman Chowdhury, Moazzem Hossain

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments This paper is accepted for presentation in TRB annual meeting 2026. The version presented here is the preprint version before peer review process

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01069 2025-10-02 cs.AI 57%

Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning

Elija Perrier

机构 * Centre for Quantum Software and Information(量子软件与信息中心) University of Technology Sydney(悉尼技术大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12838 2025-10-02 cs.CL 57%

Are Knowledge and Reference in Multilingual Language Models Cross-Lingually Consistent?

Xi Ai, Mahardika Krisna Ihsani, Min-Yen Kan

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments EMNLP'25 Findings Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26610 2025-10-01 cs.LG 57%

Uncertainty Quantification for Regression using Proper Scoring Rules

Alexander Fishkov, Kajetan Schweighofer, Mykyta Ielanskyi, Nikita Kotelevskii, Mohsen Guizani, Maxim Panov

机构 * Department of Machine Learning, MBZUAI(机器学习部门,MBZUAI) ELLIS Unit Linz and LIT AI Lab, Institute for Machine Learning, Johannes Kepler University Linz(林兹ELLIS单元和LIT AI实验室,机器学习研究所,林兹约翰内斯·开普勒大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25669 2025-10-01 cs.AI 57%

GroundSight: Augmenting Vision-Language Models with Grounding Information and De-hallucination

Xinxi Chen, Tianyang Chen, Lijia Hong

机构 * Independent Researcher(独立研究者)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04943 2025-10-01 cs.CV cs.CL 57%

ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding

Jianjiang Yang, Yanshu li, Ziyan Huang

机构 * University of Bristol(布里斯托大学) Brown University(布朗大学) South China University of Technology(华南理工大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments Accepted by conference EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23088 2025-09-30 cs.CL 57%

The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models

Esteban Garces Arias, Julian Rodemann, Christian Heumann

机构 * Department of Statistics, LMU Munich(统计系,慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍兹研究中心)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments Accepted at the 2nd UncertaiNLP Workshop @ EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21593 2025-09-29 cs.AI physics.soc-ph 57%

GeoEvolve: Automating Geospatial Model Discovery via Multi-Agent Large Language Models

Peng Luo, Xiayin Lou, Yu Zheng, Zhuo Zheng, Stefano Ermon

机构 * Massachusetts Institute of Technology(麻省理工学院) Technical University of Munich(慕尼黑技术大学) Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21336 2025-09-29 cs.IR cs.CL 57%

HetaRAG: Hybrid Deep Retrieval-Augmented Generation across Heterogeneous Data Stores

Guohang Yan, Yue Zhang, Pinlong Cai, Ding Wang, Song Mao, Hongwei Zhang, Yaoze Zhang, Hairong Zhang, Xinyu Cai, Botian Shi

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18792 2025-09-24 cs.CL 57%

Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing

Sabri Boughorbel, Fahim Dalvi, Nadir Durrani, Majd Hawasly

机构 * Qatar Computing Research Institute, HBKU(卡塔尔计算研究所,哈瓦那大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments 12 pages, accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18132 2025-09-24 cs.AI 57%

Position Paper: Integrating Explainability and Uncertainty Estimation in Medical AI

Xiuyi Fan

机构 * Lee Kong Chian School of Medicine, College of Computing Data Science, Nanyang Technological University, Singapore

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments Accepted at the International Joint Conference on Neural Networks, IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18128 2025-09-24 cs.LG 57%

Accounting for Uncertainty in Machine Learning Surrogates: A Gauss-Hermite Quadrature Approach to Reliability Analysis

Amirreza Tootchi, Xiaoping Du

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16742 2025-09-23 cs.AI 57%

Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories

Mohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang, Zichao Wang, Chandan Reddy, Ming Jin, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Virginia Tech(弗吉尼亚理工大学) Meta AI Adobe Research(Adobe研究)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16369 2025-09-23 cs.IR cs.AI cs.CE 57%

Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction

Akshay Govind Srinivasan, Ryan Jacob George, Jayden Koshy Joe, Hrushikesh Kant, Harshith M R, Sachin Sundar, Sudharshan Suresh, Rahul Vimalkanth, Vijayavallabh

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 14 Pages, 8 Tables, 2 Figures. Accepted and to be published in the proceedings of FinNLP, Empirical Methods in Natural Language Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07755 2025-09-23 cs.CL cs.CR 57%

Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts

Rochana Prih Hastuti, Rian Adam Rajagede, Mansour Al Ghanim, Mengxin Zheng, Qian Lou

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 Findings. Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12406 2025-09-17 cs.LG stat.ML 57%

Bayesian Parametric Matrix Models: Principled Uncertainty Quantification for Spectral Learning

Mohammad Nooraiepour

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12034 2025-09-16 cs.AI 57%

Human-AI Use Patterns for Decision-Making in Disaster Scenarios: A Systematic Review

Emmanuel Adjei Domfeh, Christopher L. Dancy

机构 * Department of Computer Science and Engineering(计算机科学与工程系) The Pennsylvania State University(宾夕法尼亚州立大学) Department of Industrial and Manufacturing Engineering(工业与制造工程系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08777 2025-09-11 cs.CV cs.CL 57%

Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

Eric Slyman, Mehrab Tanjim, Kushal Kafle, Stefan Lee

机构 * Adobe(Adobe公司) Oregon State University(俄勒冈州立大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments 17 pages, 8 figures, Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08679 2025-09-11 cs.LG 57%

Signal Fidelity Index-Aware Calibration for Dementia Predictions Across Heterogeneous Real-World Data

Jingya Cheng, Jiazi Tian, Federica Spoto, Alaleh Azhir, Daniel Mork, Hossein Estiri

机构 * Department of Medicine, Massachusetts General Hospital(麻省总医院内科部) Department of Biostatistics, Harvard T.H. Chan School of Public Health(哈佛T.H. Chan公共卫生学院生物统计学部) Department of Medicine, Brigham and Women’s Hospital(布里洛妇产科医院内科部)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13107 2025-09-11 cs.CL cs.IR 57%

All for law and law for all: Adaptive RAG Pipeline for Legal Research

Figarri Keisha, Prince Singh, Pallavi, Dion Fernandes, Aravindh Manivannan, Ilham Wicaksono, Faisal Ahmad, Wiem Ben Rim

机构 * University College London(伦敦大学学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments submitted to NLLP 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05557 2025-09-09 cs.AI 57%

MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media

Rui Lu, Jinhe Bi, Yunpu Ma, Feng Xiao, Yuntao Du, Yijun Tian

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04735 2025-09-08 cs.CV cs.AI 57%

Enhancing Self-Driving Segmentation in Adverse Weather Conditions: A Dual Uncertainty-Aware Training Approach to SAM Optimization

Dharsan Ravindran, Kevin Wang, Zhuoyuan Cao, Saleh Abdelrahman, Jeffery Wu

机构 * Queen's University(女王大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04664 2025-09-08 cs.CL 57%

Why Language Models Hallucinate

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang

机构 * OpenAI Georgia Tech(佐治亚理工学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02129 2025-09-03 cs.LG cs.CV 57%

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang

机构 * Hong Kong University of Science(香港科学与技术大学) South China Normal University, Shanwei, Guangdong, China(华南师范大学,汕尾,广东,中国) Shenzhen Technology University, Shenzhen, Guangdong, China(深圳科技大学,深圳,广东,中国) University of Science(科学大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00869 2025-09-03 cs.CL 57%

Exploring and Mitigating Fawning Hallucinations in Large Language Models

Zixuan Shangguan, Yanjie Dong, Lanjun Wang, Xiaoyi Fan, Victor C. M. Leung, Xiping Hu

机构 * School of Medical Technology, Beijing Institute of Technology(北京理工大学医学技术学院) Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(深圳MSU-BIT大学人工智能研究院) Tianjin University(天津大学) The Hong Kong University of Science(香港科学大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏