arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1732 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1732 篇

2510.01237 2025-10-03 cs.CL cs.AI 62%

Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation

Nandakishor M

机构 * AI Safety Research(人工智能安全研究) Convai Innovations(Convai创新)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23585 2025-10-02 cs.LG cs.AI cs.CV 62%

EVO-LRP: Evolutionary Optimization of LRP for Interpretable Model Explanations

Emerald Zhang, Julian Weaver, Samantha R Santacruz, Edward Castillo

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23497 2025-09-30 cs.AI cs.HC cs.LG 62%

Dynamic Trust Calibration Using Contextual Bandits

Bruno M. Henrique, Eugene Santos

机构 * Thayer School of Engineering(泰勒工程学院) Dartmouth College(达特茅斯学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23146 2025-09-30 cs.CL cs.LG 62%

Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models

Zichao Yu, Ming Li, Wenyi Zhang, Weiguo Gao

机构 * University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19375 2025-09-25 cs.LG cs.AI stat.ML 62%

Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation

Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil

机构 * Faculty of Dental Medicine and Oral Health Sciences, McGill University(牙医学院与口腔健康科学学院,麦吉尔大学) McGill University(麦吉尔大学) Mila–Quebec Artificial Intelligence Institute(魁北克人工智能研究所) School of Computer Science, McGill University(计算机科学学院,麦吉尔大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17671 2025-09-23 cs.CL cs.AI 62%

Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications

Selva Taş, Mahmut El Huseyni, Özay Ezerceli, Reyhan Bayraktar, Fatma Betül Terzioğlu

机构 * Hidden for Review(保密)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16696 2025-09-23 cs.CL cs.LG 62%

Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models

Wataru Hashimoto, Hidetaka Kamigaito, Taro Watanabe

机构 * Nara Institute of Science and Technology(奈良科学技术大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13702 2025-09-18 cs.CL cs.AI 62%

DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models

Xiao Zheng

机构 * School of Computing and Technology(计算机学院) China University of Petroleum(中国石油大学) Qingdao(青岛)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13334 2025-09-18 cs.AI cs.LG 62%

FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness

Anand Swaroop, Akshat Nallani, Saksham Uboweja, Adiliia Uzdenova, Michael Nguyen, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Vasu Sharma, Maheep Chaudhary

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07082 2025-09-16 cs.CV cs.AI cs.LG 62%

On the Generalization of Representation Uncertainty in Earth Observation

Spyros Kondylatos, Nikolaos Ioannis Bountos, Dimitrios Michail, Xiao Xiang Zhu, Gustau Camps-Valls, Ioannis Papoutsis

机构 * National Observatory of Athens(雅典国家天文台) National Technical University of Athens(雅典技术大学) University of Valencia(瓦伦西亚大学) Harokopio University of Athens(雅典惠克罗波利斯大学) Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Archimedes/Athena RC(阿基米德/雅典RC)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01222 2025-09-16 cs.LG cs.AI 62%

Calibration in Deep Learning: A Survey of the State-of-the-Art

Cheng Wang

机构 * Amazon(亚马逊)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10208 2025-09-15 cs.CL cs.AI 62%

SI-FACT: Mitigating Knowledge Conflict via Self-Improving Faithfulness-Aware Contrastive Tuning

Shengqiang Fu

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15850 2025-09-12 cs.LG cs.AI 62%

Uncertainty Estimation by Human Perception versus Neural Models

Pedro Mendes, Paolo Romano, David Garlan

机构 * Software and Societal Systems Department, Carnegie Mellon University(卡内基梅隆大学软件与社会系统部门) INESC-ID and Instituto Superior Técnico, Universidade de Lisboa(里斯本大学INESC-ID和理工学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00300 2025-09-11 cs.HC cs.AI cs.LG 62%

MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems

Shruthi Chari, Oshani Seneviratne, Prithwish Chakraborty, Pablo Meyer, Deborah L. McGuinness

机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院) Amazon Science(亚马逊科学) IBM Research(IBM研究院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07475 2025-09-10 cs.CL cs.AI 62%

HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention

Saumya Goswami, Siddharth Kurra

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06596 2025-09-09 cs.CL cs.AI 62%

HAVE: Head-Adaptive Gating and ValuE Calibration for Hallucination Mitigation in Large Language Models

Xin Tong, Zhi Lin, Jingya Wang, Bo Jin

机构 * Xin Tong School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) Zhi Lin School of Safety Science Tsinghua University(安全科学学院 清华大学) Jingya Wang School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) Bo Jin* The Third Research Institute of the Ministry of Public Security of China(公安部第三研究所)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02499 2025-09-09 cs.CL cs.AI 62%

MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds

Junxi Wu, Jinpeng Wang, Zheng Liu, Bin Chen, Dongjian Hu, Hao Wu, Shu-Tao Xia

机构 * Nankai University(南开大学) Tsinghua University(清华大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室) Shenzhen ShenNong Information Technology Co., Ltd.(深圳深农信息技术有限公司)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00461 2025-09-08 cs.CL cs.AI 62%

TECP: Token-Entropy Conformal Prediction for LLMs

Beining Xu, Yongming Lu

机构 * Department of School of Engineering, Shenzhen MSU-BIT University, Shenzhen, China, 518000(深圳MSU-BIT大学工程学院部门) MSU-BIT-SMBU Joint Research Center of Applied Mathematics, Shenzhen MSU-BIT University, Shenzhen, China, 518000(应用数学联合研究中心)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13748 2025-09-04 cs.CL cs.LG 62%

Learn and Unlearn: Addressing Misinformation in Multilingual LLMs

Taiming Lu, Philipp Koehn

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00646 2025-09-03 cs.CY cs.AI 62%

RAG-PRISM: A Personalized, Rapid, and Immersive Skill Mastery Framework with Adaptive Retrieval-Augmented Tutoring

Gaurangi Raul, Yu-Zheng Lin, Karan Patel, Bono Po-Jen Shih, Matthew W. Redondo, Banafsheh Saber Latibari, Jesus Pacheco, Soheil Salehi, Pratik Satam

机构 * College of Information Science, University of Arizona, Tucson, AZ, USA(信息科学学院,亚利桑那大学) Department of Electrical and Computer Engineering, University of Arizona, Tucson, AZ, USA(电气与计算机工程系,亚利桑那大学) Department of Systems and Industrial Engineering, University of Arizona, Tucson, AZ, USA(系统与工业工程系,亚利桑那大学) Department of Industrial Engineering, University of Sonora, Hermosillo, Mexico(工业工程系,索诺拉大学) Leonhard Center for Enhancement of Engineering Education, The Pennsylvania State University, University Park, PA, USA(工程教育增强中心,宾夕法尼亚州立大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.CY

Comments 9 pages, 5 figures, Accepted by IEEE FIE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21228 2025-09-01 cs.CL cs.AI 62%

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection

Weizhi Gao, Xiaorui Liu, Feiyi Wang, Dan Lu, Junqi Yin

机构 * ORNL(橡树岭国家实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17753 2025-08-26 cs.RO cs.AI cs.CL cs.HC 62%

Talking to Robots: A Practical Examination of Speech Foundation Models for HRI Applications

Theresa Pekarek Rosin, Julia Gachot, Henri-Leon Kordt, Matthias Kerzel, Stefan Wermter

机构 * Knowledge Technology, Department of Informatics, University of Hamburg(知识技术,信息学院,汉堡大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at the workshop on Foundation Models for Social Robotics (FoMoSR) at ICSR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10266 2025-08-26 cs.LG cs.AI cs.SY eess.SP eess.SY 62%

Intelligent Condition Monitoring of Industrial Plants: An Overview of Methodologies and Uncertainty Management Strategies

Maryam Ahang, Todd Charter, Mostafa Abbasi, Maziyar Khadivi, Oluwaseyi Ogunfowora, Homayoun Najjaran

机构 * Department of Electrical and Computer Engineering, University of Victoria(电气与计算机工程系,维多利亚大学) Department of Mechanical Engineering, University of Victoria(机械工程系,维多利亚大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13564 2025-08-20 cs.CV cs.AI cs.LG cs.RO 62%

The 9th AI City Challenge

Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Norimasa Kobori, Munkhjargal Gochoo, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Fady Alnajjar, Jun-Wei Hsieh, Tomasz Kornuta, Xiaolong Li, Yilin Zhao, Han Zhang, Subhashree Radhakrishnan, Arihant Jain, Ratnesh Kumar, Vidya N. Murali, Yuxing Wang, Sameer Satish Pusegaonkar, Yizhou Wang, Sujit Biswas, Xunlei Wu, Zhedong Zheng, Pranamesh Chakraborty, Rama Chellappa

机构 * NVIDIA Corporation(NVIDIA公司) Santa Clara University(圣克拉拉大学) University at Albany, SUNY(纽约州立大学阿尔巴尼分校) Iowa State University(爱荷华州立大学) Woven by Toyota, Japan(日本丰田公司) United Arab Emirates University(阿拉伯联合酋长国大学) National Yang-Ming Chiao-Tung University(国家阳明交通大学) University of Macau(澳门大学) Indian Institute of Technology Kanpur(印度理工学院坎普尔分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments Summary of the 9th AI City Challenge Workshop in conjunction with ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09085 2025-08-13 cs.NI cs.AI cs.LG 62%

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

Zihan Fang, Zheng Lin, Senkang Hu, Yihang Tao, Yiqin Deng, Xianhao Chen, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong(香港JC STEM实验室及城市大学计算机科学系)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06017 2025-08-11 cs.SE cs.CL cs.LG 62%

Position: Intelligent Coding Systems Should Write Programs with Justifications

Xiangzhe Xu, Shiwei Feng, Zian Su, Chengpeng Wang, Xiangyu Zhang

机构 * Purdue University(普渡大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06306 2025-08-11 cs.CL cs.AI cs.HC 62%

Humans overrely on overconfident language models, across languages

Neil Rathi, Dan Jurafsky, Kaitlyn Zhou

机构 * Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10583 2025-08-08 cs.SE cs.AI cs.CY 62%

$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection

Daniil Orel, Indraneil Paul, Iryna Gurevych, Preslav Nakov

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(Mohamed bin Zayed人工智能大学) Ubiquitous Knowledge Processing Lab (UKP Lab)(普及知识处理实验室) Department of Computer Science(计算机科学系) National Research Center for Applied Cybersecurity ATHENE(应用网络安全国家研究中心ATHENE)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03092 2025-08-06 cs.AI cs.CL 62%

Toward Verifiable Misinformation Detection: A Multi-Tool LLM Agent Framework

Zikun Cui, Tianyi Huang, Chia-En Chiang, Cuiqianhe Du

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23628 2025-08-04 cs.CL cs.AI 62%

AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora

Jiaxin Bai, Wei Fan, Qi Hu, Qing Zong, Chunyang Li, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Haoyu Huang, Xiao Zhou, Feng Qin, Tianshi Zheng, Xi Peng, Xin Yao, Huiwen Yang, Leijie Wu, Yi Ji, Gong Zhang, Renhai Chen, Yangqiu Song

机构 * CSE, HKUST(香港科技大学计算机科学与工程系) CSE, CUHK(香港城市大学计算机科学与工程系) Theory Lab, Huawei(华为理论实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 9 pages, preprint, code: https://github.com/HKUST-KnowComp/AutoSchemaKG

详情

展开后加载摘要…

URL PDF HTML 收藏