arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1738 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1738 篇

2506.23164 2025-07-01 cs.RO cs.AI 57%

Mode Collapse Happens: Evaluating Critical Interactions in Joint Trajectory Prediction Models

Maarten Hugenholtz, Anna Meszaros, Jens Kober, Zlatan Ajanovic

机构 * Cognitive Robotics Department, Delft University of Technology(代尔夫特理工大学认知机器人学系) Computer Science Department, RWTH Aachen University(亚琛工业大学计算机科学系)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

Comments 12 pages, 8 figures, submitted to a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19513 2025-06-25 cs.CV cs.LG 57%

Visual hallucination detection in large vision-language models via evidential conflict

Tao Huang, Zhekun Liu, Rui Wang, Yang Zhang, Liping Jing

机构 * Beijing Key Lab of Traffic Data Mining(北京交通数据挖掘与具身智能重点实验室) State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室) School of Computer Science and Technology(计算机科学与技术学院) Beijing Jiaotong University(北京交通大学) School of Automation and Intelligence(自动化与智能学院) School of Electronic and Information Engineering(电子与信息工程学院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Journal ref International Journal of Approximate Reasoning, Volume 186, November 2025, Article 109507

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18559 2025-06-24 cs.AI cs.LO 57%

T-CPDL: A Temporal Causal Probabilistic Description Logic for Developing Logic-RAG Agent

Hong Qing Yu

机构 * University of Derby(德比大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00958 2025-06-24 cs.LG cs.CV 57%

Navigating Conflicting Views: Harnessing Trust for Learning

Jueqing Lu, Wray Buntine, Yuanyuan Qi, Joanna Dipnall, Belinda Gabbe, Lan Du

机构 * Department of Data Science \& AI, Monash University School of Public Health Preventive Medicine, Monash University College of Engineering Computer Science, VinUniversity

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

Journal ref Proceedings of the 42nd International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15889 2025-06-23 cs.CL 57%

Entropy-Driven Pre-Tokenization for Byte-Pair Encoding

Yifan Hu, Frank Liang, Dachuan Zhao, Jonathan Geuter, Varshini Reddy, Craig W. Schmidt, Chris Tanner

机构 * Harvard University, Cambridge, MA(哈佛大学) Kempner Institute, Allston, MA(凯姆纳研究所) Kensho Technologies, Cambridge, MA(Kensho技术公司)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09272 2025-06-12 cs.LG stat.ML 57%

G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration

Samuel Holt, Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar

机构 * University of Cambridge(剑桥大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

Comments Accepted at the 42nd International Conference on Machine Learning (ICML 2025). 9 pages, 3 figures

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08422 2025-06-12 cs.AI 57%

Transforming Expert Knowledge into Scalable Ontology via Large Language Models

Ikkei Itoku, David Theil, Evelyn Eichelsdoerfer Uehara, Sreyoshi Bhaduri, Junnosuke Kuroda, Toshi Yumoto, Alex Gil, Natalie Perez, Rajesh Cherukuri, Naumaan Nayyar

机构 * Amazon New York, USA(亚马逊纽约分公司) Amazon Arlington, USA(亚马逊阿灵顿分公司) Amazon Seattle, USA(亚马逊西雅图分公司) Amazon Honolulu, USA(亚马逊檀香山分公司)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13949 2025-06-11 cs.CL cs.CV 57%

Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence

Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) National University of Singapore(新加坡国立大学) Southeast University(东南大学) Wuhan AI Research(武汉人工智能研究所)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments ACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07227 2025-06-10 cs.CV cs.CL 57%

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

Tianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun, Jiayi Song, Junlin Han, Zichen Liu, Conghui He, Wentao Zhang, Binhang Yuan

机构 * HKUST(香港科技大学) Shanghai AI Lab(上海人工智能实验室) Peking University(北京大学) HKUST(GZ)(香港科技大学(广州)) Oxford University(牛津大学) Imperial College London(伦敦帝国理工学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05826 2025-06-09 cs.LG 57%

Learning Along the Arrow of Time: Hyperbolic Geometry for Backward-Compatible Representation Learning

Ngoc Bui, Menglin Yang, Runjin Chen, Leonardo Neves, Mingxuan Ju, Rex Ying, Neil Shah, Tong Zhao

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01616 2025-06-03 cs.AI 57%

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Xiao Yang, Jiawei Chen, Jun Luo, Zhengwei Fang, Yinpeng Dong, Hang Su, Jun Zhu

机构 * Department of Computer Science & Technology, Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint Center for ML(计算机科学与技术系、人工智能研究院、BNRist中心、THBI实验室、清华-博世联合机器学习中心) Shanghai Key Lab. of Multidimensional Info. Processing(上海多维信息处理重点实验室)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24649 2025-06-02 cs.CV cs.AI 57%

BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models

Huu-Thien Tran, Thanh-Dat Truong, Khoa Luu

机构 * CVIU Lab, University of Arkansas(CVIU实验室,阿肯色大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments CVPRW 2025, 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24377 2025-06-02 cs.CL 57%

LLM Inference Enhanced by External Knowledge: A Survey

Yu-Hsuan Lin, Qian-Hui Chen, Yi-Jie Cheng, Jia-Ren Zhang, Yi-Hung Liu, Liang-Yu Hsia, Yun-Nung Chen

机构 * National Taiwan University(国立台湾大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23792 2025-06-02 cs.CR cs.AI 57%

Zero-Trust Foundation Models: A New Paradigm for Secure and Collaborative Artificial Intelligence for Internet of Things

Kai Li, Conggai Li, Xin Yuan, Shenghong Li, Sai Zou, Syed Sohail Ahmed, Wei Ni, Dusit Niyato, Abbas Jamalipour, Falko Dressler, Ozgur B. Akan

机构 * School of Electrical Engineering and Computer Science, TU Berlin(柏林技术大学电气与计算机工程学院) Real-Time and Embedded Computing Systems Research Centre (CISTER)(实时与嵌入式计算系统研究中心) Data61, CSIRO(数据61,澳大利亚联邦科学与工业研究组织) School of Computer Science and Engineering, the University of New South Wales(新南威尔士大学计算机科学与工程学院) State Key Laboratory of Public Big Data, College of Big Data and Information Engineering, Guizhou University(公共大数据国家重点实验室,贵州大学大数据与信息工程学院) Department of Computer Engineering, Qassim University(科威特大学计算机工程系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Division of Electrical Engineering, Department of Engineering, University of Cambridge(剑桥大学工程系电子工程系) Center for NeXt-Generation Communications (CXC), Koç University(下一代通信中心(CXC),科隆大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21958 2025-05-29 cs.CL 57%

Resolving Knowledge Conflicts in Domain-specific Data Selection: A Case Study on Medical Instruction-tuning

Qihuang Zhong, Liang Ding, Fei Liao, Juhua Liu, Bo Du, Dacheng Tao

机构 * Department of Gastroenterology, Renmin Hospital, Wuhan University(消化内科,仁医医院,武汉大学) School of Computer Science, Wuhan University(计算机学院,武汉大学) School of Computer Science, Faculty of Engineering, The University of Sydney(计算机学院,工程学院,悉尼大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20903 2025-05-28 cs.CL 57%

Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?

Ziming Wang, Zeyu Shi, Haoyi Zhou, Shiqi Gao, Qingyun Sun, Jianxin Li

机构 * SKLCCSE, School of Computer Science and Engineering, Beihang University(SKLCCSE,计算机科学与工程学院,北航) School of Software, Beihang University(软件学院,北航) Zhongguancun Laboratory, Beijing(中关村实验室,北京)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

Comments Accepted to ACL2025 Main; The code will be released soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20880 2025-05-28 cs.CL 57%

MSA at SemEval-2025 Task 3: High Quality Weak Labeling and LLM Ensemble Verification for Multilingual Hallucination Detection

Baraa Hikal, Ahmed Nasreldin, Ali Hamdi

机构 * Faculty of Computer Science, MSA University(计算机科学学院,MSA大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04615 2025-05-28 cs.CL 57%

HalluCounter: Reference-free LLM Hallucination Detection in the Wild!

Ashok Urlana, Gopichand Kanumolu, Charaka Vinayak Kumar, Bala Mallikarjunarao Garlapati, Rahul Mishra

机构 * IIIT Hyderabad(IIIT海得拉巴) TCS Research, Hyderabad, India(TCS研究, 海得拉巴, 印度) University of Oslo, Norway(奥斯陆大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments 30 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19511 2025-05-27 cs.CL 57%

Causal Distillation: Transferring Structured Explanations from Large to Compact Language Models

Aggrey Muhebwa, Khalid K. Osman

机构 * Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08739 2025-05-27 cs.LG q-bio.QM 57%

Bayesian Comparisons Between Representations

Heiko H. Schütt

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17520 2025-05-26 cs.AI 57%

Optimizing Retrieval-Augmented Generation for Electrical Engineering: A Case Study on ABB Circuit Breakers

Salahuddin Alawadhi, Noorhan Abbas

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

Comments 17 pages, 4 figures, published in CSIT Vol. 15, 2025. DOI: 10.5121/csit.2025.150905

Journal ref Computer Science and Information Technology, Vol. 15, 2025, pp. 59-77

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13501 2025-05-21 cs.LG physics.comp-ph 57%

SPIEDiff: robust learning of long-time macroscopic dynamics from short-time particle simulations with quantified epistemic uncertainty

Zequn He, Celia Reina

机构 * Department of Mechanical Engineering and Applied Mechanics University of Pennsylvania(机械工程与应用力学系 费城大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13312 2025-05-20 cs.CL 57%

GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection

Zhijie Deng, Chris Yuhao Liu, Zirui Pang, Xinlei He, Lei Feng, Qi Xuan, Zhaowei Zhu, Jiaheng Wei

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) UC Santa Cruz(加州大学圣克鲁兹分校) UIUC(伊利诺伊大学香槟分校) Southeast University(东南大学) BIAI, ZJUT(脑科学与类脑智能研究院,浙江大学) D5Data.ai

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13123 2025-05-20 cs.CV cs.AI 57%

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training

Xinsong Zhang, Yarong Zeng, Xinting Huang, Hu Hu, Runquan Xie, Han Hu, Zhanhui Kang

机构 * Tencent(腾讯)

专题命中 幻觉与事实性 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07309 2025-05-13 cs.LG 57%

Uncertainty Profiles for LLMs: Uncertainty Source Decomposition and Adaptive Model-Metric Selection

Pei-Fu Guo, Yun-Da Tsai, Shou-De Lin

机构 * National Taiwan University(国立台湾大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04709 2025-05-12 cs.AI 57%

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, Thomas Icard

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03467 2025-05-07 cs.CL 57%

Uncertainty-Aware Large Language Models for Explainable Disease Diagnosis

Shuang Zhou, Jiashuo Wang, Zidu Xu, Song Wang, David Brauer, Lindsay Welton, Jacob Cogan, Yuen-Hei Chung, Lei Tian, Zaifu Zhan, Yu Hou, Mingquan Lin, Genevieve B. Melton, Rui Zhang

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

Comments 22 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01487 2025-05-07 cs.AI 57%

FastRM: An efficient and automatic explainability framework for multimodal generative models

Gabriela Ben-Melech Stan, Estelle Aflalo, Man Luo, Shachar Rosenman, Tiep Le, Sayak Paul, Shao-Yen Tseng, Vasudev Lal

机构 * Intel Labs(英特尔实验室) Hugging Face

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01958 2025-05-06 cs.CV cs.CL 57%

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models

Liqiang Jing, Guiming Hardy Chen, Ehsan Aghazadeh, Xin Eric Wang, Xinya Du

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Massachusetts at Amherst(马萨诸塞大学阿姆赫斯特分校) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02440 2025-05-06 cs.LG 57%

Artificial Intelligence in Reactor Physics: Current Status and Future Prospects

Ruizhi Zhang, Shengfeng Zhu, Kan Wang, Ding She, Jean-Philippe Argaud, Bertrand Bouriquet, Qing Li, Helin Gong

机构 * School of Mathematical Sciences, Shanghai Key Laboratory of Pure Mathematics and Mathematical Practice, East China Normal University(数学学院、上海纯粹数学与数学实践重点实验室、华东师范大学) Department of Engineering Physics, Tsinghua University(工程物理系、清华大学) Électricité de France, R&D(法国电力公司、研发部门) Électricité de France, DQI(法国电力公司、DQI部门) Science and Technology on Reactor System Design Technology Laboratory, Nuclear Power Institute of China(反应堆系统设计技术实验室、中国核工业集团核电研究院) Paris Elite Institute of Technology, Shanghai Jiao Tong University(巴黎精英技术学院、上海交通大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments 33 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏