arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-15 至 2026-04-15 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 4 篇

2510.25512 2026-04-15 cs.LG cs.AI cs.CV 62%

FaCT: Faithful Concept Traces for Explaining Neural Network Decisions

FaCT:用于解释神经网络决策的忠实概念追踪

Amin Parchami-Araghi, Sukrut Rao, Jonas Fischer, Bernt Schiele

机构 * Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出FaCT模型,通过共享类间概念实现忠实解释,引入C²-Score评估概念一致性,提升可解释性同时保持ImageNet性能。

Comments 35 pages, 23 figures, 2 tables, Neural Information Processing Systems (NeurIPS) 2025; Code is available at https://github.com/m-parchami/FaCT

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12610 2026-04-15 cs.CL 57%

Transforming External Knowledge into Triplets for Enhanced Retrieval in RAG of LLMs

将外部知识转化为三元组以增强LLM中RAG的检索

Xudong Wang, Chaoning Zhang, Qigan Sun, Zhenzhen Huang, Chang Lu, Sheng Zheng, Zeyu Ma, Caiyan Qin, Yang Yang, Hengtao Shen

机构 * School of Computing, Kyung Hee University(韩国庆熙大学计算机学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Robotics and Advanced Manufacture, Harbin Institute of Technology(哈尔滨工业大学机器人与先进制造学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

AI总结 Tri-RAG通过结构化三元组构建提升检索效率,减少冗余信息,提高生成质量与资源利用率。

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12543 2026-04-15 cs.AI 57%

A Two-Stage LLM Framework for Accessible and Verified XAI Explanations

一个两阶段LLM框架用于可访问且验证的XAI解释

Georgios Mermigkis, Dimitris Metaxakis, Marios Tyrovolas, Argiris Sofotasios, Nikolaos Avgeris, Panagiotis Hadjidoukas, Chrysostomos Stylios

机构 * Department of Computer Engineering and Informatics, University of Patras(帕特拉大学计算机工程与信息学系) Department of Informatics and Telecommunications, University of Ioannina(伊奥安尼纳大学信息与电信系) Industrial Systems Institute, Athena Research Center(雅典研究中心工业系统研究所) Archimedes Unit, Athena Research Center(雅典研究中心阿基米德单位)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出两阶段LLM框架,通过解释器和验证器LLM生成并验证XAI解释,提升解释的准确性与可访问性。

Comments 8 pages, 8 figures, Accepted for publication at the 2026 IEEE World Congress on Computational Intelligence (WCCI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12184 2026-04-15 cs.AI 57%

TRUST Agents: A Collaborative Multi-Agent Framework for Fake News Detection, Explainable Verification, and Logic-Aware Claim Reasoning

TRUST代理:一种用于虚假新闻检测、可解释验证和逻辑感知声明推理的协作多智能体框架

Gautama Shastry Bulusu Venkata, Santhosh Kakarla, Maheedhar Omtri Mohan, Aishwarya Gaddam

机构 * George Mason University(乔治·马歇尔大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

AI总结 TRUST代理通过协作多智能体框架提升虚假新闻检测的可解释性和逻辑推理能力,引入分解器、陪审团和逻辑聚合器提升复杂声明的验证效果。

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏