arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1732 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1732 篇

2505.23968 2025-06-02 cs.CR cs.AI cs.CY cs.LG stat.ML 67%

Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention

Stephan Rabanser, Ali Shahin Shamsabadi, Olive Franzese, Xiao Wang, Adrian Weller, Nicolas Papernot

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Proceedings of the 42nd International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23854 2025-06-02 cs.CL cs.AI cs.LG 67%

Revisiting Uncertainty Estimation and Calibration of Large Language Models

Linwei Tao, Yi-Fan Yeh, Minjing Dong, Tao Huang, Philip Torr, Chang Xu

机构 * School of Computer Science University of Sydney(悉尼大学计算机科学学院) City University of Hong Kong(香港城市大学) Shanghai Jiao Tong University(上海交通大学) Department of Engineering Science University of Oxford(牛津大学工程科学系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22655 2025-05-29 cs.LG cs.AI cs.CL 67%

Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

Michael Kirchhof, Gjergji Kasneci, Enkelejda Kasneci

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17671 2025-05-16 cs.CL cs.AI cs.LG 67%

Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction

Yuanchang Ye, Weiyan Wen

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by ICIPCA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06529 2025-02-12 cs.AI cs.CL cs.LG 67%

Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity

Kaiqu Liang, Zixu Zhang, Jaime Fernández Fisac

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04428 2025-02-10 cs.CL cs.AI cs.LG 67%

Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization

Yu-Neng Chuang, Leisheng Yu, Guanchu Wang, Lizhe Zhang, Zirui Liu, Xuanting Cai, Yang Sui, Vladimir Braverman, Xia Hu

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12215 2025-01-28 cs.LG cs.CL cs.CY 67%

SoK: Machine Learning for Misinformation Detection

Madelyne Xiao, Jonathan Mayer

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07961 2024-12-12 cs.CL cs.AI cs.LG 67%

Forking Paths in Neural Text Generation

Eric Bigelow, Ari Holtzman, Hidenori Tanaka, Tomer Ullman

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00760 2024-12-03 eess.AS cs.AI cs.CL cs.ET cs.LG 67%

Automating Feedback Analysis in Surgical Training: Detection, Categorization, and Assessment

Firdavs Nasriddinov, Rafal Kocielnik, Arushi Gupta, Cherine Yang, Elyssa Wong, Anima Anandkumar, Andrew Hung

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted as a proceedings paper at Machine Learning for Health 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23054 2024-11-25 cs.LG cs.AI cs.CL cs.CV 67%

Controlling Language and Diffusion Models by Transporting Activations

Pau Rodriguez, Arno Blaas, Michal Klein, Luca Zappella, Nicholas Apostoloff, Marco Cuturi, Xavier Suau

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00499 2024-11-19 cs.CL cs.AI cs.LG 67%

ConU: Conformal Uncertainty in Large Language Models with Correctness Coverage Guarantees

Zhiyuan Wang, Jinhao Duan, Lu Cheng, Yue Zhang, Qingni Wang, Xiaoshuang Shi, Kaidi Xu, Hengtao Shen, Xiaofeng Zhu

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14259 2024-11-19 cs.CL cs.AI cs.LG 67%

Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond

Zhiyuan Wang, Jinhao Duan, Chenxi Yuan, Qingyu Chen, Tianlong Chen, Yue Zhang, Ren Wang, Xiaoshuang Shi, Kaidi Xu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by Engineering Applications of Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04331 2024-11-06 cs.CL cs.AI cs.IR cs.LG 67%

PaCE: Parsimonious Concept Engineering for Large Language Models

Jinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Darshan Thaker, Aditya Chattopadhyay, Chris Callison-Burch, René Vidal

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted in NeurIPS 2024. GitHub repository at https://github.com/peterljq/Parsimonious-Concept-Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02062 2024-10-28 cs.CL cs.AI cs.LG 67%

Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?

Wataru Hashimoto, Hidetaka Kamigaito, Taro Watanabe

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16592 2024-10-23 cs.LG cs.CL cs.CY 67%

ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding

Andrew Kan, Christopher Kan, Zaid Nabulsi

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10474 2024-08-21 cs.SE cs.AI cs.CL cs.CR cs.LG 67%

LeCov: Multi-level Testing Criteria for Large Language Models

Xuan Xie, Jiayang Song, Yuheng Huang, Da Song, Fuyuan Zhang, Felix Juefei-Xu, Lei Ma

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07453 2024-08-15 cs.CL cs.AI cs.LG 67%

Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals

Tobias A. Opsahl

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 10 pages, 3 figures, appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12836 2024-07-19 cs.CL cs.AI cs.LG 67%

OSPC: Artificial VLM Features for Hateful Meme Detection

Peter Grönquist

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00030 2024-06-04 cs.CL cs.AI cs.LG 67%

Large Language Model Pruning

Hanjuan Huang, Hao-Jia Song, Hsing-Kuo Pao

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 17 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20003 2024-05-31 cs.LG cs.AI cs.CL 67%

Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities

Alexander Nikitin, Jannik Kossen, Yarin Gal, Pekka Marttinen

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13788 2024-04-25 cs.CL cs.AI cs.CR cs.HC cs.LG 67%

Can LLM-Generated Misinformation Be Detected?

Canyu Chen, Kai Shu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Proceedings of ICLR 2024. 9 pages for main paper, 40 pages including appendix. The code, results, dataset for this paper and more resources on "LLMs Meet Misinformation" have been released on the project website: https://llm-misinformation.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08705 2024-04-16 cs.CL cs.AI cs.LG 67%

Introducing L2M3, A Multilingual Medical Large Language Model to Advance Health Equity in Low-Resource Regions

Agasthya Gangavarapu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05973 2024-03-12 cs.CL cs.AI cs.LG 67%

Calibrating Large Language Models Using Their Generations Only

Dennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun, Seong Joon Oh

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03170 2024-03-12 cs.MM cs.AI cs.CL cs.CV cs.CY 67%

SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

Peng Qi, Zehong Yan, Wynne Hsu, Mong Li Lee

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments To appear in CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09114 2024-02-27 cs.CL cs.AI cs.LG 67%

Ever: Mitigating Hallucination in Large Language Models through Real-Time Verification and Rectification

Haoqiang Kang, Juntong Ni, Huaxiu Yao

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09731 2024-02-19 cs.CL cs.AI cs.LG 67%

Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Genglin Liu, Xingyao Wang, Lifan Yuan, Yangyi Chen, Hao Peng

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07023 2024-02-13 cs.CL cs.AI cs.CV cs.HC cs.LG 67%

Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations

Ankit Pal, Malaikannan Sankarasubbu

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint version, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04782 2023-12-05 cs.CL cs.AI cs.LG 67%

HistAlign: Improving Context Dependency in Language Generation by Aligning with History

David Wan, Shiyue Zhang, Mohit Bansal

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2023 (20 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08401 2023-11-15 cs.CL cs.AI cs.LG 67%

Fine-tuning Language Models for Factuality

Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, Chelsea Finn

专题命中 幻觉与事实性 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10346 2023-09-20 cs.LG cs.AI cs.CL 67%

Explaining Agent Behavior with Large Language Models

Xijia Zhang, Yue Guo, Simon Stepputtis, Katia Sycara, Joseph Campbell

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Human Multi-Robot Interaction Workshop at IROS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏