arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 527 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 527 篇

2601.10173 2026-01-16 cs.CR cs.AI cs.CL 92%

ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack

ReasAlign: 基于推理增强的安全对齐以抵御提示注入攻击

Hao Li, Yankai Yang, G. Edward Suh, Ning Zhang, Chaowei Xiao

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) NVIDIA(NVIDIA公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 提示注入 :alignment(title,abstract);safety(title,abstract);prompt injection(title,abstract);分类 cs.CL、cs.AI

AI总结 ReasAlign通过结构化推理和测试时缩放机制,有效防御提示注入攻击,实现安全与效用的最佳平衡。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14827 2025-09-16 cs.CR cs.AI cs.CL cs.LG 89%

Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment

Zedian Shao, Hongbin Liu, Jaden Mu, Neil Zhenqiang Gong

机构 * Georgia Institute of Technology(佐治亚理工学院) Duke University(杜克大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 提示注入 :alignment(title,abstract);prompt injection(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16682 2024-12-24 cs.CR cs.AI cs.CL cs.LG 89%

The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents

Feiran Jia, Tong Wu, Xin Qin, Anna Squicciarini

专题命中 提示注入 :alignment(title,abstract);prompt injection(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06477 2026-08-10 cs.CR cs.AI cs.CL 新提交 88%

StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

StepJack:针对多步骤间接提示注入的计算机使用智能体安全基准测试

Zhuoxin Zhan, Akbar Rafiey, Avery Ma, Leila Pishdad, Layla El Asri

专题命中 提示注入 :safety(title,abstract);prompt injection(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出多步骤间接提示注入新型攻击,构建含480个测试样例的StepJack基准,评估六款先进CUAs,发现多步骤攻击可提升部分CUAs的攻击成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15441 2026-06-16 cs.CR cs.AI 新提交 88%

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

通过推理启用的任务对齐防御自适应提示注入攻击

Lipeng He, Yihan Wang, Jiawen Zhang, N. Asokan

机构 * University of Waterloo(滑铁卢大学) Zhejiang University(浙江大学) KTH Royal Institute of Technology(皇家理工学院)

专题命中 提示注入 :prompt injection(title,abstract);alignment(title);safety(abstract);分类 cs.AI

AI总结 提出RETA方法,通过基于用户任务的多目标强化学习训练防御器,利用思维链推理验证行动一致性,并采用字典学习多样性奖励生成对抗样本,在六种自适应攻击下平均攻击成功率低于4%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09315 2026-06-09 cs.CR cs.AI 新提交 88%

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

脑提示注入:BCI-LLM代理的路径安全审计

Jianwei Tai

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 提示注入 :safety(title,abstract);prompt injection(title,abstract);分类 cs.AI

AI总结 提出路径安全审计契约,通过分离定理和共形校准量化BCI-LLM代理中脑信号注入攻击的风险,实验证明确认通道可降低路由风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05566 2026-06-05 cs.AI cs.CR 88%

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

GuardNet: 用于鲁棒提示注入和越狱检测的浅层神经网络集成策略

Paulo Ricardo Ferreira Neves, Edson Rodrigues da Cruz Filho, Paulo Henrique Eleuterio Falsetti, João Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague, Patrick Vieira Laturrague, Guilherme Nielsen Dias, Marccello Wilson Perez Berto, Gustavo Voltani Von Atzingen

机构 * Quickium Technology Ltd.(Quickium技术有限公司) Federal University of São Carlos (UFSCar)(萨尔瓦多·卡罗斯联邦大学) Federal Institute of Education, Science and Technology of São Paulo (IFSP)(圣保罗教育、科学和技术联邦研究所)

专题命中 提示注入 :jailbreak(title,abstract);prompt injection(title,abstract);分类 cs.AI

AI总结 提出GuardNet,一种基于浅层神经网络(BiLSTM)集成的护栏系统,通过多样性示例覆盖和阈值校准实现对抗鲁棒性,在低延迟下达到与轻量检测器竞争的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17634 2026-05-19 cs.CR cs.CL cs.CY 88%

AI Agents May Always Fall for Prompt Injections

AI Agents May Always Fall for Prompt Injections

Sahar Abdelnabi, Eugene Bagdasarian

机构 * ELLIS Institute Tübingen & MPI-IS & Tübingen AI Center(图宾根ELLIS研究所及MPI-IS与图宾根人工智能中心) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 提示注入 :prompt injection(title,title_cn);alignment(abstract);分类 cs.CL、cs.CY

AI总结 本文基于上下文完整性理论重新审视提示注入问题,揭示了现有防御机制的不足,并提出了一种新的评估框架来设计更安全的自主代理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13236 2024-10-18 cs.CL cs.AI 88%

SPIN: Self-Supervised Prompt INjection

Leon Zhou, Junfeng Yang, Chengzhi Mao

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract);jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01462 2026-05-05 cs.CR 88%

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training

LocalAlign: 通过生成接近目标的对抗示例实现通用的提示注入防御

Yuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng Wang

专题命中 提示注入 :alignment(title,abstract);prompt injection(title,abstract)

AI总结 LocalAlign通过生成接近正确响应但错误的对抗示例,提升提示注入防御的泛化能力,采用margin-aware对齐算法增强训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05206 2026-04-15 cs.CR cs.AI cs.CL cs.CV 87%

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

大规模安全:大型模型和智能体安全的全面综述

Xingjun Ma, Yifeng Gao, Yixu Wang, Ruofan Wang, Xin Wang, Ye Sun, Yifan Ding, Hengyuan Xu, Yunhao Chen, Yunhan Zhao, Hanxun Huang, Yige Li, Yutao Wu, Jiaming Zhang, Xiang Zheng, Yang Bai, Zuxuan Wu, Xipeng Qiu, Jingfeng Zhang, Yiming Li, Xudong Han, Haonan Li, Jun Sun, Cong Wang, Jindong Gu, Baoyuan Wu, Siheng Chen, Tianwei Zhang, Yang Liu, Mingming Gong, Tongliang Liu, Shirui Pan, Cihang Xie, Tianyu Pang, Yinpeng Dong, Ruoxi Jia, Yang Zhang, Shiqing Ma, Xiangyu Zhang, Neil Gong, Chaowei Xiao, Sarah Erfani, Tim Baldwin, Bo Li, Masashi Sugiyama, Dacheng Tao, James Bailey, Yu-Gang Jiang

机构 * Fudan University(复旦大学) The University of Melbourne(墨尔本大学) Singapore Management University(新加坡国立大学) Deakin University(德肯大学) Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) University of Oxford(牛津大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) The University of Sydney(悉尼大学) Griffith University(格里菲斯大学) University of California, Santa Cruz(加州大学圣克鲁兹分校) Sea AI Lab(Sea AI实验室) Tsinghua University(清华大学) Virginia Tech(弗吉尼亚理工大学) CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全部) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) Purdue University(普渡大学) Duke University(杜克大学) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) RIKEN(理化学研究所) The University of Tokyo(东京大学)

专题命中 提示注入 :safety(title,abstract);jailbreak(abstract);prompt injection(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了大型模型和智能体的安全性,分析了各类攻击威胁及防御策略,指出安全评估、防御机制和数据实践的重要性,强调研究社区和国际合作的必要性。

Comments 706 papers, 60 pages, 3 figures, 14 tables; GitHub: https://github.com/xingjunm/Awesome-Large-Model-Safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13026 2026-07-24 cs.LG cs.CR 版本更新 86%

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

PISmith: 基于强化学习的提示注入防御红队测试

Chenlong Yin, Runpeng Geng, Yanting Wang, Jinyuan Jia

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 提示注入 :prompt injection(title,abstract);red teaming(title);分类 cs.LG

AI总结 本文提出PISmith框架,通过强化学习评估现有提示注入防御的鲁棒性,发现标准GRPO在攻击强防御时表现不佳,引入自适应熵正则化和动态优势加权以提升探索能力,实验表明最新防御仍易受适应性攻击影响。

Comments To appear in COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23557 2025-12-30 cs.CR cs.AI 86%

Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks

迈向可信的代理AI:一种多模态框架用于防止提示注入攻击

Toqeer Ali Syed, Mishal Ateeq Almutairi, Mahmoud Abdel Moaty

机构 * Faculty of Computer and Information System(计算机与信息系统系) Islamic University of Madinah(麦地那伊斯兰大学) Arab Open University-Bahrain(巴林阿拉伯开放大学)

专题命中 提示注入 :prompt injection(title,abstract);trustworthy(title);分类 cs.AI

AI总结 本文提出一种多模态框架,通过溯源感知机制防止代理AI中的提示注入攻击,提升系统安全性和稳定性。

Comments It is accepted in a conference paper, ICCA 2025 in Bahrain on 21 to 23 December

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13597 2026-02-24 cs.CR 86%

AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks

AlignSentinel: 一种考虑对齐的提示注入攻击检测方法

Yuqi Jia, Ruiqi Wang, Xilong Wang, Chong Xiang, Neil Gong

专题命中 提示注入 :prompt injection(title,abstract);alignment(title)

AI总结 AlignSentinel通过分析LLM注意力图,有效区分包含错位指令、对齐指令和非指令输入,提升提示注入攻击检测的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21151 2026-07-28 cs.AI 版本更新 85%

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

V-DEAL:将视频安全去校准诊断为理解-拒绝耦合失败

Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai

机构 * University of Queensland(昆士兰大学) University of California, Merced(加州大学默塞德分校)

专题命中 提示注入 :safety(title,abstract);alignment(abstract);prompt injection(abstract);分类 cs.AI

AI总结 研究视频大语言模型安全对齐问题,提出V-DEAL三级诊断框架分析漏洞机制,通过测试发现模型识别有害视频内容有一定准确率,但配对有害视频与良性查询时攻击成功率高,视觉理解激活拒绝倾向弱,还引入提示注入干预方法降低攻击成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10577 2026-04-20 cs.CR cs.AI 85%

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

智能体安全的盲点:良性用户指令如何暴露计算机使用智能体的关键漏洞

Xuwei Ding, Skylar Zhai, Linxin Song, Jiate Li, Taiwei Shi, Nicholas Meade, Siva Reddy, Jian Kang, Jieyu Zhao

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of Minnesota(明尼苏达大学) University of Southern California(南加州大学) McGill University(麦吉尔大学) Mila(Mila研究院) MBZUAI(MBZUAI研究院)

专题命中 提示注入 :safety(title,abstract);alignment(abstract);prompt injection(abstract);分类 cs.AI

AI总结 研究揭示了良性用户指令下智能体安全漏洞,提出OS-BLIND基准测试,发现多数智能体在攻击条件下成功率高达90%以上,且在多智能体系统中风险加剧,现有安全措施效果有限。

Comments 63 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18514 2026-02-24 cs.CR cs.AI 85%

Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models

招聘中的木马:针对标准与推理模型间接提示注入的红队案例研究

Manuel Wirth

机构 * University of Mannheim(曼海姆大学)

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究通过红队测试揭示了标准与推理模型在间接提示注入中的安全差异,发现推理模型在复杂指令下易出现元认知泄漏,而标准模型在简单攻击中表现较弱。

Comments 43 pages, 3 synthetic CV PDF's, 6 chat history PDF's and system prompts. This work was developed as part of the Responsible AI course within the Mannheim Master in Data Science (MMDS) program at the University of Mannheim

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08290 2025-12-16 cs.CR cs.AI 85%

Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem

知识体系化:模型上下文协议生态系统中的安全与安全

Shiva Gaire, Srijan Gyawali, Saroj Mishra, Suman Niroula, Dilip Thakur, Umesh Yadav

机构 * Tribhuvan University(特里布文大学) University of North Dakota(北达科他大学) Youngstown State University(亚当斯州立大学) University of Missouri(密苏里大学) University of Toledo(托莱多大学)

专题命中 提示注入 :safety(title,abstract);alignment(abstract);prompt injection(abstract);分类 cs.AI

AI总结 本文系统化分析了模型上下文协议生态系统中的安全与安全风险,提出了全面的风险分类,并探讨了从对话聊天机器人到自主代理操作系统安全过渡的路线图。

Comments All authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01634 2025-11-13 cs.CR cs.AI 85%

Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models

Daniyal Ganiuly, Assel Smaiyl

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22545 2026-07-28 cs.LG cs.AI cs.CL cs.CR 新提交 85%

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

Semalith v1.4:一个经过校准的1.84亿参数安全分类器,在参数比Llama - Guard - 3 - 8B少44倍的情况下实现了先进的提示注入检测

Tejasvi C. Addagada

专题命中 提示注入 :safety(title,abstract);prompt injection(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究针对金融服务等场景中大语言模型安全分类需求,提出Semalith v1.4分类器,能单步实现三轴安全分类,经训练和对比测试,在提示注入检测上表现出色且参数少,还给出不同场景下的部署建议。

Comments 16 pages, 8 tables, no figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12624 2026-07-15 cs.CR 新提交 85%

PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis

PVDetector:通过策略违规概念分析检测针对特定目的大语言模型代理的提示注入攻击

Junhui Wang, Hangtao Zhang, Zhirun Zheng, Li Zeng, Jiejun Xiao, Xi Luo, Lihua Yin, Saiqin Long

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract)

AI总结 研究针对特定目的LLM代理的提示注入攻击,提出PVDetector框架,通过测量与离线派生的PV概念的隐藏状态对齐来检测攻击,实验表明该方法误报率低、开销小,性能优于现有方法。

Comments Accepted to ACM MM 2026. Code: https://github.com/Claresigle/PVDetector

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29602 2026-06-30 cs.CR 85%

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios

多语言与混淆攻击场景下大语言模型提示注入漏洞的实证评估

Caglar Uysal, Baturay Birinci, Süha Orhun Mutluergil, Orçun Çetin

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract)

AI总结 本文实证评估六种大语言模型在多语言和混淆攻击下的提示注入漏洞,发现所有模型均易受攻击,非英语语言恶意合规率更高,需加强安全防御。

Comments Accepted to the AI-SS 2026 Workshop at the 21st European Dependable Computing Conference (EDCC 2026). To be published in the EDCC Companion Proceedings (EDCC-C)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05442 2025-10-08 cs.LG cs.AI cs.CL 85%

Adversarial Reinforcement Learning for Large Language Model Agent Safety

Zizhao Wang, Dingcheng Li, Vaishakh Keshava, Phillip Wallis, Ananth Balashankar, Peter Stone, Lukas Rutishauser

机构 * Google(谷歌) Google Deepmind(谷歌DeepMind) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Sony AI(索尼人工智能)

专题命中 提示注入 :safety(title,abstract);prompt injection(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14348 2025-07-29 cs.CV 85%

Manipulating Multimodal Agents via Cross-Modal Prompt Injection

Le Wang, Zonghao Ying, Tianyuan Zhang, Siyuan Liang, Shengshan Hu, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北洋大学) National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) Henan University of Science and Technology(河南科技大学)

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract)

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09102 2025-05-28 cs.LG cs.AI cs.CL cs.CR 85%

Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy

Tong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu, Sanqiang Zhao, Ravi Agrawal, Sathish Reddy Indurthi, Chong Xiang, Prateek Mittal, Wenxuan Zhou

机构 * Princeton University(普林斯顿大学) Zoom Video Communications(Zoom视频通讯)

专题命中 提示注入 :safety(title,abstract);prompt injection(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint

Journal ref ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13435 2024-12-19 cs.CL cs.AI cs.LG 85%

Lightweight Safety Classification Using Pruned Language Models

Mason Sawtell, Tula Masterman, Sandi Besen, Jim Brown

专题命中 提示注入 :safety(title,abstract);prompt injection(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20597 2026-08-13 cs.LG cs.AI cs.CR 版本更新 84%

BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents

BrowseSafe: 理解和防止AI浏览器代理中的提示注入

Kaiyuan Zhang, Mark Tenenholtz, Kyle Polley, Jerry Ma, Denis Yarats, Ninghui Li

机构 * Purdue University(普渡大学) Perplexity AI

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI、cs.LG

AI总结 BrowseSafe通过构建现实场景下的提示注入攻击基准,提出多层次防御策略,旨在提升AI浏览器代理的安全性。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05746 2026-06-11 cs.LG cs.AI 版本更新 84%

Learning to Inject: Automated Prompt Injection via Reinforcement Learning

学习注入:通过强化学习实现自动化提示注入

Xin Chen, Jie Zhang, Florian Tramèr

机构 * ETH Zürich(苏黎世联邦理工学院)

专题命中 提示注入 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI、cs.LG

AI总结 提出AutoInject,一种基于强化学习的黑盒框架,自动学习对抗性后缀进行提示注入,在AgentDojo上优于模板攻击和多种自适应攻击,并突破专门防御模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12359 2026-06-08 cs.CR cs.AI cs.CL 交叉投稿 84%

Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

零样本嵌入漂移检测:一种轻量级防御对抗提示注入的LLM方法

Anirudh Sekar, Mrinal Agarwal, Rachel Sharma, Akitsugu Tanaka, Jasmine Zhang, Arjun Damerla, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院) Berkeley(伯克利大学)

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出ZEDD,通过量化嵌入空间中良性与可疑输入之间的语义变化,实现对直接和间接提示注入的检测。该方法无需模型内部访问或先验知识,具有低工程开销,能高效部署于多种LLM架构,准确率达93%以上。

Comments Accepted to NeurIPS 2025 Lock-LLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25415 2026-05-26 cs.CL cs.CY cs.ET 84%

LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

LLM-as-a-Reviewer: 基准测试它们作为论文审稿人的能力、分歧和提示注入抵抗性

Lingyao Li, Junjie Xiong, Changjia Zhu, Runlong Yu, Chen Chen, Junyu Wang, Renkai Ma, Zhicong Lu

机构 * University of South Florida(佛罗里达南大学) Missouri University of Science and Technology(密苏里科技大学) University of Alabama(阿拉巴马大学) Florida International University(佛罗里达国际大学) University of Cincinnati(辛辛那提大学) George Mason University(乔治·梅森大学)

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);分类 cs.CL、cs.CY

AI总结 本研究通过一个系统基准测试,评估了12个大型语言模型在论文评审中的表现,包括评分校准、与人类审稿人的分歧以及对不可见字体映射攻击的抵抗性,发现LLMs存在系统性高估弱论文、与人类关注点不同以及易受提示注入攻击等问题。

详情

展开后加载摘要…

URL PDF HTML 收藏