arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-05-29 至 2026-05-29 共收录 535 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 29 篇

2605.29864 2026-05-29 cs.RO 89%

LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation

LLM引导的未来假设用于多步机器人操作中的视野感知探索

Mohammad Khoshnazar, Andrew Melnik, Michael Beetz

机构 * Institute of Artificial Intelligence, University of Bremen(人工智能研究所,不莱梅大学)

专题命中 知识编辑与模型理解 :LLM(title,title_cn)

AI总结 提出未来经验条件化(FEC)框架,利用LLM生成短期未来视频作为结构化先验,结合行为克隆和强化学习微调,提升多步机器人操作中的探索和策略适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29354 2026-05-29 cs.CR cs.LG 89%

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

无害却有害:针对Agent技能中隐蔽幻觉引导的中性提示攻击

Chia-Yi Hsu, Chia-Mu Yu, Chun-Ying Huang, Jun Sakuma

机构 * Department of Computer Science(计算机科学系) National Yang Ming Chiao Tung University(阳明交通大学) Department of Electronics and Electrical Engineering(电子与电气工程系) School of Computing(计算学院) Institute of Science Tokyo(东京科学研究所)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);prompting(title,abstract);分类 cs.LG

AI总结 本文提出中性提示攻击(NPA),通过语义上看似无害的指令(如鼓励想象和详尽性)增加代码生成Agent的包幻觉倾向,从而引入软件供应链风险,并评估了其对多种编码LLM的有效性和逃避防御的能力。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29826 2026-05-29 cs.CL cs.AI 88%

Towards Localized and Disentangled Knowledge Editing for Multimodal Large Language Models

面向多模态大语言模型的局部化与解耦知识编辑

Leijiang Gu, Zhen Zeng, Feng Li, Xinjian Gao, Zenglin Shi

机构 * Hefei University of Technology(合肥工业大学) Tongji University(同济大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 针对多模态知识编辑中因果错位和特征纠缠问题,提出LDKE框架,通过快速定位关键层和解耦分类器实现精准泛化编辑并保持高局部性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28828 2026-05-29 cs.CL cs.AI 88%

Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models

微宏检索:减少大语言模型中的长文本幻觉

Yujie Feng, Jian Li, Zhihan Zhou, Pengfei Xu, Yujia Zhang, Xiaoyu Li, Xiaohui Zhou, Alan Zhao, Xi Chen, Xiao-Ming Wu

机构 * Solar System of OVB, Tencent, China(OVB太阳系,腾讯,中国) The Hong Kong Polytechnic University, Hong Kong S.A.R.(香港理工大学,香港特别行政区) Jilin University, China(吉林大学,中国)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 提出微宏检索(M2R)框架,通过宏观检索外部粗粒度证据和微观检索推理中关键信息库,解决长文本生成中关键信息与输出距离过远导致的幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28825 2026-05-29 cs.CL 88%

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

MechELK:一种用于激发大型语言模型中潜在知识的机制可解释性框架

Ji-jun Park, Soo-joon Choi, Jiwon Jeong, Taeyang Yoon, Ju-Wan Lee

机构 * Dongguk University(东国大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 提出MechELK框架,通过定位、验证和激发三个阶段,利用稀疏自编码器特征分析和因果探测等方法,从大型语言模型中提取隐藏知识,在TruthfulQA等基准上平均激发准确率达84.7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03134 2026-05-29 cs.CL 86%

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

对话式诈骗的剖析:基于主题的LLM多轮交互红队分析

Xiangzhe Yuan, Zhenhao Zhang, Haoming Tang, Siying Hu

机构 * Department of Computer Science, University of Iowa(爱荷华大学计算机科学系) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)

专题命中 知识编辑与模型理解 :LLM(title_cn,summary_cn);分类 cs.CL

AI总结 通过LLM间模拟框架研究多轮社交工程对话中的对抗动态,分析攻击与防御策略,发现跨模型和跨语言的结果差异及策略转换的结构性变化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16178 2026-05-29 cs.CL 85%

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

理解语言模型中的事实回忆:为什么两阶段训练鼓励记忆而混合训练教授知识

Ying Zhang, Benjamin Heinzerling, Dongyuan Li, Kentaro Inui

机构 * RIKEN Center for Advanced Intelligence Project(日本理化学研究所高级智能项目中心) Tohoku University(东北大学) The University of Tokyo(东京大学) MBZUAI

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract_cn);large language model(abstract);分类 cs.CL

AI总结 通过比较2.8~4B语言模型中的两阶段训练与混合训练,发现混合训练通过联合优化目标实现存储与查询格式间的梯度一致性,驱动表征一致性并建立格式不变的检索过程,从而泛化回忆未见查询中的事实。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28837 2026-05-29 cs.CL cs.AI 85%

SERC: LDPC-Inspired Semantic Error Correction for Retrieval-Augmented Generation

SERC: 受LDPC启发的检索增强生成语义纠错方法

Gyumin Kim, Juhwan Park, Jaeha Kim, Seunggyun Han, Kyungrak Son, Ikbeom Jang

机构 * Department of Information Communications Engineering, Hankuk University of Foreign Studies, Republic of Korea(韩国外国语大学信息通信工程系) Division of Computer Engineering, Hankuk University of Foreign Studies, Republic of Korea(韩国外国语大学计算机工程系) Department of Statistics, Hankuk University of Foreign Studies, Republic of Korea(韩国外国语大学统计学系)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 针对大语言模型幻觉问题,提出受LDPC码启发的语义纠错框架SERC,通过稀疏验证策略高效检测和纠正生成文本中的错误。

Comments 15 pages, 2 figures, 6 tables. To appear in the Proceedings of the 28th International Conference on Pattern Recognition (ICPR 2026). Code available at https://github.com/labhai/SERC

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29395 2026-05-29 stat.ME stat.ML 85%

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons

低秩排序:稀疏成对比较下不确定性感知的任务特定大语言模型排名

Jiachun Li, David Simchi-Levi, Will Wei Sun

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 提出一种低秩框架,通过稀疏成对比较进行任务特定的大语言模型排名,利用任务-模型能力矩阵的低秩结构实现跨任务信息共享,并开发了不确定性量化方法以提供置信区间和排名证书。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04765 2026-05-29 cs.CL cs.AI cs.LG physics.comp-ph 83%

Differential syntactic and semantic encoding in LLMs

大型语言模型中句法与语义的差异编码

Santiago Acevedo, Alessandro Laio, Marco Baroni

机构 * Catalan Institute of Research and Advanced Studies (ICREA) and Universitat Pompeu Fabra (UPF)(加泰罗尼亚研究与高级科学研究所(ICREA)和庞培法华大学(UPF))

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过平均共享句法结构或语义的句子隐藏表示向量,发现大型语言模型(以DeepSeek-V3为例)的内部层表示中句法和语义信息至少部分线性编码,且两者编码轮廓不同,可一定程度解耦。

Comments Published as conference paper at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30042 2026-05-29 cs.AI 81%

Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection

学会选择:一种基于赋权与语义通信的自适应方法选择多智能体系统

Geremy Loachamín-Suntaxi, Robert Lazar, Dimitrios G. Giovanis, Ioannis G. Kevrekidis, Eleni D. Koronaki

机构 * Faculty of Science, Technology and Medicine(科学、技术与医学学院) University of Luxembourg(卢森堡大学) Johns Hopkins University(约翰霍普金斯大学) Luxembourg Institute of Science and Technology(卢森堡科学与技术研究院)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种结合上下文赌博机、结构化智能体间通信和语义检查点的多智能体框架,通过保持动作-结果因果一致性来提升科学计算工作流中自适应决策的收敛性、鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29737 2026-05-29 cs.CR cs.CL cs.SE 81%

Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

最小提示扰动导致代码漏洞:编码大语言模型中的提示脆弱性和隐藏状态信号

Alexander Sternfeld, Andrei Kucharavy, Ljiljana Dolamic

机构 * IEM, HES-SO, Le Foyer, Techno-Pôle 1, Sierre, Switzerland(瑞士苏黎世联邦理工学院(HES-SO)技术园区1号,西尔尔) Cyber-Defence Campus, armasuisse Science and Technology, Thun, Switzerland(瑞士图恩 Cyber-Defence 营地,armasuisse 科学与技术)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 本文通过token级突变实验,发现微小提示扰动(如单字符变化)即可使LLM生成代码从安全变为脆弱,并利用隐藏状态分析揭示输入处理漏洞比安全默认值漏洞更可预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29881 2026-05-29 cs.CV cs.AI 79%

Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering

通过屏障调控自适应闭式引导缓解视觉语言模型中的幻觉

Soumyadeep Jana, Pulkit Mittal, Sanasam Ranbir Singh

机构 * Indian Institute of Technology Guwahati(印度理工学院果阿班加)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 提出BRACS框架,通过监测视觉注意力并仅在接地退化时进行闭式修正,无需训练即可有效减少LVLM中的物体幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23853 2026-05-29 cs.AI cs.MA 79%

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

SCoOP: 多视觉-语言模型系统中用于不确定性量化的语义一致意见池化

Chung-En Johnny Yu, Brian Jalaian, Nathaniel D. Bastian

机构 * University of West Florida(西佛罗里达大学) United States Military Academy(美国军事学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 提出SCoOP框架,通过不确定性加权的线性意见池化聚合多个视觉-语言模型的输出,实现无训练的不确定性量化,有效检测幻觉并支持高不确定性样本的弃权。

Comments Accepted to ICLR 2026 Workshop on Agentic AI in the Wild: From Hallucinations to Reliable Autonomy

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30161 2026-05-29 cs.CV 78%

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

为什么远处看起来在上方:探究视觉-语言模型中的空间表征

Cheolhong Min, Jaeyun Jung, Daeun Lee, Hyeonseong Jeon, Yu Su, Jonathan Tremblay, Chan Hee Song, Jaesik Park

机构 * Seoul National University(首尔国立大学) The Ohio State University(俄亥俄州立大学) NVIDIA(英伟达)

专题命中 知识编辑与模型理解 :language model(title,abstract)

AI总结 通过最小对比对分析,发现视觉-语言模型存在垂直-距离纠缠(将图像垂直位置与距离混淆),这种透视偏差导致性能差距,并随数据规模扩大而加剧,而具有良好分离空间轴的模型更鲁棒。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29956 2026-05-29 cs.IR 75%

Uncertainty Quantification for Multimodal Retrieval Augmented Generation

多模态检索增强生成的不确定性量化

Simon Binz, Heydar Soudani, Faegheh Hasibi

专题命中 知识编辑与模型理解 :LLM(abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出 LeMUQ 方法,通过多模态和检索感知的概率信号建模不确定性,提升多模态 RAG 系统的可靠性,AUROC 平均提升 3.8%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27272 2026-05-29 cs.CL cs.AI cs.LG 75%

When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks

当2D任务遇到1D序列化:结构化任务中的序列化摩擦

Chung-Hsiang Lo, Lu Li, Diji Yang, Tianyu Zhang, Yunkai Zhang, Yoshua Bengio, Yi Zhang

机构 * Northeastern University(东北大学) University of Pennsylvania(宾夕法尼亚大学) UC Santa Cruz(加州大学圣克鲁兹分校) Mila - Quebec AI Institute(魁北克人工智能研究所) University of Montreal(蒙特利尔大学) BAIR, UC Berkeley(伯克利大学BAIR实验室)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过矩阵转置、康威生命游戏和LU分解三个任务,发现将二维布局任务序列化为一维文本会因表示不匹配导致性能下降,且错误呈现空间结构模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28920 2026-05-29 cs.LG cs.AI stat.ML 73%

Conf-Gen: Conformal Uncertainty Quantification for Generative Models

Conf-Gen: 生成模型的共形不确定性量化

Gabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc T. Law, Kin Kwan Leung

机构 * layer6ai-labs(layer6ai实验室)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出Conf-Gen框架,通过共形风险控制适配生成任务,统一并扩展了共形预测在大型语言模型等生成模型中的应用,并在图像生成、对话AI和AI代理等新领域提供了形式化保证。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28990 2026-05-29 cs.LG 70%

Learning Robust and Task-Invariant Functional Representation from fMRI through Siamese Self-Supervised Learning

通过孪生自监督学习从fMRI中学习鲁棒且任务不变的功能表示

Jiyao Wang, Peiyu Duan, Nicha C. Dvornek, Lawrence H. Staib, Denis Sukhodolsky, Pamela Ventola, James S. Duncan

机构 * organization= Department of Biomedical Engineering , addressline= Yale University , city= New Haven , state= CT , country= USA organization= Radiology \& Biomedical Imaging , addressline= Yale School of Medicine , city= New Haven , state= CT , country= USA organization= Electrical Engineering , addressline= Yale University , city= New Haven , state= CT , country= USA organization= Child Study Center , addressline= Yale School of Medicine , city= New Haven , state= CT , country= USA

专题命中 知识编辑与模型理解 :foundation model(abstract);pretraining(abstract);分类 cs.LG

AI总结 提出轻量级自监督框架BrainSimSiam,利用正样本对学习鲁棒且通用的fMRI表示,在多个下游任务中超越全监督基线,接近大规模模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02116 2026-05-29 cs.LG 70%

Statistical Consistency and Generalization of Contrastive Representation Learning

对比表示学习的统计一致性与泛化性

Yuanfan Li, Xiyuan Wei, Tianbao Yang, Yiming Ying

机构 * University of Sydney Texas A\&M University

专题命中 知识编辑与模型理解 :language model(abstract);foundation model(abstract);分类 cs.LG

AI总结 本文提出统一的统计学习理论,证明对比损失与最优排序统计一致,并推导出随负样本数增加而改善的泛化界,解释了大负样本集的经验优势。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29628 2026-05-29 cs.SD cs.AI cs.CL cs.LG eess.AS 67%

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

COMET:音频-文本多模态对比嵌入中模态间隙的概念空间剖析

Yonggang Zhu, Liting Gao, Aidong Men, Wenwu Wang

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Centre for Vision, Speech, and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)

专题命中 知识编辑与模型理解 :pretraining(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出COMET框架,通过PLS-SVD分解揭示CLAP模型中模态间隙主要由少数共享概念轴贡献,并基于谱截断方法无训练地缓解间隙,实现零样本音频字幕接近全监督性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08722 2026-05-29 cs.LG cs.AI 62%

The Impact of Semantic Pairs on Self-Supervised Representation Learning

语义对自监督表示学习的影响

Mohammad Alkhalefi, Georgios Leontidis, Mingjun Zhong

机构 * Department of Computing Science University of Aberdeen(计算科学系大学阿伯丁) Department of Physics and Technology UiT The Arctic University of Norway(物理与技术系UiT北极大学)

专题命中 知识编辑与模型理解 :pretraining(abstract);分类 cs.AI、cs.LG

AI总结 通过控制实验研究语义正对(不同同类实例)相比增强正对在自监督学习中的效果,发现语义对能提升泛化性能,尤其对比学习受益最大。

Comments 19 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28870 2026-05-29 cs.LG cs.AI 62%

Representation Alignment Rests on Linear Structure

表示对齐依赖于线性结构

Kiril Bangachev, Guy Bresler, Yury Polyanskiy

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 知识编辑与模型理解 :LLM(abstract_cn);分类 cs.AI、cs.LG

AI总结 本文通过信号、偏差和噪声的三部分统计框架研究柏拉图表示假说,提出对齐源于对象与属性的线性关系,并通过稀疏自编码器提取线性特征、中心化和归一化减少偏差、以及数据稀缺导致噪声等证据支持该框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27518 2026-05-29 cs.CL 57%

Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs

过度拒绝与表示子空间:对齐大语言模型中任务条件拒绝的机制分析

Utsav Maskey, Mark Dras, Usman Naseem

机构 * Macquarie University(麦考瑞大学)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL

AI总结 通过分析有害拒绝和过度拒绝的表示几何,发现过度拒绝方向是任务相关的且存在于良性任务表示簇中,解释了为何全局方向消融无法解决过度拒绝,并表明需要任务特定的几何干预。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13600 2026-05-29 cs.CV 50%

SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification

SAVAA: 通过逐步自适应视觉注意力放大减轻LVLMs中的幻觉

Jiacheng Zhang, Feng Liu, Chao Du, Tianyu Pang

机构 * Sea AI Lab(海思人工智能实验室) The University of Melbourne(墨尔本大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 提出SAVAA框架,通过视觉接地熵估计幻觉风险并自适应调整视觉注意力放大因子,在多个基准上显著减轻大型视觉语言模型的幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他LLM 30 篇

2605.28840 2026-05-29 cs.CL cs.AI cs.SE 92%

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

LLM代理的一致性如何?测量多步工具调用流水线中的行为可重复性

Abel Yagubyan

机构 * Independent Researcher(独立研究者)

专题命中 其他LLM :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究多步工具调用LLM代理在重复相同调用时是否选择相同工具、顺序和参数,通过系统实验测量行为一致性,并发现代理存在显著不一致性。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30157 2026-05-29 stat.AP 92%

Leveraging Large Language Models to Improve Precision in Randomized Controlled Trials

利用大型语言模型提高随机对照试验的精度

Jaylin Lowe, Adam Sales, Johann A. Gagnon-Bartsch

专题命中 其他LLM :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract)

AI总结 本文探索如何安全、严谨地利用大型语言模型(LLM)的预测来提升随机对照试验(RCT)的精度,并通过三个案例验证其有效性。

Comments Submitted to Machine Learning and Artificial Intelligence for Causal Inference in the Behavioral and Social Sciences: Methodological Advances and Applications, a topical issue of the Zeitschrift für Psychologie

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30096 2026-05-29 cs.CR cs.AI 92%

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

AI攻击者对固定脆弱目标的可靠性如何?LLM渗透测试一致性的400次运行实证研究

Galip Tolga Erdem

机构 * Independent Researcher(独立研究者)

专题命中 其他LLM :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 通过400次自主渗透测试运行(4个模型各100次),研究LLM在固定目标上攻击行为的一致性,发现模型间成功率差异显著且失败模式独特。

Comments 41 pages, 7 figures. Code and 400-run dataset: https://doi.org/10.5281/zenodo.20421592

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26506 2026-05-29 cs.CL cs.CR 92%

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

SafeReview: 防御基于LLM的评审系统免受对抗性隐藏提示攻击

Yuan Xin, Yixuan Weng, Minjun Zhu, Ying Ling, Chengwei Qin, Michael Backes, Yue Zhang, Linyi Yang

机构 * CISPA Westlake University(西交利物浦大学) Southern University of Science and Technology(南方科技大学) HKUST (Guangzhou)(香港科技大学(广州))

专题命中 其他LLM :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出SafeReview,一种共进化对抗训练框架,通过联合训练生成器和防御者模型,增强基于LLM的同行评审系统对对抗性隐藏提示的鲁棒性。

Comments 17 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29062 2026-05-29 cs.CL 92%

Bosses, Kings, and the Commons: Cooperation Under Power Asymmetry in LLM Societies

老板、国王与公地:LLM 社会中权力不对称下的合作

Abhilekh Borah

专题命中 其他LLM :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究通过引入不对称权力代理(老板或国王)的多智能体模拟框架 SovSim,发现权力不对称导致 LLM 社会中合作与可持续性严重崩溃,生存率较对称设置下降高达 87.3%。

Comments Paper under review

详情

展开后加载摘要…

URL PDF HTML 收藏