arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7539 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7539 篇

2606.29600 2026-06-30 cs.CV cs.AI 83%

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models

一场景,两深度:探究单目基础模型中的几何歧义性

Xiaohao Xu, Feng Xue, Xiang Li, Haowei Li, Shusheng Yang, Tianyi Zhang, Matthew Johnson-Roberson, Xiaonan Huang

机构 * University of Michigan(密歇根大学) Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学) Vanderbilt University(范德比大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);prompting(abstract);分类 cs.AI

AI总结 本文提出稀疏双层序数基准MD-3k,用于测量单目深度基础模型的深度层偏好和多层空间关系准确性,发现不同模型对同一分层几何结构有不同解析,且拉普拉斯视觉提示可改变冻结模型的输出层。

Comments 49 pages, 25 figures; Accepted by European Conference on Computer Vision (ECCV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13884 2026-06-15 cs.AI 新提交 83%

Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents

能力最小化作为安全原语:风险感知因果门控实现最小特权LLM代理

Laxmipriya Ganesh Iyer, Rahul Suresh Babu

机构 * Independent Researcher(独立研究者) United States of America(美国)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);分类 cs.AI

AI总结 提出风险感知因果门控(RACG)框架,通过因果效应估计与校准风险控制决定是否采纳模型预测,显著降低高成本错误,同时保持非门控策略的大部分效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11400 2026-06-11 cs.SD cs.AI eess.AS 新提交 83%

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models

引导听哪里:基于指令的激活操控重定向大型音频语言模型中的时间注意力

Tsung-En Lin, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(国立清华大学人工智能研究中心(NTU AI-CoRE))

专题命中 知识编辑与模型理解 :language model(title,abstract);prompting(abstract);分类 cs.AI

AI总结 提出基于指令的向量操控方法,通过对比不同指令下的激活来重定向音频令牌的时间注意力,实现无需训练的声音事件定位,显著优于直接提示和随机基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05557 2026-06-05 cs.CL 83%

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents

AURA: 面向情境化LLM代理中隐式需求挖掘的意图导向探测

Yang Li, Jiaxiang Liu, Jiang Cai, Mingkun Xu

机构 * Guangdong Institute of Intelligence Science and Technology(广东省智能科学与技术研究院)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);分类 cs.CL

AI总结 提出AURA方法,通过在场景感知和工具使用之间插入意图推理步骤生成IntentFrame,以结构化估计隐式需求并控制探测预算,在隐式意图基准上提升覆盖率达+0.07,同时减少82%的探测次数并避免隐私违规。

Comments Submitted to EMNLP 2026. Code, simulator, and benchmark: https://github.com/innovation64/AURA

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01844 2026-06-05 cs.CL 83%

The Cylindrical Representation Hypothesis for Language Model Steering

语言模型引导的圆柱表示假说

Lang Gao, Jinghui Zhang, Wei Liu, Fengxian Ji, Chenxi Wang, Zirui Song, Akash Ghosh, Youssef Mohamed, Preslav Nakov, Xiuying Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 提出圆柱表示假说(CRH),通过放宽线性表示假说(LRH)的正交性假设,解释语言模型引导中的不稳定性和不确定性。

Comments ICML 2026 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03052 2026-06-02 cs.CL 83%

How Language Models Process Negation

语言模型如何处理否定

Zhejian Zhou, Tianyi Zhou, Robin Jia, Jonathan May

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 研究大型语言模型处理否定的内部机制,发现模型内部存在正确处理的组件,但后期注意力层导致错误,通过消融可提升准确率;模型同时采用抑制和构建两种机制,其中构建机制更显著。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22068 2026-06-01 cs.LG 83%

Quantifying the Uncertainty of Foundation Models with Singular Value Ensembles

用奇异值集合量化基础模型的不确定性

Mehmet Ozgur Turkoglu, Dominik J. Mühlematter, Alexander Becker, Konrad Schindler, Helge Aasen

机构 * ETH Zürich, Photogrammetry \& Remote Sensing

专题命中 知识编辑与模型理解 :foundation model(title,abstract);pretraining(abstract);分类 cs.LG

AI总结 提出奇异值集成(SVE)方法,通过冻结奇异向量并仅训练每个成员的奇异值,以极小的参数开销实现隐式集成,从而有效量化基础模型的不确定性。

Comments Accepted at ICML 2026 (camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05483 2026-05-26 cs.AI 83%

MMUEChange: A Generalized LLM Agent Framework for Intelligent Multi-Modal Urban Environment Change Analysis

MMUEChange:面向智能多模态城市环境变化分析的通用LLM智能体框架

Zixuan Xiao, Jun Ma, Siwei Zhang

机构 * Department of Urban Planning and Design, The University of Hong Kong(香港大学城市规划与设计系)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);分类 cs.AI

AI总结 提出MMUEChange多模态智能体框架,通过模块化工具包和模态控制器实现异构城市数据灵活集成与跨模态对齐,在三个城市案例中任务成功率提升46.7%并有效缓解幻觉。

Journal ref Applied Soft Computing 190 (2026) 114576

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23189 2026-05-25 cs.LG 83%

Empirical Bayes Conformal Prediction for Vision and Language Models

视觉与语言模型的经验贝叶斯共形预测

Jiapeng Zeng, Yogesh Prabhu, Zhanpeng Zeng, Michael A. Newton, Vikas Singh

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of California San Diego(加州大学圣地亚哥分校) Xiamen University(厦门大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract_cn);分类 cs.LG

AI总结 提出经验贝叶斯共形预测框架,利用r值将得分变异性转化为不确定性感知的非一致性得分,在保持目标覆盖的同时减少高方差假候选的包含。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02028 2026-05-19 cs.CL 83%

Language models fail at extended rule following

语言模型在扩展规则遵循中表现不佳

Tianxiang Dai, Jonathan Fan

机构 * Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 研究发现语言模型在重复应用规则时无法保持精确状态,即使增加模型规模和计算资源也无法克服这一缺陷,表明需要新的模型架构来实现可靠的规则遵循。

Comments for accessing the data and code for reproduction of the study, see https://txdai.github.io/counting-reliability/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02734 2026-05-18 q-bio.BM cs.AI q-bio.GN 83%

SAE-RNA: A Sparse Autoencoder Model for Interpreting RNA Language Model Representations

SAE-RNA:一种用于解释RNA语言模型表示的稀疏自编码器模型

Taehan Kim, Sangdae Nam

机构 * Department of Computer Science, University of California, Berkeley(加州大学伯克利分校计算机科学系) Department of Development Engineering, University of California, Berkeley(加州大学伯克利分校发展工程系)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 本文提出SAE-RNA模型,通过分析RiNALMo表示并映射到人类水平的生物特征,探索稀疏自编码器在RNA语言模型表示中的可解释性及局限性。

Comments 12 pages, 7 figures. v2: Updated bibliography to improve reference accuracy and reflect updated publication venues. Refined claims for better alignment with results and added an Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11410 2026-05-15 cs.AI 83%

What Do EEG Foundation Models Capture from Human Brain Signals?

EEG基础模型从人类大脑信号中捕捉了什么?

Ling Tang, Qian Chen, Jilin Mei, Houshi Xu, Quanshi Zhang, Jing Shao, Na Zou, Xia Hu, Dongrui Liu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) Tongji University(同济大学) University of Houston(休斯顿大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);pretraining(abstract);分类 cs.AI

AI总结 研究探讨EEG基础模型如何从原始信号中学习,并通过实验验证其表示与特征工程基线的对齐情况,揭示了模型学习、使用及解释性方面的关键发现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10415 2026-05-14 cs.CL 83%

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis

对主观分析中语言模型不确定性与人类分歧的对齐

Junyu Lu, Deyi Ji, Xuanyi Liu, Lanyun Zhu, Bo Xu, Liang Yang, Xian-Sheng Hua, Hongfei Lin

机构 * School of Computer Science and Technology, Dalian University of Technology, China(大连理工大学计算机科学与技术学院) Tencent(腾讯) Peking University(北京大学) Tongji University(同济大学)

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出DPUA框架,通过感知分歧和对齐不确定性,提升模型在主观分析中的可靠性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10337 2026-05-12 cs.AI eess.SP 83%

CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings

CORTEG:基础模型促进从头皮到脑内记录的跨模态表示迁移

Liuyin Yang, Qiang Sun, Bob Van Dyck, Eva Calvo Merino, Marc M. Van Hulle

机构 * Laboratory for Neuro- & Psychophysiology, Department of Neurosciences, KU Leuven(神经与心理生理实验室,神经科学系,比利时鲁文大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);pretraining(abstract);分类 cs.AI

AI总结 CORTEG通过预训练的头皮EEG基础模型实现跨模态表示迁移,提升脑机接口的跨患者学习和解码性能,同时在单GPU上10-30分钟内校准新患者。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00323 2026-05-04 cs.CV cs.LG 83%

Online Self-Calibration Against Hallucination in Vision-Language Models

在线对抗幻觉的视觉-语言模型自我校准

Minghui Chen, Chenxu Yang, Hengjie Zhu, Dayan Wu, Zheng Lin, Qingyi Si

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);preference optimization(abstract);分类 cs.LG

AI总结 本文提出OSCAR框架,通过蒙特卡洛树搜索和双粒度奖励机制在线校准视觉-语言模型,减少幻觉并提升多模态能力。

Comments IJCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27209 2026-05-04 cs.SE cs.AI 83%

Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves

构建中的理论:协调语言模型用于研究软件,其中规范在演变

Halley Young, Nikolaj Björner

机构 * Microsoft Research(微软研究院)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 本文提出Comet-H框架,通过迭代提示自动机协调语言模型在研究软件中的应用,解决规范演变中的 hallucination accumulation 和 desynchronization 问题,并通过实验证明其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17469 2026-05-01 cs.CL cs.HC 83%

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability

跨语言情感不一致:审计多语言语言模型以检测反向风险、方言表示和情感稳定性

Nusrat Jahan Lia, Shubhashis Roy Dipta

机构 * Institute of Information Technology University of Dhaka(达卡大学信息科技学院) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 知识编辑与模型理解 :language model(title);LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本文通过控制基准框架评估多语言Transformer模型在平行孟加拉语-英语句子对上的表现,揭示了模型在情感反向、同理心不对称和方言表示中的问题,提出需引入情感稳定性指标以提升跨语言可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23079 2026-04-28 cs.CV cs.AI 83%

From Pixels to Explanations: Interpretable Diabetic Retinopathy Grading with CNN-Transformer Ensembles, Visual Explainability and Vision-Language Models

从像素到解释:结合CNN-Transformer集成、视觉可解释性和视觉语言模型的可解释性糖尿病视网膜病变分级

Pir Bakhsh Khokhar, Carmine Gravino, Fabio Palomba, Sule Yildirim Yayilgan, Sarang Shaikh

机构 * Department of Informatics, University of Salerno(萨勒诺大学信息学院) Department of Information Security and Communication Technology (IIK), Norwegian University of Science and Technology(挪威科技大学信息安全部)

专题命中 知识编辑与模型理解 :language model(title,abstract);prompting(abstract);分类 cs.AI

AI总结 本文提出结合强判别模型与多模态解释的方法,利用APTOS 2019基准评估六种CNN和Transformer模型,通过集成策略提升分级一致性,并利用视觉语言模型生成可解释的视网膜病变分级结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19052 2026-04-22 cs.CL 83%

Cell-Based Representation of Relational Binding in Language Models

基于细胞的语言模型中关系绑定的表示

Qin Dai, Benjamin Heinzerling, Kentaro Inui

机构 * Tohoku University(东大大学) RIKEN(日本科学技術研究所) AIP(Advanced Institute for Prognostic Studies) MBZUAI(马克斯·普朗克研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 研究揭示语言模型通过细胞绑定表示进行关系绑定,通过多句子数据验证了细胞子空间的可解码性及跨上下文迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05437 2026-04-22 cs.CL 83%

Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models

一旦正确,仍错误:多语言视觉-语言模型中的反事实幻觉

Basel Mousi, Fahim Dalvi, Shammur Chowdhury, Firoj Alam, Nadir Durrani

机构 * Qatar Computing Research Institute, HBKU(卡塔尔计算研究所,哈姆丹·本·哈马德大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);prompting(abstract);分类 cs.CL

AI总结 研究探讨多语言视觉-语言模型在文化背景下反事实幻觉的问题,提出M²CQA基准测试,通过17个中东国家图像和多语言对比陈述评估模型性能,发现阿拉伯语尤其在方言中反事实幻觉率显著上升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13114 2026-04-16 cs.SE cs.AI 83%

The Code Whisperer: LLM and Graph-Based AI for Smell and Vulnerability Resolution

代码低语者:结合图神经网络和大语言模型的代码和漏洞解决方法

Mohammad Baqar, Raji Rustamov, Alexander Hughes

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出Code Whisperer框架,结合图分析与大语言模型,统一检测和修复代码维护性和安全性问题,提升检测性能和修复建议实用性。

Comments 10 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10151 2026-04-15 cs.CL 83%

Nationality encoding in language model hidden states: Probing culturally differentiated representations in persona-conditioned academic text

语言模型隐藏状态中的国籍编码:探究基于人格条件的学术文本中的文化差异表示

Paul Jackson, Ruizhe Li, Elspeth Edelstein

机构 * Language Centre, School of Language, Literature, Music and Visual Culture, University of Aberdeen(语言中心,语言文学音乐与视觉文化学院,阿伯丁大学) School of Natural and Computing Sciences, University of Aberdeen(自然与计算科学学院,阿伯丁大学) School of Language, Literature, Music and Visual Culture, University of Aberdeen(语言文学音乐与视觉文化学院,阿伯丁大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 研究测试Gemma-3-4b-it在生成英文学术文本时是否在隐藏状态中编码文化差异的国籍信息,通过分析隐藏状态激活并发现国籍编码在不同层中呈现非单调轨迹。

Comments 42 pages, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09364 2026-04-14 cs.CV cs.CL 83%

Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts

仲裁失败,而非感知失明:视觉-语言模型如何解决视觉-语言冲突

Farhad Nooralahzadeh, Omid Rohanian, Yi Zhang, Jonathan Fürst, Kurt Stockinger

机构 * Institute of Computer Science, Zurich University of Applied Sciences(苏黎世应用科技大学计算机科学研究所) University of Oxford(牛津大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 研究探讨视觉-语言模型在视觉与语言冲突中的仲裁机制,发现编码与基础的脱节,通过多模态仲裁交叉分析揭示视觉属性在早期层可线性解码,最终层logit与基础结果相关性达0.847,表明需针对性干预提升视觉基础能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18893 2026-04-14 cs.AI 83%

Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation

语言模型中的定量内省:跨对话跟踪情感状态

Nicolas Martorell, Bruno Bianchi

机构 * Universidad de Buenos Aires. Facultad de Ciencias Exactas y Naturales. Departamento de Computación(布宜诺斯艾利斯大学。精确与自然科学学院。计算机系) CONICET-Universidad de Buenos Aires. Instituto de Ciencias de la Computación (ICC)(国家科学与技术研究理事会-布宜诺斯艾利斯大学。计算机科学研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 本文研究语言模型通过数字自报告跟踪情感状态的方法,发现通过计算logit基自报告可揭示内省能力,该指标能有效追踪内部状态变化,并在不同模型规模下表现出良好的一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08885 2026-04-13 cs.LG 83%

Uncertainty-Aware Transformers: Conformal Prediction for Language Models

具有不确定性的转换器:语言模型的置信预测

Abhiram Vellore, Niraj K. Jha

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.LG

AI总结 本文提出CONFIDE框架,通过置信预测提升语言模型的可解释性和准确性,在BERT-tiny上测试准确率提升4.09%,并在资源受限和高风险任务中表现出鲁棒性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19297 2026-04-08 cs.LG 83%

CLaRE-ty Amid Chaos: Quantifying Representational Entanglement to Predict Ripple Effects in LLM Editing

在混乱中保持清晰:量化表征纠缠以预测LLM编辑中的涟漪效应

Manit Baser, Alperen Yildiz, Dinil Mon Divakaran, Mohan Gurusamy

机构 * National University of Singapore(新加坡国立大学) A*STAR Institute for Infocomm Research (A*STAR I2 R)(新加坡科技研究局资讯通信研究院)

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出CLaRE技术,通过量化事实间的纠缠关系,识别LLM编辑中的潜在涟漪效应,提升模型编辑的稳定性和效率。

Comments Accepted to ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19959 2026-04-08 cs.AR cs.AI 83%

From Concept to Practice: an Automated LLM-aided UVM Machine for RTL Verification

从概念到实践:一种自动化的LLM辅助UVM机器用于RTL验证

Junhao Ye, Yuchen Hu, Ke Xu, Dingrong Pan, Qichun Chen, Jie Zhou, Shuai Zhao, Xinwei Fang, Xi Wang, Nan Guan, Zhe Jiang

机构 * School of Integrated Circuits, Southeast University(东南大学集成电路学院) National Center of Technology Innovation for EDA(国家EDA技术创新中心) College of Economics, Shenzhen University(深圳大学经济学院) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Department of Computer Science, University of York(约克大学计算机科学系) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出UVM^2框架,利用大语言模型自动生成并优化UVM测试平台,显著降低手动工作量,提升验证效率与覆盖率。

Comments Accepted by the IEEE/ACM International Conference on Computer-Aided Design (ICCAD) 2025. This version includes the camera-ready manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19315 2026-04-06 cs.CL 83%

AutoPCR: Automated Phenotype Concept Recognition by Prompting

AutoPCR: 通过提示实现自动表型概念识别

Yicheng Tao, Yuanhao Huang, Yiqun Wang, Xin Luo, Jie Liu

机构 * University of Michigan(密歇根大学)

专题命中 知识编辑与模型理解 :prompting(title);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出AutoPCR,一种基于提示的表型概念识别方法,无需领域特定训练即可自动泛化至新本体和未见数据,实验显示其在多个数据集上表现最佳。

Comments Accepted at ISMB 2026 (Proceedings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01375 2026-04-03 cs.AI 83%

Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges

对齐以错位:基于元优化LLM评判者的自动LLM劫持

Hamin Koo, Minseon Kim, Jaehyung Kim

机构 * Yonsei University(延世大学) Microsoft Research(微软研究院)

专题命中 知识编辑与模型理解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出AMIS框架,通过双层结构联合优化劫持提示和评分模板,提升LLM劫持效果,实测在AdvBench和JBB-Behaviors上取得最佳性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00228 2026-04-02 cs.CL 83%

Do Language Models Know When They'll Refuse? Probing Introspective Awareness of Safety Boundaries

语言模型是否知道何时会拒绝?探测安全边界内的自我意识

Tanay Gondil

机构 * Purdue University(普渡大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 研究探讨语言模型能否准确预测自身拒绝行为,通过系统分析发现模型在安全边界处敏感度下降,且通过置信度评分可提升安全部署的准确性。

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏