arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7539 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7539 篇

2304.02202 2023-04-06 cs.CV cs.HC cs.LG 83%

Towards Self-Explainability of Deep Neural Networks with Heatmap Captioning and Large-Language Models

Osman Tursun, Simon Denman, Sridha Sridharan, Clinton Fookes

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01736 2022-07-06 cs.CL 83%

Probing via Prompting

Jiaoda Li, Ryan Cotterell, Mrinmaya Sachan

专题命中 知识编辑与模型理解 :prompting(title,abstract);language model(abstract);分类 cs.CL

Comments NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01113 2021-10-05 cs.CL 83%

Probing Language Models for Understanding of Temporal Expressions

Shivin Thukral, Kunal Kukreja, Christian Kavouras

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

Comments Accepted for publication in BlackboxNLP (EMNLP 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.14388 2021-09-13 cs.CL 83%

Universal Sentence Representation Learning with Conditional Masked Language Model

Ziyi Yang, Yinfei Yang, Daniel Cer, Jax Law, Eric Darve

专题命中 知识编辑与模型理解 :language model(title,abstract);post-training(abstract);分类 cs.CL

Comments Accepted to the 2021 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04718 2026-08-13 cs.LG cs.AI cs.CL 版本更新 83%

Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

通过语言模型中几乎正交的特征实现孤立干预

Moritz Miller, Florent Draye, Bernhard Schölkopf

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究基于机械可解释性前提,针对语言模型特征纠缠问题,受“独立因果机制”启发,提出将内部特征约束为几乎正交,通过形式化问题、界定干扰传播并关联正则化,实现对数学推理概念更孤立的干预且保留模型性能。

Comments Published as a conference paper at the Conference on Language Modeling (COLM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21075 2026-06-23 cs.CL cs.AI cs.LG 新提交 83%

FiLM-Coordinated Dual-Branch Transformer for Global-Local Dependency Modeling in Language Modeling

FiLM协调的双分支Transformer用于语言建模中的全局-局部依赖建模

Zhiqiang Zhou, Xu Ling, Junliang Dai

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出FiLM协调的双分支Transformer,通过特征级线性调制动态协调全局与局部分支,在轻量级设置下优于单分支基线。

Comments 14 pages, 7 figures, 7 tables. Small-scale language modeling study on FiLM-coordinated dual-branch Transformer architectures, including multi-seed evaluation, cross-dataset validation, ablation studies, efficiency analysis, and parameter-matched fairness baselines

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07564 2025-12-11 cs.CV cs.AI cs.CL cs.LG 83%

Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models

迈向更可靠的人工智能:减少视觉-语言模型中的幻觉

Kassoum Sanogo, Renzo Ardiccioni

机构 * Department of CS AI and Data Science(计算机科学与数据科学系) ESEO Engineering School(ESEO工程学院) Faculty of Law, Economy, Management(法学院、经济与管理学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出无需训练的自我纠正框架,通过不确定性引导的视觉再注意力减少视觉-语言模型中的幻觉,提升响应准确性。

Comments 24 pages, 3 figures, 2 tables. Training-free self-correction framework for vision-language models. Code and implementation details will be released at: https://github.com/kassoumsanogo1/self-correcting-vlm-re-Attention.git

Journal ref The 4th National and International Academic Conference Celebrating the 20th Anniversary of Rajapruk University (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06073 2025-11-11 cs.CL cs.AI cs.LG cs.LO 83%

Stemming Hallucination in Language Models Using a Licensing Oracle

Simeon Emanuilov, Richard Ackermann

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 23 pages, 4 figures, 8 tables. Introduces the Licensing Oracle, an architectural solution for eliminating hallucinations in language models through formal SHACL validation against knowledge graphs. All datasets and models are available at https://huggingface.co/collections/s-emanuilov/licensing-oracle-experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16970 2026-08-19 cs.CR cs.AI cs.LG 新提交 82%

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

探测预填充:通过潜在激活检测代码漏洞

Alizishaan Khatri

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.AI、cs.LG

AI总结 该研究从四个LLM提取预填充标记激活训练MLP探测器,在四个代码漏洞基准上实现平均F1值41.7%,证明LLM自身代码表示含漏洞信息,为轻量原生筛查提供方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13328 2026-08-14 cs.CL cs.AI 新提交 82%

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

提问方式的重要性:大语言模型(LLMs)中与性别相关的语言偏差

Katherine Van Koevering, Anjalie Field

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.CL、cs.AI

AI总结 该研究发现LLMs会因提示中女性常用的语言特征生成更短欠正式的回应,其影响远大于明确性别线索,且事后缓解困难,呼吁关注语言变异以缓解LLM介导职场交流的差异化影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08024 2026-08-11 cs.CL cs.AI 新提交 82%

Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States

提示嵌入探针(PEP):基于隐藏状态的大语言模型幻觉检测

Zakhar Mrykhin, Valentin Malykh

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出提示嵌入探针(PEP),一种白盒幻觉检测方法,通过少量可学习提示嵌入扩充标准线性探针,在TriviaQA等数据集的Qwen3模型上验证其在生成前、跨模型等场景的有效性,仅需少量参数即可提升检测性能。

Comments 10 pages, 7 figures. Code available at https://github.com/zazamrykh/internal_probing

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04286 2026-08-06 cs.CL cs.LG 新提交 82%

Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks

通过语义等价对抗攻击引出大语言模型的内在幻觉

Atri Vivek Sharma, Brian Formento, Alessio Lomuscio

机构 * Imperial College London(帝国理工学院) Safe Intelligence(安全智能研究院)

专题命中 知识编辑与模型理解 :LLM(summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出新框架,通过语义等价对抗攻击测试LLMs鲁棒性,发现先进模型易受语义保留扰动影响致忠实度大幅下降,凸显需开发鲁棒接地的LLM架构与训练目标。

Comments To be presented at COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01046 2026-08-04 cs.CL cs.AI 新提交 82%

DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text

DeBERTa-Sentinel:实现透明且可信的AI生成文本检测

Muhammad Yousaf Rehman, Muhammad Islam

机构 * University of Hertfordshire(赫特福德大学) James Cook University(詹姆斯库克大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 DeBERTa-Sentinel是基于DeBERTa-v3的透明AI生成文本检测框架,在GLC-AIText数据集上实现优异检测性能,可公开token级解释以支持利益相关者审计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22646 2026-07-28 cs.AI cs.LG 新提交 82%

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models

从预训练语言模型中提取算法:以隐马尔可夫模型为例

Yijia Dai, Zhaolin Gao, Yahya Sattar, Jennifer J. Sun, Sarah Dean

机构 * Cornell University(康奈尔大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究预训练语言模型从隐马尔可夫模型预测下一个观测值的能力背后算法,通过三阶段流程,先实证比较缩小候选算法范围,再推导理论联系并验证,最后引入主激活探针揭示低维线性表示及变化,将其上下文行为与内部机制联系起来。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21443 2026-07-20 cs.CL cs.AI 82%

Domain Knowledge-Enhanced LLMs for Fraud and Concept Drift Detection

领域知识增强的大语言模型用于欺诈和概念漂移检测

Ali Şenol, Garima Agrawal, Huan Liu

机构 * School of Computing and Augmented Intelligence (SCAI), Arizona State University (ASU)(计算与增强智能学院(SCAI),亚利桑那州立大学) Department of Computer Engineering, Tarsus University(计算机工程系,塔鲁斯大学) Minerva CQ and HumaConn AI Consulting(Minerva CQ和HumaConn人工智能咨询) School of Computing and Augmented Intelligence (SCAI), Arizona State University(计算与增强智能学院(SCAI),亚利桑那州立大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出一种领域知识增强的大语言模型框架,通过集成结构化领域知识和漂移检测单元,实现高准确率的欺诈对话检测和概念漂移分类。

Journal ref Electronics 2026, 15(3), 534

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01901 2026-07-03 cs.LG cs.AI cs.MM 新提交 82%

SABER: A Semantic-Aligned Brain Network Analysis Framework via Multi-scale Hypergraphs

SABER: 一种基于多尺度超图的语义对齐脑网络分析框架

Yidan Xu, Xiangmin Han, Rundong Xue, Huihui Ye

机构 * Hangzhou Dianzi University(杭州电子科技大学) Tsinghua University(清华大学) Xi’an Jiaotong University(西安交通大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出SABER框架,通过多尺度超图建模功能子网络与高阶依赖,并引入决策级语义对齐机制,将大语言模型语义直接融入预测,在ABIDE和ADHD-200数据集上取得最优性能,尤其在小样本场景下。

Comments Accepted to IEEE International Conference on Multimedia and Expo (ICME) 2026;

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30578 2026-06-30 cs.CL cs.LG 82%

Uncertainty-Aware Generation and Decision-Making Under Ambiguity

模糊性下的不确定性感知生成与决策

Nico Daheim, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt(普遍知识处理实验室(UKP实验室),计算机科学系,达姆施塔特技术大学) National Research Center for Applied Cybersecurity ATHENE, Germany(应用网络安全国家研究中心ATHENE,德国)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 基于贝叶斯决策理论和风险规避决策,提出不确定性感知算法用于辅导和同行评审任务,通过共形预测提供策略和分数的保证,实验表明贝叶斯方法优于风险规避规则。

Comments Code available under https://github.com/UKPLab/arXiv2026-uncertainty-aware

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23885 2026-06-24 cs.CV cs.AI cs.CL cs.MM 新提交 82%

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

注意头:多模态大语言模型的拓扑表示对齐

Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi

机构 * University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学) University of Pisa(比萨大学) AMD Silo AI

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出头级表示对齐(HeRA)方法,在注意力头级别强制跨模态对齐,通过对比目标匹配局部拓扑结构,选择对齐最差的头进行训练,有效提升视觉任务性能并减少幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14080 2026-06-23 cs.CL cs.AI 版本更新 82%

Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality

空货架还是丢钥匙?回忆是参数化事实性的瓶颈

Nitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek, Gal Yona

机构 * Google Research(Google研究) Technion – Israel Institute of Technology(以色列技术学院)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.CL、cs.AI

AI总结 提出行为框架区分LLM事实错误源于知识缺失或访问失败,通过WikiProfile基准测试发现前沿模型编码饱和但回忆是主要瓶颈,思考可恢复大量失败。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06333 2026-06-05 cs.LG cs.AI 82%

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability

子空间感知稀疏自编码器用于有效的机制可解释性

Seyed Arshan Dalili, Mehrdad Mahdavi

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 针对稀疏自编码器将特征假设为一维导致特征分裂的问题,提出子空间感知稀疏自编码器(SASA),通过学习解码器子空间、块稀疏门控和核范数正则化,在GPT-2和Mistral-7B上减少特征分裂和吸收,提高单义性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26366 2026-06-03 cs.AI cs.LG 82%

Automatic Layer Selection for Hallucination Detection

幻觉检测的自动层选择

Xinpeng Wang, William X. Cao, Andrew Gordon Wilson, Zhe Zeng

机构 * University of Washington(华盛顿大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 针对大语言模型中幻觉检测的层选择问题,提出无需训练的FEPoID准则,自动识别最优中间层,并结合截断策略提升检测性能。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02237 2026-06-02 cs.LG cs.AI 82%

Concept Heterogeneity-aware Representation Steering

概念异质性感知表示引导

Laziz U. Abdullaev, Noelle Y. L. Wong, Ryan T. Z. Lee, Shiqi Jiang, Khoi N. M. Nguyen, Tan M. Nguyen

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 针对大语言模型表示非均匀导致全局引导脆弱的问题,提出基于最优传输的输入依赖引导方法CHaRS,通过高斯混合模型和离散最优传输实现更有效的行为控制。

Journal ref ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09730 2026-06-02 cs.CL cs.LG 82%

Interpreto: An Explainability Library for Transformers

Interpreto:一个用于Transformer的可解释性库

Antonin Poché, Thomas Mullor, Gabriele Sarti, Frédéric Boisnard, Corentin Friedrich, Charlotte Claye, François Hoofd, Raphael Bernas, Nicholas Asher, Céline Hudelot, Fanny Jourdan

机构 * IRT Saint Exupéry Toulouse(伊尔杜夫圣埃克苏佩里图卢斯) IRIT Toulouse(图卢兹IRIT) Khoury College of Computer Sciences(科赫里计算机科学学院) Ampere(阿姆佩尔) MICS, CentraleSupélec(MICS,中央超导学院) Scienta Lab(科学实验室) Thales Avionics(泰勒斯航空电子) ANITI

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract_cn);language model(abstract);分类 cs.CL、cs.LG

AI总结 Interpreto是一个开源Python库,通过归因方法和基于概念的解释,为HuggingFace语言模型(从早期BERT变体到LLM)提供统一的解释工作流,其端到端基于概念的流水线是主要创新。

Comments Accepted to ACL 2026 System Demonstration. Equal contribution: Poché and Jourdan

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30790 2026-06-01 cs.IR cs.AI cs.CL 82%

On the impact of retrieved content representations in RAG Pipelines

关于检索内容表示对RAG管道的影响

Jonathan J Ross, Bevan Koopman, Anton van der Vegt, Guido Zuccon

机构 * The University of Queensland(昆士兰大学) CSIRO(澳大利亚联邦科学与工业研究组织)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 通过控制变量实验,研究检索文档的不同表示(选择、摘要、改写等)对RAG生成准确性的影响,发现答案保留是主要决定因素。

Comments 23 pages, 15 figures, submitted to ACL May 2026 ARR

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27402 2026-05-28 cs.CY cs.AI cs.CL 82%

REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading

REC-CBM:面向可信开放评分的基于规则感知的错误修正概念瓶颈模型

Chengshuai Zhao, Fan Zhang, Kumar Satvik Chaudhary, Yiwen Li, Lo Pang-Yun Ting, Ying-Chih Chen, Huan Liu

机构 * School of Computing and Augmented Intelligence, Arizona State University, USA(计算与增强智能学院,亚利桑那州立大学,美国) Mary Lou Fulton Teachers College, Arizona State University, USA(玛丽·卢·福洛顿教师学院,亚利桑那州立大学,美国) Department of Computer Science, National Yang Ming Chiao Tung University, TW(国立阳明交通大学计算机科学系,台湾)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出REC-CBM模型,通过规则感知概念编码器、序数成对校准目标和潜在概念错误修正模块,解决开放评分中标准概念瓶颈模型无法建模细粒度规则维度、忽略评分序数语义和概念标注不可靠的问题,在提升评分性能的同时保持可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26808 2026-05-27 cs.LG cs.AI cs.IT math.IT 82%

Innovation: An Almost Characterization of Hallucination

创新:幻觉的几乎刻画

Nishant P. Das, Piyush Srivastava

机构 * School of Technology and Computer Science, Tata Institute of Fundamental Research, Mumbai, Maharashtra - 400 005, India(技术与计算机科学学院,塔塔基础研究机构,孟买,马哈拉施特拉邦 - 400 005, 印度)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文引入“创新”属性来刻画大语言模型幻觉的必然性,证明创新与幻觉几乎等价,并基于创新率给出新的幻觉率下界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20242 2026-05-21 cs.LG cond-mat.mtrl-sci cs.AI physics.chem-ph 82%

LEAP: A closed-loop framework for perovskite precursor additive discovery

LEAP:一种用于钙钛矿前驱体添加剂发现的闭环框架

Xin-De Wang, Zhi-Rui Chen, Ze-Feng Gao, Peng-Jie Guo, Cheng Mu, Zhong-Yi Lu

机构 * School of Physics, Renmin University of China(中国人民大学物理学院) School of Chemistry and Life Resource, Renmin University of China(中国人民大学化学与生命资源学院)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 该研究提出LEAP框架,结合大语言模型和主动学习,通过文献驱动的机制相关描述符和贝叶斯优化,实现了钙钛矿太阳能电池添加剂的高效发现,实验验证显示其在性能提升方面优于通用模型。

Comments 30 pages; 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20084 2026-05-20 cs.CL cs.AI 82%

BalanceRAG: Joint Risk Calibration for Cascaded Retrieval-Augmented Generation

BalanceRAG: 为级联检索增强生成进行联合风险校准

Zijun Jia, Yuanchang Ye, Sen Jia, Yiyao Qian, Haoning Wang, Baojie Chen, Diyin Tang, Jinsong Yu, Zhiyuan Wang

机构 * Beihang University(北航) Shenzhen Institute of Advanced Technology(深圳先进技术研究院) Zhejiang University of Finance & Economics(浙江财经大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出BalanceRAG,一种用于级联检索增强生成的联合风险校准方法,通过在二维晶格上确定安全操作点,实现风险自适应的阈值校准,从而在控制系统级错误率的同时保留更多示例,并扩展到多风险校准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19735 2026-05-20 cs.CL cs.AI 82%

ContextRAG: Extraction-Free Hierarchical Graph Construction for Retrieval-Augmented Generation

ContextRAG: 无提取的分层图构建用于检索增强生成

Roman Prosvirnin, Sergei Kuznetsov, Seungmin Jin

机构 * HSE University(俄罗斯高等经济大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出ContextRAG,一种无需大型语言模型提取实体和关系的图检索增强生成系统,通过残差量化k均值和Formal Concept Analysis方法构建模糊概念图,在130个任务的UltraDomain子集中实现了33.6%的F1分数,显著优于传统方法。

Comments Preprint. 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21702 2026-05-18 cs.LG cs.CL 82%

Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities

超越遗忘:机器去学习引发可控的副作用和能力

Tien Dang, The-Hai Nguyen, Dinh Mai Phuong, Nguyen Minh Phuong, Anh Bui, Hoang Thanh-Tung, Le-Minh Nguyen, Naoya Inoue

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究所) Quantum AI and Cyber Security Institute(量子人工智能与网络安全研究所) FPT Corporation(FPT公司) Monash University(墨尔本大学) RIKEN(理化学研究所)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文探讨了通过去学习引发可控副作用和能力的现象,通过线性表示假设验证了该现象在行为控制和能力增强任务中的有效性。

Comments 36 pages, 19 tables, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏