arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-05-19 至 2026-05-19 共收录 717 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 61 篇

2601.18664 2026-05-19 cs.IR 67%

S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation

S$^2$GR: 潜在空间中基于分步语义引导的生成推荐

Zihao Guo, Jian Wang, Ruxin Zhou, Youhua Liu, Jiawei Guo, Jun Zhao, Xiaoxiao Xu, Yongqi Liu, Kaiqiao Zhan

专题命中 领域大模型 :large language model(abstract);language model(abstract)

AI总结 本文提出S$^2$GR框架,通过在生成SID代码前插入思考令牌,建立稳健的语义基础,提升生成推荐的推理能力与性能。

Comments Accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16775 2026-05-19 cs.CV cs.AI cs.LG 62%

VolTA-3D: Self-Supervised Learning for Brain MRI using 3D Volumetric Token Alignment

VolTA-3D: 基于3D体积分块对齐的脑MRI自监督学习

Amy Makawana, Abhijeet Parida, Marius George Linguraru, Julia Ive, Syed Muhammad Anwar

机构 * Institute of Health Informatics(健康信息学研究所) Sheikh Zayed Institute for Pediatric Surgical Innovation(谢赫扎耶德儿童外科创新研究所) School of Medicine and Health Sciences(医学与健康科学学院)

专题命中 领域大模型 :pretraining(abstract);分类 cs.AI、cs.LG

AI总结 本文提出VolTA-3D,一种用于脑MRI自监督学习的3D视觉Transformer框架,通过联合对齐全局类风格标记和局部块标记,增强体积分块表示的可迁移性,从而在多个下游任务中表现出更好的泛化能力和鲁棒性。

Comments Accepted at EMBC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16351 2026-05-19 cs.LG cs.AI 62%

PIMSM: Physics-Informed Multi-Scale Mamba for Stable Neural Representations under Distribution Shift

PIMSM:基于物理的多尺度Mamba用于在分布偏移下稳定的神经表示

Sangyoon Bae, Shinjae Yoo, Jiook Cha

机构 * Interdisciplinary Program in Artificial Intelligence(人工智能交叉学科项目) Seoul National University(首尔国立大学) Computational Science Initiative(计算科学倡议) Brookhaven National Laboratory(布鲁赫斯国家实验室) Department of Psychology(心理学系)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出PIMSM,一种基于物理的多尺度Mamba架构,通过时间尺度对齐提升科学基础模型在分布偏移下的鲁棒性和表示稳定性,实验证明其在fMRI和气象预测中的有效性。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18547 2026-05-19 cs.AI 57%

VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

VISAFF: 以说话者为中心的视觉情感特征学习用于对话中的情感识别

Linan ZHU, Zihao Zhai, Xiao Han, Yuqian Fu, Xiangfan Chen, Xiangjie Kong, Guojiang Shen

机构 * Zhejiang University of Technology(浙江工业大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 本文提出VISAFF框架,通过以说话者为中心的视觉情感特征学习方法,解决对话中情感识别中的复杂场景问题,提升计算效率并避免大规模模型微调的高成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18007 2026-05-19 cs.CL 57%

Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labeling

推理时对修辞角色标注中困难示例的语义重排序

Anas Belfathi, Nicolas Hernandez, Laura Monceaux, Warren Bonnard, Richard Dufour

机构 * Nantes Université, École Centrale Nantes, CNRS, LS2N, UMR 6004(南特大学,中央理工学院,国家科学研究中心,LS2N,UMR 6004) University of Lorraine(洛林大学)

专题命中 领域大模型 :language model(abstract);分类 cs.CL

AI总结 本文提出RISE框架,在推理时利用标签语义对修辞角色标注中的困难示例进行重排序,提升模型预测的准确性和鲁棒性。

Comments Accepted at ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17812 2026-05-19 cs.AI 57%

Going Headless? On the Boundaries of Vertical AI Firms

going headless?关于垂直AI企业的边界

Muhammad Zia Hydari, Farooq Muzaffar

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 领域大模型 :prompting(abstract);分类 cs.AI

AI总结 本文探讨了垂直AI企业在会计、法律、医疗、采购等领域中,将工作流、领域逻辑和责任整合到单一应用中的传统模式,以及通用AI代理如何解构这种模式,促使企业采取"going headless"策略。文章指出,这种策略对某些企业有益,对另一些企业则可能造成破坏,并提出了基于任务-责任制度的三类分类体系及规则债务的概念。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17071 2026-05-19 cs.AI 57%

AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation

AnchorDiff: 基于拓扑结构的掩码扩散模型与基于置信度的重写方法用于放射学报告生成

Shiying Yu, Jielei Wang, Guoming Lu

机构 * University of Electronic Science and Technology of China(电子科技大学)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 本文提出AnchorDiff,一种首个结合临床锚点的掩码扩散框架,用于生成放射学报告。该方法通过拓扑感知训练策略和推理时的重写策略,有效缓解了固定顺序自回归解码的局限性,实现了最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16973 2026-05-19 cs.CV cs.LG 57%

SHED: Style-Homogenized Embedding Alignment for Domain Generalization

SHED: 风格均质化嵌入对齐用于领域泛化

Kai Gan, Tong Wei

机构 * School of Computer Science and Engineering, Southeast University, Nanjing 210096, China(1 东南大学计算机科学与工程学院,南京 210096,中国) Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China(2 教育部计算机网络与信息集成重点实验室(东南大学),中国)

专题命中 领域大模型 :language model(abstract);分类 cs.LG

AI总结 本文提出SHED方法,通过均质化嵌入对齐来解决领域泛化中的信息不对称问题,实验表明其在多个基准测试中取得了最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16966 2026-05-19 cs.AI 57%

Harnessing AI for Inverse Partial Differential Equation Problems: Past, Present, and Prospects

利用人工智能解决逆偏微分方程问题:过去、现在与展望

Zhentao Tan, Yuze Hao, Boyi Zou, Mingsheng Long, Yi Yang, Gang Bao

机构 * Collaborative Innovation Center of Artificial Intelligence (CCAI), Zhejiang University(人工智能协同创新中心(CCAI),浙江大学) School of Mathematical Sciences, Zhejiang University(浙江大学数学科学学院) Tsinghua University(清华大学) Center for Interdisciplinary Applied Mathematics, School of Mathematical Sciences, Zhejiang University(浙江大学数学科学学院交叉应用数学中心)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 本文综述了利用人工智能解决逆偏微分方程问题的最新进展,涵盖了逆问题、逆设计和控制问题三大类,总结了科学和工业领域中的典型应用,并讨论了开放挑战和未来前景。

Comments 35 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13270 2026-05-19 cs.CV cs.AI 57%

RadGame: An AI-Powered Platform for Radiology Education

RadGame:一种基于人工智能的放射学教育平台

Mohammed Baharoon, Siavash Raissi, John S. Jun, Thibault Heintz, Mahmoud Alabbad, Ali Alburkani, Sung Eun Kim, Kent Kleinschmidt, Abdulrahman O. Alhumaydhi, Mohannad Mohammed G. Alghamdi, Jeremy Francis Palacio, Mohammed Bukhaytan, Noah Michael Prudlo, Rithvik Akula, Brady Chrisler, Benjamin Galligos, Mohammed O. Almutairi, Mazeen Mohammed Alanazi, Nasser M. Alrashdi, Joel Jihwan Hwang, Sri Sai Dinesh Jaliparthi, Luke David Nelson, Nathaniel Nguyen, Sathvik Suryadevara, Steven Kim, Mohammed F. Mohammed, Yevgeniy R. Semenov, Kun-Hsing Yu, Abdulrhman Aljouie, Hassan AlOmaish, Adam Rodman, Pranav Rajpurkar

机构 * Harvard Medical School(哈佛医学院) Mass General Brigham(麻省总医院) Maastricht University(马斯特里赫特大学) Department of Medical Imaging, King Abdulaziz Medical City, Ministry of National Guard, Riyadh, Saudi Arabia(国王阿卜杜勒-阿齐兹医疗城医学影像科,沙特阿拉伯) National Strategic Technology Research Institute, Seoul National University Hospital(全国战略技术研究所,首尔国立大学医院) Saint Louis University School of Medicine(圣路易斯大学医学院) College of Medicine, King Saud bin Abdulaziz University for Health Sciences(国王萨勒曼·本·阿卜杜勒阿齐兹大学医学院) Tufts University School of Medicine(塔夫茨大学医学院) Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 RadGame通过结合游戏化与大规模公开数据集,提供AI驱动的反馈,提升放射学教育中的定位和报告撰写能力,显著提高学习效果。

Comments ML4H Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06933 2026-05-19 cs.LG physics.comp-ph physics.flu-dyn 57%

Neural equilibria for long-term prediction of nonlinear conservation laws

神经均衡用于非线性守恒律的长期预测

J. Antonio Lara Benitez, Kareem Hegazy, Junyi Guo, Ivan Dokmanić, Michael W. Mahoney, Maarten V. de Hoop

机构 * Rice University(里士大学) ICSI and University of California at Berkeley(ICSI和加州大学伯克利分校) University of Basel(巴塞尔大学) ICSI, LBNL, and University of California at Berkeley(ICSI、劳伦斯伯克利国家实验室和加州大学伯克利分校)

专题命中 领域大模型 :foundation model(abstract);分类 cs.LG

AI总结 本文提出NeurDE方法,通过结合守恒律与神经网络,实现对非线性守恒律系统更精确的长期预测,优于现有SciML方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16639 2026-05-19 cs.LG 57%

MedMIX: Modality-Internal Expert Fusion for Multimodal Medical Diagnosis

MedMIX:多模态医学诊断中的模态内部专家融合

Seungik Cho, Anqi Li, Wei Qiu

机构 * Department of Physics and Astronomy(物理与天文学系) Department of Electrical and Computer Engineering(电气与计算机工程系) Rice University(里奇大学)

专题命中 领域大模型 :foundation model(abstract);分类 cs.LG

AI总结 MedMIX通过融合模态内部专家、跨模态学习融合及大-小模型协作,提升多模态医学预测的鲁棒性,适用于缺失模态的场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16475 2026-05-19 cs.DL cs.CL 57%

Generative Artificial Intelligence for Literature Reviews

生成式人工智能用于文献综述

Gerit Wagner, Julian Prester, Reza Mousavi, Roman Lukyanenko, Guy Pare

机构 * School of Strategy, Innovation and Technology, University of Sydney(悉尼大学战略、创新与技术学院) McIntire School of Commerce, University of Virginia(弗吉尼亚大学麦克蒂尔商学院)

专题命中 领域大模型 :language model(abstract);分类 cs.CL

AI总结 本文探讨生成式人工智能在文献综述中的应用,分析其在文本摘要、问答、数据提取和翻译等方面的能力,提出使用通用和专用工具进行综述的方法,并讨论其机遇与风险,以及对科学进步的影响。

Journal ref Journal of Information Technology, 02683962261425675 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21787 2026-05-19 cs.CV 50%

Benchmarking Recurrent Event-Based Object Detection for Industrial Multi-Class Recognition on MTevent

在MTevent上评估用于工业多类识别的循环事件基目标检测基准

Lokeshwaran Manohar, Moritz Roidl

机构 * Chair of Material Handling and Warehousing, TU Dortmund University, Dortmund, Germany(物料搬运与仓储学系,杜伊斯堡-艾森大学,多特蒙德,德国)

专题命中 领域大模型 :pretraining(abstract)

AI总结 本文研究了在MTevent数据集上使用循环ReYOLOv8s进行工业多类识别的性能,并通过非循环YOLOv8s作为基线分析时间记忆的影响,发现事件域预训练对性能提升更有效。

Comments Accepted at the Neuromorphic Field Robotics and Automation Workshop, ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06384 2026-05-19 eess.IV cs.CV 50%

Mitigating 3D Prostate Biparametric MRI Data Scarcity through Domain Adaptation using Locally-Trained Latent Diffusion Models for Prostate Cancer Detection

通过使用本地训练的潜在扩散模型进行领域适应以缓解3D前列腺双参数MRI数据稀缺问题

Emerson P. Grabke, Babak Taati, Masoom A. Haider

机构 * Institute of Biomedical Engineering, University of Toronto(多伦多大学生物医学工程研究所) Lunenfeld-Tanenbaum Research Institute, Mount Sinai Hospital(圣心医院卢内尔-塔内本研究所) KITE Research Institute, Toronto Rehabilitation Institute, University Health Network(多伦多康复研究所、KITE研究所在大学健康网络) Joint Department of Medical Imaging, University of Toronto, Princess Margaret Hospital, and Sinai Health systems(多伦多大学联合医学影像部门、玛格丽特医院及辛纳医疗系统) Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Faculty Affiliate of the Vector Institute, Toronto(向量研究所教职员工)

专题命中 领域大模型 :pretraining(abstract)

AI总结 本文提出CCELLA++,一种新的潜在扩散模型流程,用于同时生成3D双参数前列腺MRI(bpMRI),包括轴向T2加权(AxT2)、高b值扩散系列(HighB)和表观扩散系数图(ADC),以克服数据稀缺问题。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00470 2026-05-19 cs.CV 50%

FG-TreeSeg: Flow-Guided Tree Crown Segmentation without Instance Annotations

FG-TreeSeg:基于流引导的树冠分割无需实例标注

Pengyu Chen, Fangzheng Lyu, Sicheng Wang, Cuizhen Wang

机构 * Department of Geography, University of South Carolina(南卡罗来纳大学地理系) Department of Geography, Virginia Polytechnic Institute and State University(弗吉尼亚理工大学地理系)

专题命中 领域大模型 :foundation model(abstract)

AI总结 本文提出FG-TreeSeg,通过将树冠建模为拓扑流场中的星形凸对象,利用Cellpose-SAM实现无需标注的树冠实例分割,实验表明其在不同传感器和冠层密度下均具有良好的泛化能力。

Comments 5 pages, 8 figures

Journal ref IEEE Geoscience and Remote Sensing Letters, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 知识编辑与模型理解 34 篇

2603.23638 2026-05-19 cs.AI 92%

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

LLM代理能成为CFO吗?在不确定的企业环境中评估长期资源分配

Yi Han, Yan Wang, Lingfei Qian, Haohang Li, Yupeng Cao, Yueru He, Xueqing Peng, Nanhan Shen, Yitao Xu, Yankai Chen, Dongji Feng, Jimin Huang, Xue Liu, Jian-Yun Nie, Sophia Ananiadou

机构 * Georgia Institute of Technology(佐治亚理工学院) The Fin AI Stevens Institute of Technology(史蒂文斯理工学院) Columbia University(哥伦比亚大学) George Mason University(乔治·马歇尔大学) McGill University(麦吉尔大学) Mohamed bin Zayed University of Artificial Intelligence(莫扎大学人工智能学院) California State University, Monterey Bay(加州州立大学蒙特雷湾分校) University of Manchester(曼彻斯特大学) Mila – Quebec Artificial Intelligence Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过EnterpriseArena模拟器评估LLM在不确定环境下的长期资源分配能力,发现现有模型在复杂任务中表现不足,仅15.4%的试验能持续完整周期。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14125 2026-05-19 cs.CL 89%

Polar probe linearly decodes semantic structures from LLMs

极性探针线性解码LLM中的语义结构

Pablo J. Diego-Simón, Pierre Orhan, Emmanuel Chemla, Yair Lakretz, Jean-Rémi King

机构 * LSCP, ENS, PSL, EHESS, CNRS(LSCP、ENS、PSL、EHESS、CNRS) Paris Brain Institute(巴黎脑研究所) Earth Species Project(地球物种计划) Meta AI

专题命中 知识编辑与模型理解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究通过极性探针线性恢复LLM中的语义结构,发现其基于嵌入距离和方向表示实体存在与关系类型,且在中层表现更优,能泛化至新实体但随语义结构规模下降。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16699 2026-05-19 cs.CL cs.AI 88%

Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents

校准后再行动:在大语言模型代理中考虑成本的探索

Wenxuan Ding, Nicholas Tomlin, Greg Durrett

机构 * New York University(纽约大学)

专题命中 知识编辑与模型理解 :LLM(title,summary_cn);分类 cs.CL、cs.AI

AI总结 本文提出Calibrate-Then-Act框架,使LLM代理在不确定环境下显式平衡成本与不确定性,从而更优地决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20562 2026-05-19 cs.CL cs.AI 87%

Permutation-Consensus Listwise Judging for Robust Factuality Evaluation

排列一致性列表判断用于鲁棒事实性评估

Tianyi Huang, Nathan Huang, Justin Tang, Wenqian Chen, Elsa Fan

机构 * App-In Club(App-In俱乐部) Carnegie Mellon University(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出PCFJudge方法,通过多排列重跑列表事实性提示以提高LLM事实性判断的鲁棒性,实验显示其在RewardBench 2 Factuality上显著提升准确率。

Comments Accepted at the Fifth Workshop on Natural Language Generation, Evaluation, and Metrics at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17193 2026-05-19 cs.MA 87%

Multi-LLM Systems Exhibit Robust Semantic Collapse

多语言模型系统表现出稳健的语义崩溃

Weiyi Kong, Shiyang Lai, Jinghua Piao, James Evans

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 研究探讨了多语言模型系统在闭环设置中维持开放知识生产的基本限制,发现系统在语义表示上出现系统性收敛,尽管表面上有词汇变化。

Comments 64 pages, 8 figures, 7 tables; includes Supplementary Information

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16686 2026-05-19 cs.LG 86%

Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates

基于张量结构更新的混合专家LLM可扩展知识编辑

Roman Maksimov, Vladimir Aletov, Dmitry Bylinkin, Daniil Medyakov, Vladimir Solodkin, Aleksandr Beznosikov

机构 * OpenAI DeepSeek-AI Qwen Team(Qwen团队) Shazeer et al.(Shazeer等人) Molodtsov et al.(Molodtsov等人)

专题命中 知识编辑与模型理解 :LLM(title_cn,summary_cn);分类 cs.LG

AI总结 本文提出一种针对混合专家架构LLM的知识编辑方法,通过张量结构和Woodbury矩阵恒等式实现高效参数更新,提升编辑效率6倍,扩展了知识编辑的应用范围。

Comments 17 pages, 3 architectures, 1 figure, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01235 2026-05-19 cs.SD cs.AI 86%

MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention

MindMelody:一种基于EEG的闭环个性化音乐干预系统

Yimeng Zhang, Yueru Sun, Haoyu Gu, Zhanpeng Jin

机构 * South China University of Technology(南方科技大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出MindMelody系统,通过EEG信号实时生成个性化音乐,结合Transformer-GNN和RAG-LLM实现情绪感知与音乐生成的闭环控制,提升情感适应性与用户参与度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13339 2026-05-19 cs.CL cs.AI 86%

Probing Persona-Dependent Preferences in Language Models

探测语言模型中依赖人格的偏好

Oscar Gilg, Pierre Beckmann, Daniel Paleka, Patrick Butlin

机构 * MATS EPFL(瑞士联邦理工学院) ETH Zürich(苏黎世联邦理工学院) Eleos AI Research(Eleos AI研究)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);post-training(abstract);分类 cs.CL、cs.AI

AI总结 本文通过训练线性探针来探测语言模型中不同人格下的偏好表示,发现偏好向量在不同人格间具有共享特性,并能通过调整偏好向量来影响模型的输出选择。

Comments 41 pages, 45 figures. Code: https://github.com/oscar-gilg/Preferences. Earlier write-up on LessWrong: https://www.lesswrong.com/posts/pxC2RAeoBrvK8ivMf/models-have-linear-representations-of-what-tasks-they-like-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10134 2026-05-19 cs.CR cs.AI cs.CL 86%

Reverse-Engineering Model Editing on Language Models

语言模型上的逆向工程模型编辑

Zhiyu Sun, Minrui Luo, Yu Wang, Zhili Chen, Tianxing He

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract_cn);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了语言模型中参数编辑的漏洞,提出了一种名为KSTER的逆向工程攻击方法,通过利用参数更新的低秩结构恢复编辑数据,并提出subspace camouflage防御策略以降低重建风险。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18163 2026-05-19 cs.AI cs.CL 85%

TRACE: Trajectory Correction from Cross-layer Evidence for Hallucination Reduction

TRACE: 通过跨层证据进行轨迹修正以减少幻觉

Tej Sanibh Ranade

机构 * Independent Researcher(独立研究者)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);pretraining(abstract);分类 cs.CL、cs.AI

AI总结 本文提出TRACE算法,通过跨层证据在推理时修正LLM中的幻觉,无需训练或标注,通过内部证据选择修正策略,提升多个基准测试的性能。

Comments 25 pages, 8 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02178 2026-05-19 cs.CL cs.AI cs.LG 85%

The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level

专家反击:在专家层面解读混合专家语言模型

Jeremy Herbst, Stefan Wermter, Jae Hee Lee

机构 * Department of Informatics, University of Hamburg, Hamburg, Germany(汉堡大学信息学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过k稀疏探测比较MoE专家与密集FFN,发现专家神经元更单语义,提出以专家为分析单位,揭示专家是细粒度任务专家,而非领域专家或token处理者。

Comments 8 pages, 7 Figures. Accepted at ICML 2026. Improved writing, changed author order, updated citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17435 2026-05-19 cs.CL 84%

BELIEF: Structured Evidence Modeling and Uncertainty-Aware Fusion for Biomedical Question Answering

BELIEF: 结构化证据建模与不确定性感知融合用于生物医学问答

Chang Zong, Hao Ning, Siliang Tang, Jie Huang, Jian Wan

机构 * School of Computer Science and Technology, Zhejiang University of Science and Technology(浙江理工大学计算机科学与技术学院) College of Artificial Intelligence, Zhejiang University(浙江大学人工智能学院) Zhejiang Key Laboratory of Biomedical Intelligent Computing Technology, Zhejiang University of Science and Technology(浙江理工大学生物医学智能计算技术重点实验室)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 本文提出BELIEF框架,通过结构化证据建模和不确定性感知融合,提升生物医学问答任务中检索文献的利用效率,实现对证据可靠性、不确定性以及候选假设的支持强度的显式建模。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02028 2026-05-19 cs.CL 83%

Language models fail at extended rule following

语言模型在扩展规则遵循中表现不佳

Tianxiang Dai, Jonathan Fan

机构 * Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 研究发现语言模型在重复应用规则时无法保持精确状态,即使增加模型规模和计算资源也无法克服这一缺陷,表明需要新的模型架构来实现可靠的规则遵循。

Comments for accessing the data and code for reproduction of the study, see https://txdai.github.io/counting-reliability/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18442 2026-05-19 cs.RO 82%

SG-CADVLM: A Context-Aware Decoding Powered Vision Language Model for Safety-Critical Scenario Generation

SG-CADVLM: 一种基于上下文感知解码的视觉语言模型,用于安全关键场景生成

Hongyi Zhao, Shuo Wang, Qijie He, Ziyuan Pu

机构 * School of Transportation, Southeast University(东南大学交通学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract)

AI总结 本文提出SG-CADVLM,一种结合上下文感知解码的多模态输入处理框架,用于从事故报告中生成高保真的安全关键场景,通过减少视觉语言模型的幻觉并同时生成道路几何和车辆轨迹,提升了生成场景的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏