arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-06-23 至 2026-06-23 共收录 778 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 71 篇

2606.21676 2026-06-23 cs.IR 新提交 67%

CRAwLeR -- Cross-Reference Aware Legal Retrieval

CRAwLeR -- 交叉引用感知的法律检索

Maciej Jalocha, William Michelsen

专题命中 领域大模型 :LLM(abstract,abstract_cn)

AI总结 针对法律文档中的交叉引用,提出CRAwLeR流水线,生成需要上下文的查询并构建数据集,实验表明约80%的查询真正需要上下文,最佳Recall@10为55%-59%。

Comments 26 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20466 2026-06-23 cs.HC 67%

Piloting Planetarium Visualizations with LLMs during Live Events in Science Centers

在科学中心现场活动期间利用LLM试点行星馆可视化

Mathis Brossier, Mujtaba Fadhil Jawad, Emma Broman, Ylva Selling, Julia Hallsten, Alexander Bock, Johanna Björklund, Tobias Isenberg, Anders Ynnerman, Mario Romero, Lonni Besançon

专题命中 领域大模型 :LLM(title_cn)

AI总结 研究利用AI试点提升科学中心公共展览体验,通过7名专业向导评估AI试点在减少人力负担和实现多任务处理中的作用。

Comments Submitted to Posters of CHI'26

Journal ref Short Paper Proceedings of EuroVis 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20913 2026-06-23 cs.CV cs.AI cs.LG 新提交 62%

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

PROTON: 基于原型的测试时在线OOD检测方法用于医学视觉语言模型

Abhijit Das, Nichula Wasalathilaka, Yifan Lu, Adinath Dukre, Dwarikanath Mahapatra, Shadab Khan, Imran Razzak

机构 * MBZUAI(穆罕默德·本·扎耶德人工智能大学) University of Peradeniya(佩拉德尼亚大学) Khalifa University(哈利法大学) ADIA Lab(阿布扎比投资局实验室) MedOS

专题命中 领域大模型 :language model(abstract);分类 cs.AI、cs.LG

AI总结 针对医学视觉语言模型在部署时难以检测分布外输入的问题,提出PROTON方法,通过在线原型库和自适应融合原型距离与最大概念匹配得分,无需修改模型或训练数据,在多个OOD场景下提升检测性能。

Journal ref 29th International Conference on Medical Image Computing and Computer Assisted Intervention 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23301 2026-06-23 cs.AI 新提交 57%

EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning

EHR-Complex:面向复杂临床推理的医疗智能体基准测试

Yitong Qiao, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu, Kui Ren

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团)

专题命中 领域大模型 :LLM(abstract_cn);分类 cs.AI

AI总结 提出EHR-Complex基准,基于MIMIC-IV构建约5.2万项交互式临床数据库推理任务,评估智能体在复杂SQL和代码执行中的表现,揭示主要失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22771 2026-06-23 cs.CL 新提交 57%

Learning Moral Diversity: Modelling Individual Perspectives in Moral Classification of Texts

学习道德多样性:在文本道德分类中建模个体视角

Yi Ren, Lewis Mitchell, Matthew Roughan

机构 * School of Mathematical Sciences, Adelaide University(阿德莱德大学数学科学学院)

专题命中 领域大模型 :language model(abstract);分类 cs.CL

AI总结 针对社交媒体文本道德分类中标注者主观性导致的分歧,提出在预训练语言模型中加入学习标注者特定特征的层,改进个体标注预测并揭示道德视角差异。

Comments Accepted at the Seventh Workshop on NLP and Computational Social Science. 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22486 2026-06-23 cs.CV cs.AI cs.HC 新提交 57%

Human and AI collaboration for pulmonary nodule segmentation

人类与AI协作进行肺结节分割

Hongqiao Dong, Wenhao Chi, Ruobing Liang, Xiaokui Yang, Wenhua Liang, Peng Hou, Wenjun Pu, Yipeng Zhao, Ping Chen, Haiping Liu, Jianxing He, Bo Liu

机构 * State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences(数学科学国家重点实验室,数学与系统科学研究院,中国科学院) PET/CT Center, The First Affiliated Hospital of Guangzhou Medical University(广州医科大学第一附属医院PET/CT中心) Department of Thoracic Surgery and Oncology, The First Affiliated Hospital of Guangzhou Medical University(广州医科大学第一附属医院胸外科与肿瘤科) Department of Mathematical Science, Tsinghua University(清华大学数学科学系) Yau Mathematical Sciences Center, Tsinghua University(清华大学尤金数学科学中心) China State Key Laboratory of Respiratory Disease & National Clinical Research Centre for Respiratory Disease, Guangzhou, China(中国呼吸疾病国家重点实验室及呼吸疾病临床研究中心,广州,中国) School of Mathematical Sciences, University of Chinese Academy of Sciences(中国科学院大学数学科学学院) National Center for Respiratory Medicine, National Clinical Research Center for Respiratory Disease, Guangzhou Institute of Respiratory Health, The First Affiliated Hospital of Guangzhou Medical University(呼吸医学国家中心、呼吸疾病临床研究中心、广州呼吸健康研究院、广州医科大学第一附属医院)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 提出Hi-Seg框架,基于SAM通过人类迭代优化提示实现肺结节分割,在12中心1179例CT上平均Dice达85%,优于多种深度学习模型,并降低标注时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22002 2026-06-23 cs.CV cs.LG 新提交 57%

One-Shot Data Selection for Medical Image Classification via Graph Coverage

基于图覆盖的医学图像分类一次性数据选择

Zahiriddin Rustamov, Nadia Badawi, Rafat Damseh, Nazar Zaki

机构 * United Arab Emirates University(阿拉伯联合酋长国大学) KU Leuven(鲁汶大学)

专题命中 领域大模型 :foundation model(abstract);分类 cs.LG

AI总结 提出一种基于图的一次性数据选择方法,利用预训练编码器的k近邻图构建热扩散核,通过贪婪设施位置选择最大化数据流形覆盖的子集,在五个MedMNIST数据集上优于基线方法。

Comments Accepted at MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20873 2026-06-23 cs.CL 新提交 57%

SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding

SciLens: 多模态科学声明验证的智能蕴含与归因框架

Yueming Wang, Tianshi Zheng, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See

机构 * The Hong Kong University of Science and Technology(香港科技大学) Hong Kong Baptist University(香港浸会大学) NVIDIA AI Technology Center(英伟达人工智能技术中心)

专题命中 领域大模型 :language model(abstract);分类 cs.CL

AI总结 提出SciLens框架,通过将声明分解为原子命题并归因到表格/图表证据,实现多模态科学声明验证,在SciClaimEval上达到79.2%宏F1和63.1%配对准确率。

Comments KDD 2026 SciSoc Agents & LLMs (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20677 2026-06-23 cs.AI cs.CV 新提交 57%

Democratizing and accelerating AI-driven pathology research through agentic intelligence

通过智能体智能实现AI驱动病理学研究的民主化与加速

Jiabo Ma, Cheng Jin, Yihui Wang, Hao Jiang, Ling Liang, Yingxue Xu, Junlin Hou, Zhengrui Guo, Zhengyu Zhang, Yifei Xia, Hongyi Wang, Fengtao Zhou, Zhe Xu, Huajun Zhou, Jiarui Ouyang, Qian Zeng, On Ki Tang, Eunhyang Park, Carolyn Glass, Ronald Cheong Kin Chan, Li Liang, Hao Chen

机构 * Hong Kong University of Science and Technology(香港科技大学) Southern Medical University(南方医科大学) Nanfang Hospital(南方医院) Guangdong Province Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理重点实验室) Jinfeng Laboratory(金凤实验室)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 提出PathLab框架,通过结构化组合领域技能和工具,将自然语言研究目标转化为可执行的计算病理学工作流,在12个数据集上达到与专家实现相当的性能,并显著降低编程门槛。

Comments 29 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21177 2026-06-23 eess.IV cs.AI cs.CV physics.med-ph 新提交 57%

Anatomically Consistent TMJ Disc Segmentation via Semantic Anchoring and Clinical Priors

基于语义锚定和临床先验的解剖一致性颞下颌关节盘分割

Dayun Ju, Chanyoung Kim, Sunyoung Jung, Hyo-Jung Jung, Chena Lee, Younjung Park, Seong Jae Hwang

机构 * Department of Artificial Intelligence, Yonsei University(燕山大学人工智能学院) Department of Computer Science, Emory University(埃默里大学计算机科学系) College of Dentistry, Yonsei University(燕山大学牙科学院)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 提出TISC框架,通过原型语义锚定模块定位关节盘,并利用临床元数据(开口受限指标)细化边界,在2488例MRI上实现Dice提升4.96,生成解剖一致的分割。

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04579 2026-06-23 cs.AI 版本更新 57%

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

SCI-PRM:用于科学推理验证的工具感知过程奖励模型

Xiangyu Zhao, Henry Hengyuan Zhao, Yiheng Wang, Wanghan Xu, Yuhao Zhou, Qinglong Cao, Zhiwang Zhou, Lei Bai, Wenlong Zhang, Xiao-Ming Wu

机构 * The Hong Kong Polytechnic University(香港理工大学) Shanghai AI Lab(上海人工智能实验室) National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) Sichuan University(四川大学) Tongji University(同济大学)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 针对科学推理中工具使用和事实一致性问题,提出Sci-PRM模型,通过构建包含工具链轨迹的数据集SCIPRM70K并训练过程奖励模型,在测试时扩展和强化学习中提供细粒度监督,提升基础模型性能。

Comments Accepted by KDD 2026 AI4Science Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21563 2026-06-23 cs.LG 版本更新 57%

Embedding-Based Federated Learning with Runtime Governance for Iron Deficiency Prediction

基于运行时治理的嵌入式联邦学习用于缺铁预测

Fan Zhang, Simon Deltadahl, Majid Lotfian Delouee, Daniel Kreuter, Joseph Taylor, Allerdien Visser, BloodCounts Consortium, James H. F. Rudd, Nicholas S. Gleadall, Suthesh Sivapalaratnam, Folkert Asselbergs, Martijn C. Schut, Michael Roberts

机构 * Theoretical Physics University of Cambridge Cambridge, UK Translational AI Laboratory, Dept. of Laboratory Medicine Amsterdam UMC Amsterdam, The Netherlands Precision Health University Research Institute Queen Mary Univ. of London London, UK Department of Medicine University of Cambridge Cambridge Biomedical Campus Cambridge, UK Transplant Cambridge Biomedical Campus Cambridge, UK Dept. of Cardiology Amsterdam Cardiovascular Sciences Amsterdam UMC Amsterdam, The Netherlands

专题命中 领域大模型 :foundation model(abstract);分类 cs.LG

AI总结 本文提出了一种基于嵌入的联邦学习框架,用于从常规全血计数数据中预测缺铁,并在两个临床环境中部署,展示了个性化聚合方法在处理不同样本量和任务相关性时的优越性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16625 2026-06-23 cs.AI 版本更新 57%

MedBayes-Lite: A Clinical Uncertainty Governance Layer for Risk-Aware Medical Decision Support

MedBayes-Lite: 用于风险感知医疗决策支持的临床不确定性治理层

Elias Hossain, Md Mehedi Hasan Nipu, Maleeha Sheikh, Tasfia Nuzhat, Rajib Rana, Subash Neupane, Björn W. Schuller, Niloofar Yousefi

机构 * College of Engineering and Computer Science, University of Central Florida(中央佛罗里达大学工程与计算机科学学院) Department of Computer Science and Engineering, North South University(北方南大学计算机科学与工程系) Department of Electrical and Computer Engineering, Purdue University Fort Wayne(普渡大学福克斯堡分校电气与计算机工程系) School of Mathematics, Physics and Computing, University of Southern Queensland(南方昆士兰大学数学、物理与计算学院) Meharry Medical College(梅哈里医学学院) CHI – Chair of Health Informatics, Technical University of Munich (TUM)(慕尼黑技术大学健康信息学系) GLAM – Group on Language, Audio, & Music, Imperial College London(伦敦帝国理工学院语言、音频与音乐组)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 提出无需重训练的MedBayes-Lite层,结合MC Dropout、校准和置信度引导的弃权,减少临床问答中高置信度错误,将校准误差降低0.23-0.33,高严重性错误降至近零。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05161 2026-06-23 cs.CV 版本更新 50%

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

基于VLM的医学图像分布外检测的Wasserstein对齐定位

Bernhard Kainz, Johanna P Mueller, Matthew Baugh, Cosmin Bercea

机构 * Department of Computing, Imperial College London, UK(伦敦帝国理工学院计算机系) Technical University Munich, DE(慕尼黑技术大学) Munich Center for Machine Learning (MCML), DE(慕尼黑机器学习中心(MCML))

专题命中 领域大模型 :language model(abstract)

AI总结 提出WALDO框架,利用最优传输理论通过熵加权切片Wasserstein距离、Goldilocks区域采样和自一致性聚合实现零样本异常定位,在NOVA脑MRI基准上mAP@30达43.5%,相对提升19%。

Comments submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22038 2026-06-23 eess.SP 新提交 50%

Full-Domain Coupler: A Wireless Native Neural Backbone for Channel Representation and Deduction

全域耦合器:一种用于信道表示与推演的无线原生神经骨干网络

Zirui Chen, Ziqing Xing, Zhaoyang Zhang, Hongning Ruan, Yuzhi Yang, Zhaohui Yang, Chongwen Huang, Merouane Debbah

专题命中 领域大模型 :foundation model(abstract)

AI总结 针对无线数据多域深度耦合问题,提出无线原生AI神经骨干网络Coupler,通过逐层域分解与交错级联实现高效信道表示,在信道推演任务中取得显著性能提升。

Comments The code and data supporting this work are available at https://github.com/XIronMan0220/Coupler-Channel-Deduction

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17090 2026-06-23 q-bio.CB 版本更新 50%

Intracellular Measurement-Informed Multiscale Modeling for Scalable iPSC Manufacturing

细胞内测量驱动的多尺度建模用于可扩展iPSC制造

Fuqiang Cheng, Zahra Foroozan Jahromi, Keqi Wang, Thomas C. Caldwell, Grace Cai, Keilung Choy, Jared Auclair, Jeffrey L. Campbell, Youbo Zhao, Jane Ring, Seongkyu Yoon, Sarah W. Harcum, Wei Xie

专题命中 领域大模型 :foundation model(abstract)

AI总结 针对3D聚集体培养中的空间和代谢异质性,开发了整合分子、细胞和宏观过程的多尺度机制模型,通过结合单层动力学网络与生物系统-of-系统模型,统一了异质性数据,为可扩展iPSC生物制造提供了定量基础。

Comments 33 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24160 2026-06-23 eess.IV cs.CV 版本更新 50%

Beyond the LUMIR challenge: The pathway to foundational registration models

超越LUMIR挑战:走向基础配准模型

Junyu Chen, Shuwen Wei, Joel Honkamaa, Pekka Marttinen, Hang Zhang, Min Liu, Yichao Zhou, Zuopeng Tan, Zhuoyuan Wang, Yi Wang, Hongchao Zhou, Shunbo Hu, Yi Zhang, Qian Tao, Lukas Förner, Thomas Wendler, Bailiang Jian, Benedikt Wiestler, Tim Hable, Jin Kim, Dan Ruan, Frederic Madesta, Thilo Sentker, Wiebke Heyer, Lianrui Zuo, Yuwei Dai, Jing Wu, Jerry L. Prince, Harrison Bai, Yong Du, Yihao Liu, Alessa Hering, Reuben Dorent, Lasse Hansen, Mattias P. Heinrich, Aaron Carass

机构 * The Russell H. Morgan Department of Radiology(Russell H. Morgan放射科) Radiological Science, Johns Hopkins Medical School(约翰霍普金斯医学院放射科学) Department of Computer Science, Aalto University(阿尔托大学计算机科学系) Cornell University(康奈尔大学) Canon Medical Systems (China) Co. Ltd.(佳能医疗系统(中国)有限公司) School of Biomedical Engineering, Shenzhen University Medical School(深圳大学医学院生物医学工程学院) Department of Imaging Physics, Delft University of Technology(代尔夫特理工大学成像物理系) Technical University of Munich(慕尼黑技术大学) Radboud University Medical Center(拉德伯德大学医学中心) Inria, Paris, France(法国巴黎Inria)

专题命中 领域大模型 :foundation model(abstract)

AI总结 提出LUMIR挑战,通过大规模无监督脑MRI配准任务,验证深度学习方法在生成解剖合理变形场和跨域鲁棒性上的优势,推动通用医学图像配准基础模型的发展。

Comments Accepted to Medical Image Analysis ((c) MedIA). Code available at https://github.com/JHU-MedImage-Reg/LUMIR_L2R

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15205 2026-06-23 cs.CY 版本更新 50%

A Peek Behind the Curtain: Using Step-Around Prompt Engineering to Identify Bias and Misinformation in GenAI Models

幕后窥探:使用逐步提示工程识别生成式AI模型中的偏见与虚假信息

Don Hickerson, Mike Perkins

专题命中 领域大模型 :prompting(abstract)

AI总结 本文探讨逐步提示工程作为对抗性提示技术,用于测试生成式AI的安全护栏和偏见缓解机制,并提出了学术伦理治理框架。

Comments v3 Substantial text changes and addition of structured framework

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 知识编辑与模型理解 46 篇

2606.22792 2026-06-23 cs.AI 新提交 92%

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

随机性的起源:大型语言模型不确定性量化的综合研究

Xiang-Jun Ou, Shuang Liang, Xin-Yu Hu, Rong-Hao Huang, Jing Wang, Shao-Qun Zhang

机构 * National Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) School of Intelligent Science and Technology(智能科学与技术学院)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 提出细粒度不确定性分类法,将LLM不确定性归因于输入、参数、标记和解码过程,并评估21种UQ方法,发现基于共识的方法(Deg和EigV)表现最佳,且模型规模与不确定性负相关。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21616 2026-06-23 cs.CL 新提交 92%

LLM and Human Modes of Representation

LLM与人类的表征模式

Shalom Lappin

机构 * Shalom Lappin

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究比较大型语言模型(LLM)与人类在语言知识和现实推理中的表征差异,发现LLM在语言任务上表现优异但处理方式不同,且在推理任务中学习与泛化效率低于人类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21255 2026-06-23 cs.CL 新提交 92%

SCOPE: Sequential Conformal Probing for Reliable OOD Rejection in LLM Services

SCOPE: 用于LLM服务中可靠OOD拒绝的顺序共形探测

Zhuoyun Li, Boxuan Wang, Changshun Wu, Xiaowei Huang, Yi Dong

机构 * School of Computer Science and Informatics, University of Liverpool, United Kingdom(英国利物浦大学计算机科学与信息学院)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出SCOPE框架,通过选择可读隐藏层、构建共形门控和超鞅e过程,在LLM服务中实现带理论保证的OOD拒绝,实验表明优于标准最终层检测器。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09520 2026-06-23 physics.chem-ph cs.AI 新提交 92%

Closing the Prior-Posterior Loop: Self-Reflective Molecular Design with Analysis-Driven LLM Iteration

闭合先验-后验循环:基于分析驱动LLM迭代的自反性分子设计

Junyi Gong, Zijie Qiu, Ben Zhong Tang

机构 * Faculty of Chemistry, Shenzhen MSU-BIT University(深圳MSU-BIT大学化学学院) School of Science and Engineering, Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)科学与工程学院) Department of Chemistry, Hong Kong University of Science and Technology(香港科技大学化学系)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种自反性分子设计框架,用第一性原理计算的完整物化理由替代标量反馈,使LLM从随机采样器转变为因果推理器,在HOMO-LUMO能隙任务中实现0.0003 eV偏差和100%成功率。

Comments 3 tables, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21645 2026-06-23 cs.CL cs.LG 新提交 91%

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models

大型语言模型中二项式排序偏好的行为与表征证据

Zhiqing Yang, Yilun Liu, Yunpu Ma, Volker Tresp, Hinrich Schütze

机构 * Ludwig Maximilian University of Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.CL、cs.LG

AI总结 通过多语言二项式数据集和分布度量,研究LLM对梯度频率分布的建模能力,发现模型虽能恢复主导顺序,但难以精确匹配偏好分布,并通过稀疏探针验证偏好强度在中间层编码且可操纵。

Comments Code and data are publicly available at https://github.com/Zhi-qing-Yang/Linguistic-Binomials-in-Large-Language-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21195 2026-06-23 cs.CL cs.AI 新提交 91%

Beyond Hooking Onto the World: Referential Profiles and the Numerical Structure of LLM Grounding

超越钩连世界:指称轮廓与LLM基础的数字结构

Joo Yull Rhee

机构 * Sungkyunkwan University(成均馆大学)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文重新审视大语言模型的基础问题,提出指称是基于轮廓、上下文敏感、受情感影响且受规范约束的,并通过优化将语言痕迹参数化为数字结构,支持LLM拥有衍生性、语言中介的指称形式。

Comments 29 pages, no figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21048 2026-06-23 cs.CL 新提交 91%

Event Ontology Expansion via LLM-Based Conceptualization

基于LLM概念化的事件本体扩展

Weicheng Ren, Zixuan Li, Long Bai, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng

机构 * State Key Laboratory of AI Safety(人工智能安全国家重点实验室) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);prompting(abstract);分类 cs.CL

AI总结 提出ConceptE框架,利用LLM从句子和触发词中提取概念级语义,增强事件聚类和层次扩展,在ACE、ERE和MAVEN上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21102 2026-06-23 cs.CY 新提交 90%

Incoherent Values? Probing LLM Preferences Through Parametric Variation

不一致的价值?通过参数变化探测LLM的偏好

Elena Ajayi, Angelica Chowdhury, Seth Lazar

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本文通过参数化变体测试LLM的价值一致性,发现即使最强模型也存在显著不一致,且一致性不随能力提升而出现,但允许推理时间可减少不一致。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21666 2026-06-23 cs.AI cs.CL cs.MA 新提交 90%

Hallucination as Context Drift: Synchronization Protocols for Multi-Agent LLM Systems

幻觉作为上下文漂移:多智能体LLM系统的同步协议

Carson Rodrigues

机构 * Celabe

专题命中 知识编辑与模型理解 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 提出上下文分歧分数(CDS)和共享状态验证协议(SSVP),通过周期性交换压缩状态摘要来减少多智能体LLM系统中的幻觉,实验表明SSVP在旅行规划领域将幻觉率降低至0.463,且API调用减少58%。

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22741 2026-06-23 cs.LG 新提交 90%

GRADE: Graph Representation of LLM Agent Dependency and Execution

GRADE: LLM智能体依赖与执行的图表示

Yue Zhao

机构 * University of Southern California(南加州大学)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);分类 cs.LG

AI总结 提出GRADE,将LLM智能体的运行建模为具有执行边和依赖边的图,依赖边通过分级推断,在多个数据集上优于运行规模指标,并可用于故障定位。

Comments 18 pages, 5 figures, 8 tables. Code: https://github.com/yzhao062/grade

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21399 2026-06-23 cs.AI 新提交 90%

Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

校准不是控制:为什么LLM代理监督需要干预

Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) CFAR Agency for Science Technology and Research(科技研究局CFAR) IHPC Agency for Science Technology and Research(科技研究局IHPC) Department of Statistics and Operations Research UNC-Chapel Hill(北卡罗来纳大学教堂山分校统计与运筹学系)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);分类 cs.AI

AI总结 本文指出LLM代理运行时监督中常用的标量风险预测存在目标错误,提出以干预优势为决策对象,并引入前缀分支方法进行动作条件控制,实验表明该方法能显著降低控制遗憾。

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23921 2026-06-23 cs.CL cs.LG stat.ML 版本更新 90%

The Trilemma of Truth in Large Language Models

大型语言模型中真理的三难困境

Germans Savcisens, Tina Eliassi-Rad

机构 * Khoury College of Computer Sciences Northeastern University(东北大学科里学院) Network Science Institute Northeastern University(东北大学网络科学研究所) Santa Fe Institute(圣菲研究所)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL、cs.LG

AI总结 针对大型语言模型知识真实性探测方法的缺陷,提出结合多实例学习与共形预测的sAwMIL框架,实现真、假及不确定三类分类,揭示真理与虚假的非对称编码及第三种信号的存在。

Comments The main text is 9 pages long (plus 3 pages of references); supplementary material (60 pages) is included in the same PDF

详情

展开后加载摘要…

URL PDF HTML 收藏