arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12565 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 12565 篇

2603.16880 2026-04-03 eess.SP cs.CL cs.LG q-bio.NC 86%

NeuroNarrator: A Generalist EEG-to-Text Foundation Model for Clinical Interpretation via Spectro-Spatial Grounding and Temporal State-Space Reasoning

NeuroNarrator:一种通过频谱-空间定位和时间状态空间推理的通用EEG到文本基础模型,用于临床解释

Guoan Wang, Shihao Yang, Jun-en Ding, Hao Zhu, Feng Liu

机构 * Stevens Institute of Technology(史蒂文斯理工学院)

专题命中 领域大模型 :foundation model(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 NeuroNarrator通过频谱-空间定位和时间状态空间推理,将EEG信号转化为临床文本,解决传统EEG分析方法在临床解释中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18602 2026-04-01 cs.NE cs.AI cs.LG 86%

LLM-Meta-SR: In-Context Learning for Evolving Selection Operators in Symbolic Regression

LLM-Meta-SR: 上下文学习用于符号回归中演化的选择算子

Hengzhe Zhang, Qi Chen, Bing Xue, Wolfgang Banzhaf, Mengjie Zhang

机构 * Centre for Data Science and Artificial Intelligence & School of Engineering and Computer Science, Victoria University of Wellington(惠灵顿维多利亚大学数据科学与人工智能中心及工程与计算机科学学院) Department of Computer Science and Engineering, Michigan State University(密歇根州立大学计算机科学与工程系)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种元学习框架,使LLM自动设计进化符号回归算法的选择算子,解决现有方法中语义指导不足和代码膨胀问题,实验表明LLM设计的选择算子优于九个专家设计基线,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21636 2026-03-31 cs.AI cs.CL 86%

Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks

硅 bureaucracy 与 AI 考试导向教育:LLM 测试基准中的污染敏感性与分数可信度

Yiliang Song, Hongjun An, Jiangan Chen, Xuanchen Yan, Huan Song, Jiawei Shao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院(TeleAI)) Guangxi Normal University(广西师范大学) Northwestern Polytechnical University(西北工业大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文探讨了基于测试基准的LLM评估体系,指出其依赖于基准分数直接反映真实泛化能力的假设,并提出审计框架以分析污染敏感性和分数可信度。

Comments Remove the NeurIPS 2026 template

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23520 2026-03-26 cs.CL cs.AI 86%

From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM

从医生专业能力到临床代理:通过轻量级大语言模型保留、标准化和扩展医生的医学专业知识

Chanyong Luo, Jirui Dai, Zhendong Wang, Kui Chen, Jiaxi Yang, Bingjie Lu, Jing Wang, Jiaxin Hao, Bing Li, Ruiyang He, Yiyu Qiao, Chenkai Zhang, Kaiyu Wang, Zhi Liu, Zeyu Zheng, Yan Li, Xiaohong Gu

机构 * the School of Chinese Medicine, the Beijing University of Chinese Medicine, Beijing, China(中国中医药大学中医学院,北京中医药大学,北京,中国) the School of Pharmacy, Nanjing University of Chinese Medicine, Nanjing, China(南京中医药大学药学院,南京,中国) the Infectious disease department, Dongfang Hospital, Beijing University of Chinese Medicine, Beijing, China(北京中医药大学东四医院感染科,北京,中国) the Gulou Hospital of Traditional Chinese Medicine of Beijing,Beijing, China(北京中医药大学附属鼓楼中医医院,北京,中国) the Department of Computer Science, Johns Hopkins University, Baltimore, USA(约翰霍普金斯大学计算机科学系,美国马里兰州巴尔的摩市) the Department of Education, Dongzhimen Hospital, Beijing University of Chinese Medicine, Beijing, China(北京中医药大学东直门医院教育部,北京,中国) the Department of Pediatrics, Wangjing Hospital, China Academy of Chinese Medical Sciences, Beijing, China(中国医学科学院北京协和医院儿科部,北京,中国) the School of Information Engineering, Huzhou University, Huzhou, China(湖州大学信息工程学院,湖州,中国) Research Center for Scientific Data Hub, Zhejiang Lab, Hangzhou, China(浙江省实验室科学数据中心研究中心,杭州,中国) the Frontier Basic Research Center, Zhejiang Lab, Hangzhou, China(浙江省实验室前沿基础研究中心,杭州,中国) the Research Center for High Efficiency Computing Infrastructure, Zhejiang Lab, Hangzhou, China(浙江省实验室高效计算基础设施研究中心,杭州,中国)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Med-Shicheng框架,通过轻量级大语言模型系统学习并转移知名中医的辨证论治理念及案例依赖适应规则,实现医学知识的标准化和规模化应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22765 2026-03-25 cs.CL cs.AI cs.IR 86%

DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona

DALDALL:通过利用LLM-Persona实现法律领域词义多样性的数据增强

Janghyeok Choi, Jaewon Lee, Sungzoon Cho

机构 * Department of Industrial Engineering, Seoul National University, Seoul, South Korea(工业工程系,首尔国立大学,首尔,韩国)

专题命中 领域大模型 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出DALDALL框架,通过利用法律领域专业角色生成高质量合成查询,提升词义多样性并保持语义一致性,在CLERC和COLIEE基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18115 2026-03-20 cs.LG cs.AI 86%

LLM-Augmented Computational Phenotyping of Long Covid

增强型语言模型的长期新冠计算表型分析

Jing Wang, Jie Shen, Amar Sra, Qiaomin Xie, Jeremy C Weiss

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于LLM的Grace Cycle框架,通过迭代整合假设生成、证据提取和特征优化,发现长期新冠的三个临床亚型,为复杂纵向数据的表型筛查提供统计学支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16567 2026-03-18 cs.CL cs.AI 86%

Characterizing Delusional Spirals through Human-LLM Chat Logs

通过人类-大语言模型聊天日志表征妄想螺旋

Jared Moore, Ashish Mehta, William Agnew, Jacy Reese Anthis, Ryan Louie, Yifan Mai, Peggy Yin, Myra Cheng, Samuel J Paech, Kevin Klyman, Stevie Chancellor, Eric Lin, Nick Haber, Desmond C. Ong

机构 * Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学) University of Chicago(芝加哥大学) Independent Researcher(独立研究者) Harvard Belfer Center(哈佛贝尔弗中心) University of Minnesota(明尼苏达大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文通过分析19名受聊天机器人影响用户的历史对话日志,揭示了妄想螺旋中用户与聊天机器人互动模式及心理危害,提出28项代码用于评估对话中的妄想、自伤和AI拟人化现象。

Comments To appear at ACM FAccT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00917 2026-03-18 cs.CL cs.AI 86%

Prompt Sensitivity and Answer Consistency of Small Open-Source Language Models for Clinical Question Answering in Low-Resource Healthcare

小开源语言模型在低资源医疗场景中的提示敏感性与答案一致性

Shravani Hariprasad

机构 * Independent Researcher(独立研究者)

专题命中 领域大模型 :language model(title,abstract);instruction tuning(abstract);pretraining(abstract);分类 cs.CL、cs.AI

AI总结 研究评估了五种开源模型在三个医疗问答数据集上的表现,发现一致性与准确性独立,Llama 3.2在准确性与可靠性上表现最佳,但领域预训练不足以保证结构化医疗问答的正确性。

Comments 30 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08935 2026-03-12 cs.CV cs.AI cs.CL cs.DL cs.IR 86%

PathoScribe: Transforming Pathology Data into a Living Library with a Unified LLM-Driven Framework for Semantic Retrieval and Clinical Integration

PathoScribe: 将病理数据转化为一个活的图书馆:一种统一的LLM驱动框架,用于语义检索和临床整合

Abdul Rehman Akbar, Samuel Wales-McGrath, Alejadro Levya, Lina Gokhale, Rajendra Singh, Wei Chen, Anil Parwani, Muhammad Khalid Khan Niazi

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 PathoScribe通过统一的LLM框架实现病理数据的语义检索与临床整合,提升病理档案的活用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22159 2026-03-10 cs.CR cs.AI cs.CL 86%

RedSage: A Cybersecurity Generalist LLM

RedSage:一种网络安全通用型大语言模型

Naufal Suryanto, Muzammal Naseer, Pengfei Li, Syed Talal Wasim, Jinhui Yi, Juergen Gall, Paolo Ceravolo, Ernesto Damiani

专题命中 领域大模型 :LLM(title,abstract);pretraining(abstract);post-training(abstract);分类 cs.CL、cs.AI

AI总结 RedSage通过领域感知的代理增强和预训练,提升网络安全专长和通用推理能力,实现开源本地部署。

Comments Published at ICLR 2026; Project page: https://risys-lab.github.io/RedSage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06046 2026-03-10 cs.CL cs.LG 86%

Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information

健康的大语言模型?对大语言模型在英国政府公共卫生信息领域知识的评估

Joshua Harris, Fan Grayson, Felix Feldman, Timothy Laurence, Toby Nonnenmacher, Oliver Higgins, Leo Loman, Selina Patel, Thomas Finnie, Samuel Collins, Michael Borowitz

机构 * UK Health Security Agency (UKHSA)(英国卫生安全局)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本文提出PubHealthBench基准测试,评估LLM在公共卫生领域的知识,发现最新LLM在多项选择题中表现优异,但在自由回答中需进一步改进。

Comments 27 pages, 9 pages main text

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02353 2026-03-10 cs.CL cs.LG 86%

An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph

利用LLM增强的知识图谱对塞内加尔法律文本进行结构化

Oumar Kane, Mouhamad M. Allaya, Dame Samb, Mamadou Bousso

机构 * Sciences and Technologies | Economic and Social Sciences(科技 | 经济和社会科学) Iba Der Thiam University(伊巴德里亚姆大学) Social Sciences, Iba Der Thiam University, Thies, Senegal(社会科学,伊巴德里亚姆大学,泰西,塞内加尔)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 本研究利用LLM增强的知识图谱对塞内加尔法律文本进行结构化处理,提取7967篇文章并构建图数据库,提升法律信息的可访问性和理解性。

Comments 8 pages, 8 figures, 2 tables, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19678 2026-03-10 cs.AI cs.LG 86%

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review

从大语言模型推理到自主AI代理:全面综述

Mohamed Amine Ferrag, Norbert Tihanyi, Merouane Debbah

机构 * Department of Computer and Network Engineering(计算机与网络工程系) United Arab Emirates University(阿联酋大学) Technology Innovation Institute(技术创新研究院) Eötvös Loránd University(埃斯特哈齐·洛朗大学) Research Institute for Digital Future(数字未来研究院) Khalifa University(卡塔尔大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文综述了大语言模型和自主AI代理的发展,提出了涵盖多种任务的基准分类法,并讨论了未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04743 2026-03-06 cs.IR cs.AI cs.CL 86%

DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval

DARE: 通过分布感知检索对齐LLM代理与R统计生态系统

Maojun Sun, Yue Wu, Yifei Xie, Ruijian Han, Binyan Jiang, Defeng Sun, Yancheng Yuan, Jian Huang

机构 * Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hong Kong SAR, China(数据科学与人工智能系,香港理工大学,香港特别行政区,中国) Department of Applied Mathematics, The Hong Kong Polytechnic University, Hong Kong SAR, China(应用数学系,香港理工大学,香港特别行政区,中国)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 DARE通过整合数据分布信息提升R包检索效果,构建了面向R的LLM代理以实现更高效的统计分析任务。

Comments 24 pages,7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20095 2026-03-03 cs.CV cs.CL cs.LG 86%

BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models

BioCAP: 利用合成描述性注释超越标签在生物基础模型中的应用

Ziheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G. Campolongo, Matthew J. Thompson, Net Zhang, Samuel Stevens, Hilmar Lapp, Tanya Berger-Wolf, Yu Su, Wei-Lun Chao, Jianyang Gu

机构 * The Ohio State University(俄亥俄州立大学) Duke University(杜克大学) Boston University(波士顿大学)

专题命中 领域大模型 :foundation model(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 BioCAP通过生成合成描述性注释提升生物基础模型的性能,实现物种分类和文本-图像检索的高准确率。

Comments ICLR 2026; Project page: https://imageomics.github.io/biocap/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00451 2026-03-03 cs.AI cs.CL 86%

Confusion-Aware Rubric Optimization for LLM-based Automated Grading

面向混淆的评分标准优化:基于大语言模型的自动评分系统

Yucheng Chu, Hang Li, Kaiqi Yang, Yasemin Copur-Gencturk, Joseph Krajcik, Namsoo Shin, Jiliang Tang

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 CARO通过结构化分离错误信号,提升大语言模型自动评分系统的准确性和效率,有效解决评分逻辑冲突问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02090 2026-03-02 cs.CL cs.AI 86%

LEC-KG: An LLM-Embedding Collaborative Framework for Domain-Specific Knowledge Graph Construction -- A Case Study on SDGs

LEC-KG: 一种基于大语言模型嵌入的领域知识图谱构建协作框架——以可持续发展目标为例

Yikai Zeng, Yingchao Piao, Changhua Pei, Jianhui Li

机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 LEC-KG通过结合大语言模型与知识图谱嵌入的双向协作框架,有效解决领域知识图谱构建中的长尾关系问题,提升低频关系的识别与验证能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12259 2026-02-26 cs.AI cs.LG 86%

Think like a Scientist: Physics-guided LLM Agent for Equation Discovery

像科学家一样思考:基于物理的LLM代理用于方程发现

Jianke Yang, Ohm Venkatachalam, Mohammad Kianezhad, Sharvaree Vadgama, Rose Yu

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 KeplerAgent通过模拟科学家的推理过程,利用物理知识和符号回归引擎,在方程发现任务中实现了更高的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23339 2026-02-23 cs.LG cs.AI physics.chem-ph q-bio.QM 86%

VALID-Mol: a Systematic Framework for Validated LLM-Assisted Molecular Design

VALID-Mol: 一种系统化的验证框架用于验证的LLM辅助分子设计

Malikussaid, Hilal Hudan Nuha, Isman Kurniawan

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 VALID-Mol通过整合化学验证与LLM驱动的分子设计,显著提高了有效化学结构的生成率,同时确保合成可行性和目标结合亲和力的提升。

Comments 6 pages, 1 figure, 1 algorithm, 5 tables, to be published in ISPACS 2025, unabridged version exists as arXiv:2506.23339v1

Journal ref Proc. 2025 Int. Symp. on Intell. Signal Process. and Commun. Syst. (ISPACS), 2025, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11733 2026-02-05 cs.CY cs.AI cs.CL cs.HC 86%

LLM Agents for Education: Advances and Applications

教育中的LLM代理:进展与应用

Zhendong Chu, Shen Wang, Jian Xie, Tinghui Zhu, Yibo Yan, Jinheng Ye, Aoxiao Zhong, Xuming Hu, Jing Liang, Philip S. Yu, Qingsong Wen

机构 * Squirrel Ai Learning Fudan University(复旦大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了LLM代理在教育中的应用进展,探讨了其技术实现、挑战及在不同教育领域的应用。

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04206 2026-02-05 cs.CL cs.AI 86%

Enforcing Monotonic Progress in Legal Cross-Examination: Preventing Long-Horizon Stagnation in LLM-Based Inquiry

在法律质证中强制单调进展:防止基于LLM的探究中长周期停滞

Hsien-Jyh Liao

机构 * Taiwan, ROC(台湾,中华民国)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Soft-FSM架构,通过外部确定性状态控制器在法律质证中强制单调进展,显著提升任务完成率至97%以上。

Comments Submitted to ICAIL 2026. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18506 2026-02-05 cs.LG cs.AI 86%

LLM-ABBA: Understanding time series via symbolic approximation

LLM-ABBA:通过符号近似理解时间序列

Xinye Chen, Erin Carson, Cheng Kang

机构 * Sorbonne Université, CNRS, LIP6(索邦大学、法国国家科学研究中心、LIP6实验室) Department of Numerical Mathematics, Charles University(数值数学系、查尔斯大学) Department of Cybernetics, Czech Technical University in Prague(控制论系、布拉格捷克技术大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 LLM-ABBA通过符号近似方法在时间序列任务中实现高性能,适用于分类、回归和预测等多种下游任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19120 2026-02-04 cs.IR cs.AI cs.LG 86%

RobustExplain: Evaluating Robustness of LLM-Based Explanation Agents for Recommendation

RobustExplain: 评估基于LLM的推荐解释代理的鲁棒性

Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu

机构 * Workday

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 RobustExplain提出首个评估框架,用于衡量LLM生成推荐解释的鲁棒性,揭示当前模型鲁棒性较低,较大模型稳定性更高。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00699 2026-02-03 cs.AI cs.CL cs.IR 86%

From Prompt to Graph: Comparing LLM-Based Information Extraction Strategies in Domain-Specific Ontology Development

从提示到图:比较基于大语言模型的信息提取策略在领域特定本体构建中的应用

Xuan Liu, Ziyu Li, Mu He, Ziyang Ma, Xiaoxu Wu, Gizem Yilmaz, Yiyuan Xia, Bingbing Li, He Tan, Jerry Ying Hsi Fuh, Wen Feng Lu, Anders E. W. Jarfors, Per Jansson

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文比较了三种基于大语言模型的信息提取策略,用于在有限数据下构建领域特定本体,通过实验验证了最佳方法的有效性。

Comments 11 pages,8 figures,3 tables,presented at International Conference on Industry of the Future and Smart Manufacturing,2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16584 2026-02-03 cs.CL cs.AI 86%

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

从分数到步骤:诊断和改进证据医学计算中LLM的性能

Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao

机构 * Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院) Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedRaC框架,通过分步评估和代码执行提升LLM在证据医学计算中的准确性,揭示现有评估方法的不足,并推动临床可信度的提升。

Comments Equal contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18217 2026-01-27 cs.AI cs.LG 86%

Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents

支付更少的泛化税:RL训练对LLM代理跨域泛化能力的研究

Zhihan Liu, Lin Guan, Yixin Nie, Kai Zhang, Zhuoqun Hao, Lin Chen, Asli Celikyilmaz, Zhaoran Wang, Na Zhang

机构 * Meta Superintelligence Labs(Meta超智能实验室) FAIR at Meta(Meta的FAIR) Northwestern University(西北大学) The Ohio State University(俄亥俄州立大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 领域大模型 :LLM(title,abstract);post-training(abstract);SFT(abstract);分类 cs.AI、cs.LG

AI总结 本研究探讨了RL训练对LLM代理跨域泛化能力的影响,发现增加状态信息丰富度可提升泛化性能,同时指出建模选择对泛化能力的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06161 2026-01-26 cs.LG cs.AI 86%

LATTLE: LLM Attention Transplant for Transfer Learning of Tabular Data Across Disparate Domains

LATTLE: LLM注意力移植用于跨异质领域的表格数据迁移学习

Ibna Kowsar, Kazi F. Akhter, Manar D. Samad

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 LATTLE通过LLM注意力移植方法,实现了跨异质领域表格数据的高效迁移学习,无需共享特征或大规模预训练模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07890 2026-01-15 cs.MA cs.AI cs.LG stat.ME stat.ML 86%

CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models

CrowdLLM: 构建基于大语言模型的数字人群并整合生成模型

Ryan Feng Lin, Keyu Tian, Hanming Zheng, Congjing Zhang, Li Zeng, Shuai Huang

机构 * Department of Industrial and Systems Engineering, University of Washington(华盛顿大学工业与系统工程系) Department of Data Science, City University of Hong Kong(香港城市大学数据科学系)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 CrowdLLM通过整合预训练大语言模型和生成模型,提升数字人群的多样性和保真度,适用于社交模拟、众包等多领域应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06780 2026-01-14 cs.CL cs.AI 86%

Foundations of LLM Knowledge Materialization: Termination, Reproducibility, Robustness

LLM知识材料化的基础:终止性、可重现性与鲁棒性

Luca Giordano, Simon Razniewski

机构 * ScaDS.AI Dresden/Leipzig & TU Dresden, Germany(ScaDS.AI 德累斯顿/莱比锡及德累斯顿技术大学,德国)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了LLM知识材料化的终止性、可重现性和鲁棒性,通过实验揭示了不同因素对知识提取效果的影响。

Comments Accepted and published in Findings of EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05376 2026-01-12 cs.AI cs.CL 86%

The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models

医疗人设悖论:临床语言模型中的行为先验

Tassallah Abdullahi, Shrestha Ghosh, Hamish S Fraser, Daniel León Tramontini, Adeel Abbasi, Ghada Bourjeily, Carsten Eickhoff, Ritambhara Singh

机构 * Brown University(布朗大学) University of Tuebingen(图宾根大学)

专题命中 领域大模型 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 研究揭示医疗人设在临床语言模型中产生上下文依赖的性能权衡,显示其在重症护理任务中提升表现,但在初级护理中降低性能,且互动风格影响风险倾向。

详情

展开后加载摘要…

URL PDF HTML 收藏