arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 18837 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 18837 篇

2412.14368 2026-04-28 cs.CL 88%

Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs

超越数学:故事作为LLM中受记忆限制推理的测试平台

Yuxuan Jiang, Francis Ferraro

机构 * Department of Computer Science and Electrical Engineering(计算机科学与电气工程系) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 推理与问题求解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文通过故事测试平台探讨LLM在记忆受限推理中的表现,提出减少机械记忆的方法,发现记忆驱动性能下降,揭示现有基准测试的数据污染问题。

Comments published on EACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22266 2026-04-27 cs.CL 88%

Large Language Models Decide Early and Explain Later

大语言模型早决策晚解释

Ayan Datta, Zhixue Zhao, Bhuvanesh Verma, Radhika Mamidi, Mounika Marreddy, Alexander Mehler

机构 * University of Sheffield, UK(谢菲尔德大学,英国) Goethe University, Frankfurt am Main, Germany(法兰克福歌德大学,德国)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 研究发现大语言模型在生成过程中,只有32%的查询在中间阶段改变最终答案,且稳定后生成额外760个推理token。通过早停策略可减少500个token使用,仅导致2%的准确率下降。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19815 2026-04-23 cs.AI 88%

Large Language Models Meet Biomedical Knowledge Graphs for Mechanistically Grounded Therapeutic Prioritization

大型语言模型与生物医学知识图谱结合用于机制性指导的治疗优先级排序

Chih-Hsuan Wei, Chi-Ping Day, Zhizheng Wang, Christine C. Alewine, Betty Tyler, Hasan Slika, David Saraf, Chin-Hsien Tai, Joey Chan, Robert Leaman, Zhiyong Lu

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出DrugKLM框架,结合生物医学知识图谱与大语言模型的机制推理,提升治疗优先级排序的准确性与生物学依据。

Comments 24 pages, 5 figures in main text

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20523 2026-04-23 cs.SE cs.AI 88%

Early-Stage Product Line Validation Using LLMs: A Study on Semi-Formal Blueprint Analysis

利用LLM进行早期阶段产品线验证:对半正式蓝图分析的研究

Viet-Man Le, Thi Ngoc Trang Tran, Sebastian Lubos, Alexander Felfernig, Damian Garber

机构 * Graz University of Technology(格拉茨技术大学)

专题命中 推理与问题求解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究LLM能否直接对半正式文本蓝图执行特征模型分析操作,通过12种先进LLM和16种标准操作,对比其与求解器FLAMA的输出,发现优化推理模型在准确性上接近求解器,揭示了结构解析和约束推理的系统性错误及精度-成本权衡。

Comments The 41st ACM/SIGAPP Symposium on Applied Computing (SAC '26), March 23--27, 2026, Thessaloniki, Greece DOI: 10.1145/3748522.3779903

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06168 2026-04-22 cs.AI 88%

Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models

思维链作为镜像:评估大语言模型与人类偏好之间结构化推理的对齐

Boxuan Wang, Zhuoyun Li, Xinmiao Huang, Xiaowei Huang, Yi Dong

机构 * School of Computer Science and Informatics, University of Liverpool, United Kingdom(利物浦大学计算机科学与信息学学院)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出一种量化评估大语言模型多步结构化推理与人类偏好对齐的方法,引入对齐分数指标,通过构建基于语义熵的矩阵比较模型生成的思维链与人类偏好参考,发现对齐分数在2步推理时达到峰值,支持其作为诊断信号。

Comments Accepted to ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17884 2026-04-21 cs.AI 88%

SPREG: Structured Plan Repair with Entropy-Guided Test-Time Intervention for Large Language Model Reasoning

SPREG:基于熵引导的测试时干预的结构化计划修复用于大语言模型推理

Xuan Wang, Yu Ming, Xinhao Zhong, Xinyu Yu, Wenjie Wang, Shuai Chen, Wei Lin

机构 * Meituan(美团)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 SPREG通过熵引导的测试时干预,解决大语言模型在长链推理中的逻辑幻觉和随机漂移问题,通过动态修复机制提升推理准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05543 2026-04-21 cs.CL cs.SD eess.AS 88%

Closing the Modality Reasoning Gap for Speech Large Language Models

弥合语音大语言模型的模态推理差距

Chaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu, Yan Lu, Jinyu Li, Zhizheng Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Microsoft Corporation(微软公司)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文提出TARS框架,通过不对称奖励设计对齐文本和语音条件轨迹,显著缩小模态推理差距并在7B级语音大语言模型中取得最佳性能。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04509 2026-04-21 cs.CL 88%

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

ErrorRadar:通过错误检测基准测试多模态大语言模型的复杂数学推理能力

Yibo Yan, Shen Wang, Jiahao Huo, Hang Li, Boyan Li, Jiamin Su, Xiong Gao, Yi-Fan Zhang, Tianlong Xu, Zhendong Chu, Aoxiao Zhong, Kun Wang, Hui Xiong, Philip S. Yu, Xuming Hu, Qingsong Wen

机构 * Squirrel AI HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) MSU(密歇根州立大学) UCAS(中国科学技术大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 ErrorRadar通过错误检测任务评估多模态大语言模型在复杂数学推理中的能力,包含2500个高质量K-12数学问题,实验显示GPT-4o仍比人类评价差10%。

Comments Accepted by The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026, Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10681 2026-04-17 cs.CR cs.AI 88%

Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models

Critical-CoT: 一种针对大语言模型推理层后门攻击的稳健防御框架

Vu Tuan Truong, Long Bao Le

机构 * INRS, University of Quebec(INRS大学魁北克分校)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出Critical-CoT,通过双阶段微调提升大语言模型的批判性思维,使其能自动识别潜在后门并拒绝生成恶意推理步骤,有效防御推理层后门攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14525 2026-04-17 cs.AI 88%

Quantifying Cross-Query Contradictions in Multi-Query LLM Reasoning

量化多查询LLM推理中的跨查询矛盾

Rohit Kumar Salla, Ramya Manasa Amancherla, Manoj Saravanan

机构 * Virginia Tech(弗吉尼亚理工大学) Columbia University(哥伦比亚大学)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract,comments);language model(abstract,comments);分类 cs.AI

AI总结 本文研究多查询推理中的逻辑一致性,提出包含390个实例的基准测试,引入集级指标,通过提取承诺、验证全局可满足性及反例引导修复,减少跨查询矛盾并保持单查询准确性。

Comments Accepted at the ICLR 2026 Workshop on Logical Reasoning of Large Language Models. 9 pages, 6 tables, code and data at https://huggingface.co/datasets/rohitspider/cross_query_benchmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13371 2026-04-16 cs.CL 88%

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints

大语言模型在有限离散状态空间问题中复杂性诱导限制的实证证据

Md. Fahad Ullah Utsho, Mohd. Ruhul Ameen, Akif Islam, Md. Golam Rashed, Dipankar Das

机构 * Department of Information and Communication Engineering, University of Rajshahi, Bangladesh(拉贾沙希大学信息与通信工程系,孟加拉国) College of Engineering and Computer Sciences, Marshall University, Huntington, WV, USA(马歇尔大学工程与计算机科学学院,美国,亨廷顿,西弗吉尼亚) Department of Computer Science and Engineering, University of Rajshahi, Bangladesh(拉贾沙希大学计算机科学与工程系,孟加拉国)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文通过构建九个经典推理任务,系统评估大推理模型在逐渐增加的复杂性下的鲁棒性,发现模型在低复杂度时表现良好,但超过特定阈值后显著退化,提出推理崩溃现象。

Comments 45 pages, 36 figures, 7 tables, Journal Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11340 2026-04-16 cs.CL 88%

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models

神经链式推理搜索:寻找最优推理路径以增强大语言模型

Guoming Ling, Zhongzhan Huang, Yupei Lin, Junxin Li, Shanshan Zhong, Hefeng Wu, Liang Lin

机构 * Sun Yat-sen University(中山大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文提出神经链式推理搜索框架,通过动态搜索最优推理策略提升大语言模型性能,实验表明在多个基准测试中准确率提升超3.5%,生成长度减少超22%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19917 2026-04-15 cs.CL 88%

PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models

PILOT:通过内部化潜在优化轨迹进行大语言模型的规划

Haoyu Zheng, Yun Zhu, Yuqian Yuan, Bo Yuan, Wenqiao Zhang, Siliang Tang, Jun Xiao

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 PILOT通过内部化潜在优化轨迹,帮助大语言模型克服长周期任务中的推理不稳定问题,实验显示其在数学和编码基准测试中性能优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09712 2026-04-14 cs.CV cs.AI 88%

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models

LAST:利用工具作为提示以增强多模态大语言模型的空间推理能力

Shi-Yu Tian, Zhi Zhou, Kun-Yang Yu, Ming Yang, Yang Chen, Ziqiao Shang, Lan-Zhe Guo, Yu-Feng Li

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 LAST通过整合专用视觉模型,解决多模态大语言模型在复杂几何布局解析中的幻觉和不精确问题,提出统一框架提升空间推理能力。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02490 2026-04-06 cs.CR cs.AI 88%

Automated Malware Family Classification using Weighted Hierarchical Ensembles of Large Language Models

基于预训练大语言模型加权层次集成的自动化恶意软件家族分类

Samita Bai, Hamed Jelodar, Tochukwu Emmanuel Nwankwo, Parisa Hamedi, Mohammad Meymani, Roozbeh Razavi-Far, Ali A. Ghorbani

机构 * Canadian Institute for Cybersecurity, Faculty of Computer Science, University of New Brunswick(加拿大网络安全研究所,新不伦瑞克大学计算机科学学院)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出基于预训练大语言模型加权层次集成的零标签恶意软件家族分类框架,通过聚合多个互补推理能力的语言模型决策预测,提升分类鲁棒性和分析员推理一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02398 2026-04-06 cs.SE cs.AI 88%

Improving MPI Error Detection and Repair with Large Language Models and Bug References

利用大语言模型和Bug参考改进MPI错误检测与修复

Scott Piersall, Yang Gao, Shenyang Liu, Liqiang Wang

机构 * Dept. of Computer Science, University of Central Florida(中佛罗里达大学计算机科学系)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出结合Few-Shot Learning、Chain-of-Thought和Retrieval Augmented Generation技术,提升大语言模型在MPI错误检测与修复中的准确性,实验结果显示错误检测准确率从44%提升至77%。

Comments 41 pages, 8 figures

Journal ref Journal of Parallel and Distributed Computing, Volume 213, 2026, 105255, ISSN 0743-7315

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02113 2026-04-03 cs.CL 88%

Reliable Control-Point Selection for Steering Reasoning in Large Language Models

用于大语言模型转向推理的可靠控制点选择

Haomin Zhuang, Hojun Yoo, Xiaonan Luo, Kehan Guo, Xiangliang Zhang

机构 * University of Notre Dame(圣母大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文提出稳定过滤方法,通过概率模型识别稳定的行为边界,提升转向向量的稳定性与准确性,在MATH-500数据集上取得0.784准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06663 2026-03-27 cs.CV cs.AI 88%

Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting

图标记:通过基于图的视觉提示提升多模态语言模型的空间推理能力

Giacomo Frisoni, Lorenzo Molfetta, Mattia Buzzoni, Gianluca Moro

机构 * University of Bologna(博洛尼亚大学)

专题命中 推理与问题求解 :language model(title,abstract);prompting(title,abstract);分类 cs.AI

AI总结 本文提出Graph-of-Mark,一种基于图的视觉提示方法,通过在输入图像上叠加场景图来增强多模态语言模型的空间推理能力,实验表明其在视觉问答和定位任务中提升了11个百分点的准确率。

Comments Please cite the definitive, copyrighted, and peer-reviewed version of this article published in AAAI 2026, edited by Sven Koenig et al., AAAI Press, Vol. 40, No. 36, Technical Track, pp. 30726-30734, 2026. DOI: https://doi.org/10.1609/aaai.v40i36.40329

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22784 2026-03-25 cs.LG 88%

Caterpillar of Thoughts: The Optimal Test-Time Algorithm for Large Language Models

思维 caterpillar:大型语言模型的最优测试时间算法

Amir Azarmehr, Soheil Behnezhad, Alma Ghafari

机构 * Northeastern University(东北大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.LG

AI总结 本文提出Caterpillar of Thoughts算法,通过理论分析证明最优算法生成caterpillar树,从而减少生成次数并提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19558 2026-03-23 cs.CL 88%

TextReasoningBench: Does Reasoning Really Improve Text Classification in Large Language Models?

TextReasoningBench: 推理是否真的能提升大语言模型中的文本分类性能?

Xinyu Guo, Yazhou Zhang, Jing Qin

机构 * School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院) School of Nursing, The Hong Kong Polytechnic University(香港理工大学护理学院)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文通过TextReasoningBench评估推理策略在文本分类中的有效性与效率,发现推理并非总能提升性能,且多数策略效率低下。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19515 2026-03-23 cs.AI 88%

ItinBench: Benchmarking Planning Across Multiple Cognitive Dimensions with Large Language Models

ItinBench:利用大语言模型在多个认知维度上进行规划的基准测试

Tianlong Wang, Pinqiao Wang, Weili Shi, Sheng li

机构 * University of Virginia(弗吉尼亚大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出ItinBench,通过整合空间推理和传统语言推理任务,评估大语言模型在多认知维度上的规划能力,发现模型在同时处理多个维度时性能不稳定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19514 2026-03-23 cs.AI 88%

Learning to Disprove: Formal Counterexample Generation with Large Language Models

学习以反驳:利用大语言模型进行形式反例生成

Zenan Li, Zhaoyu Li, Kaiyu Yang, Xiaoxing Ma, Zhendong Su

机构 * ETH Zurich(苏黎世联邦理工学院) University of Toronto(多伦多大学) MiroMind(MiroMind公司) Nanjing University(南京大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文通过微调大语言模型生成形式反例,解决数学推理中反例发现的不足,提出符号突变策略提升训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16434 2026-03-18 cs.AI q-fin.TR 88%

From Natural Language to Executable Option Strategies via Large Language Models

从自然语言到可执行期权策略的大型语言模型

Haochen Luo, Zhengzhao Lai, Junjie Xu, Yifan Li, Tang Pok Hin, Yuan Zhang, Chen Liu

机构 * City University of Hong Kong(香港城市大学) The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)) Shanghai University of Finance and Economics(上海财经大学) University of Science and Technology of China(中国科学技术大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出Option Query Language,通过语法规则将期权市场抽象为高层 primitives,使LLM能作为语义解析器生成可执行期权策略,并通过新数据集验证其优于直接生成方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12633 2026-03-18 cs.CV cs.AI 88%

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

DiG:通过差异 grounding 提升多模态大语言模型的细粒度感知

Zhou Tao, Shida Wang, Yongxiang Hua, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出DiG框架,通过学习相似图像对的差异识别提升多模态大语言模型的细粒度感知能力,实验表明其在多个视觉感知基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00410 2026-03-17 cs.LG 88%

Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models

共奖励:用于大型语言模型中激发推理的稳定自监督强化学习

Zizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao, Zhanke Zhou, Xuan Li, Xiao Feng, Jiangchao Yao, Bo Han

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.LG

AI总结 本文提出Co-rewarding框架,通过引入互补监督提升训练稳定性,通过对比一致性和模型蒸馏方法,在多个数学推理基准上实现性能提升,尤其在Llama-3.2-3B-Instruct上表现突出。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13266 2026-03-17 cs.AI 88%

Multi-hop Reasoning and Retrieval in Embedding Space: Leveraging Large Language Models with Knowledge

嵌入空间中的多跳推理与检索:利用大型语言模型与知识

Lihui Liu

机构 * Wayne State University(韦恩州立大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出EMBRAG框架,通过生成基于知识图谱的逻辑规则并在嵌入空间中推理,提升知识图谱问答任务的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11279 2026-03-16 cs.AI 88%

AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities

AI心理测量学:通过心理测量效度评估大型语言模型的心理推理能力

Yibai Li, Xiaolin Lin, Zhenghui Sha, Zhiye Jin, Xiaobing Li

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文通过心理测量学方法评估四个主流LLM的心理推理能力和整体效度,发现高性能模型在效度上优于其 predecessors,验证了AI心理测量学的应用价值。

Comments Accepted for publication in the Proceedings of the 58th Hawaii International Conference on System Sciences (HICSS), 2025

Journal ref Proceedings of the 58th Hawaii International Conference on System Sciences (HICSS), January 2025, pp. 5189-5197

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12908 2026-03-13 cs.CV cs.AI 88%

DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning

DeepSport: 一种基于代理强化学习的多模态大语言模型,用于通过主动推理实现综合体育视频理解

Junbo Zou, Haotian Xia, Zhen Ye, Shengjie Zhang, Christopher Lai, Vicente Ordonez, Weining Shen, Hanjie Chen

机构 * Georgia Institute of Technology(佐治亚理工学院) Rice University(Rice大学) Johns Hopkins University(约翰霍普金斯大学) University of California, Irvine(加州大学 Irvine分校) University of California, Santa Barbara(加州大学圣巴巴拉分校)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 DeepSport通过代理强化学习实现多运动视频理解,首次端到端训练多任务模型,显著提升性能与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07825 2026-03-10 cs.CL 88%

Benchmarking Large Language Models for Quebec Insurance: From Closed-Book to Retrieval-Augmented Generation

魁北克保险领域大型语言模型基准测试:从闭卷到检索增强生成

David Beauchemin, Richard Khoury

机构 * Group for Research in Artificial Intelligence of Laval University (GRAIL)(拉瓦尔大学人工智能研究组)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文提出AEPC-QA基准测试,评估51个LLM在魁北克保险领域的闭卷和检索增强生成性能,揭示推理、RAG效果及专业化悖论等关键发现。

Comments Publish at the Advances in Financial AI: Towards Agentic and Responsible Systems Workshop @ ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07728 2026-03-10 cs.AI 88%

A Novel Multi-Agent Architecture to Reduce Hallucinations of Large Language Models in Multi-Step Structural Modeling

一种新型多智能体架构以减少大型语言模型在多步骤结构建模中的幻觉

Ziheng Geng, Jiachen Liu, Ran Cao, Lu Cheng, Dan M. Frangopol, Minghui Cheng

机构 * Department of Civil and Architectural Engineering, University of Miami(迈阿密大学土木与建筑工程系) HBC Engineering Company(HBC工程公司) College of Civil Engineering, Hunan University(湖南大学土木学院) Department of Computer Science, University of Illinois Chicago(伊利诺伊大学芝加哥分校计算机科学系) Department of Civil and Environmental Engineering, Lehigh University(莱斯利大学土木与环境工程系) School of Architecture, University of Miami(迈阿密大学建筑学院)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出了一种新型多智能体架构,通过并行代理协同工作,利用OpenSeesPy实现多步骤结构建模自动化,有效减少幻觉并提高计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏