arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-03-10 至 2026-03-10 共收录 443 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 17 篇

2603.07624 2026-03-10 cs.RO 78%

GeoLoco: Leveraging 3D Geometric Priors from Visual Foundation Model for Robust RGB-Only Humanoid Locomotion

GeoLoco: 利用视觉基础模型中的3D几何先验实现稳健的纯RGB人形机器人运动

Yufei Liu, Xieyuanli Chen, Hainan Pan, Chenghao Shi, Yanjie Chen, Kaihong Huang, Zhiwen Zeng, Huimin Lu

专题命中 后训练与偏好优化 :foundation model(title,abstract)

AI总结 GeoLoco通过利用视觉基础模型的3D几何先验,实现纯RGB驱动的人形机器人稳健运动,解决仿真到现实迁移难题。

Comments 8 pages, 6 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05611 2026-03-10 cs.CV 78%

FLARE: Learning Future-Aware Latent Representations from Vision-Language Models for Autonomous Driving

FLARE: 从视觉-语言模型中学习未来感知的潜在表示以实现自动驾驶

Chengen Xie, Chonghao Sima, Tianyu Li, Bin Sun, Junjie Wu, Zhihui Hao, Hongyang Li

机构 * Shanghai Jiaotong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) OpenDriveLab at The University of Hong Kong(香港大学OpenDrive实验室) Li Auto Inc.(Li汽车公司)

专题命中 后训练与偏好优化 :language model(title,abstract)

AI总结 FLARE通过自监督的未来特征预测目标,从大规模未标记数据中学习鲁棒驾驶表示,无需语言监督,提升自动驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08412 2026-03-10 cs.CL cs.AI 73%

Aligning to Illusions: Choice Blindness in Human and AI Feedback

对幻觉的对齐:人类和AI反馈中的选择盲区

Wenbin Wu

专题命中 后训练与偏好优化 :LLM(abstract);RLHF(abstract);分类 cs.CL、cs.AI

AI总结 研究揭示RLHF中偏好构建的问题,显示人类和AI反馈中的选择盲区受上下文影响,导致奖励信号失真。

Comments 16 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08312 2026-03-10 cs.CL 70%

Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder

学习多语句层面的属性表示:统一的语音编码器

Maryem Bouziane, Salima Mdhaffar, Yannick Estève

机构 * LIA - Avignon Université, France(阿维尼翁大学LIA中心,法国)

专题命中 后训练与偏好优化 :foundation model(abstract);post-training(abstract);分类 cs.CL

AI总结 本文提出统一的后训练框架,使语音基础模型能生成多种语句层面的表示,用于多语言语音检索和说话人识别。

Comments Submitted to Interspeech

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07867 2026-03-10 cs.LG cs.AI 62%

Slumbering to Precision: Enhancing Artificial Neural Network Calibration Through Sleep-like Processes

沉睡以精确:通过睡眠样过程增强人工神经网络校准

Jean Erik Delanois, Aditya Ahuja, Giri P. Krishnan, Maxim Bazhenov

机构 * Department of Computer Science \& Engineering, University of California, San Diego, La Jolla, California, USA ARTISAN, Georgia Institute of Technology, Atlanta, Georgia, USA Department of Medicine, University of California, San Diego, La Jolla, California, USA

专题命中 后训练与偏好优化 :post-training(abstract);分类 cs.AI、cs.LG

AI总结 通过睡眠样过程提出SRC方法,有效提升神经网络校准,与温度缩放结合在AlexNet和VGG19上表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07700 2026-03-10 cs.CV cs.AI 57%

TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward

TDM-R1: 通过非可微奖励强化少步扩散模型

Yihong Luo, Tianyang Hu, Weijian Luo, Jing Tang

机构 * Hong Kong University of Science and Technology(香港科技大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) hi-Lab, Xiaohongshu Inc(小红书实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 后训练与偏好优化 :post-training(abstract);分类 cs.AI

AI总结 TDM-R1通过非可微奖励强化少步扩散模型,提升文本到图像生成性能

Comments https://luo-yihong.github.io/TDM-R1-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07366 2026-03-10 cs.CL 57%

RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts

RILEC:检测和生成英语学习者文本中的L1俄语干扰错误

Darya Kharlamova, Irina Proskurina

机构 * Higher School of Economics(俄罗斯高等经济学院) Université Claude Bernard Lyon 1(克莱尔蒙特大学里昂1分校) Université Lumière Lyon 2(里昂2大学路易斯大学) ERIC(教育研究与创新中心)

专题命中 后训练与偏好优化 :language model(abstract);分类 cs.CL

AI总结 RILEC通过生成语言模型和增强方法,有效检测和生成英语学习者文本中的L1俄语干扰错误。

Comments 12 pages, 7 tables, 2 figures. Accepted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06621 2026-03-10 cs.LG 57%

Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models

奖励在攻击下:分析过程奖励模型的鲁棒性和可hack性

Rishabh Tiwari, Aditya Tomar, Udbhav Bamba, Monishwaran Maheswaran, Heng Yang, Michael W. Mahoney, Kurt Keutzer, Amir Gholami

机构 * UC Berkeley(加州大学伯克利分校) Transmute AI ICSI(国际计算机科学研究所) LBNL(劳伦斯伯克利国家实验室)

专题命中 后训练与偏好优化 :LLM(abstract);分类 cs.LG

AI总结 研究揭示了过程奖励模型在对抗性攻击下的脆弱性,指出其更像流畅性检测器而非推理验证器,并提出PRM-BiasBench用于评估模型鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08476 2026-03-10 cs.RO 50%

LAR-MoE: Latent-Aligned Routing for Mixture of Experts in Robotic Imitation Learning

LAR-MoE:用于机器人模仿学习中专家混合的潜在对齐路由

Ariel Rodriguez, Chenpan Li, Lorenzo Mazza, Rayan Younis, Ortrun Hellig, Sebastian Bodenstedt, Martin Wagner, Stefanie Speidel

机构 * Department of Translational Surgical Oncology, NCT/UCC Dresden(转化外科肿瘤学系) Cluster of Excellence-CeTI, TUD, Germany(卓越中心-CeTI) Department of Visceral, Thoracic and Vascular Surgery, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD, Germany(visceral、胸腔和血管外科系)

专题命中 后训练与偏好优化 :post-training(abstract)

AI总结 LAR-MoE通过潜在对齐路由在机器人模仿学习中实现无监督的专家混合,提升任务适应性和效率。

Comments Submitted to iROS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08030 2026-03-10 cs.CV 50%

QualiTeacher: Quality-Conditioned Pseudo-Labeling for Real-World Image Restoration

QualiTeacher: 基于质量条件的伪标签用于现实世界图像恢复

Fengyang Xiao, Jingjia Feng, Peng Hu, Dingming Zhang, Lei Xu, Guanyi Qin, Lu Li, Chunming He, Sina Farsiu

机构 * Duke University Tsinghua University \'Ecole Polytechnique F\'ed\'erale de Lausanne (EPFL) National University of Singapore Sun Yat - sen University or

专题命中 后训练与偏好优化 :preference optimization(abstract)

AI总结 QualiTeacher通过条件化伪标签质量提升图像恢复质量,解决传统方法中伪标签质量与模型泛化能力的矛盾。

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 长上下文与记忆 14 篇

2510.00444 2026-03-10 cs.CL 90%

TokMem: One-Token Procedural Memory for Large Language Models

TokMem: 一个令牌的程序性记忆用于大型语言模型

Zijun Wu, Yongchang Hao, Lili Mou

机构 * Dept. Computing Science & Alberta Machine Intelligence Institute (Amii), University of Alberta(计算科学系及阿尔伯塔机器智能研究所(Amii),阿尔伯塔大学)

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 TokMem通过单个可训练令牌实现程序性记忆,提升大型语言模型在任务回忆和多步函数调用中的性能,同时减少训练参数和上下文开销。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19747 2026-03-10 cs.CL cs.AR 88%

HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture

HaLoRA: 基于混合计算-内存架构的硬件感知低秩适应法用于大型语言模型

Taiqiang Wu, Chenchen Ding, Wenyong Zhou, Yuxin Cheng, Xincheng Feng, Shuqi Wang, Wendong Xu, Chufan Shi, Zhengwu Liu, Ngai Wong

机构 * The University of Hong Kong(香港大学) Tsinghua University(清华大学)

专题命中 长上下文与记忆 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 HaLoRA通过在混合计算-内存架构中部署低秩适应,利用RRAM和SRAM实现高效能和高精度的LLM微调。

Comments 22 pages, Accepted by TODAES (ACM Transactions on Design Automation of Electronic Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07670 2026-03-10 cs.AI 85%

Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers

自主LLM代理的记忆:机制、评估与新兴前沿

Pengfei Du

机构 * Hong Kong Research Institute of Technology(香港理工大学)

专题命中 长上下文与记忆 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文探讨了自主LLM代理中记忆的机制、评估及未来发展方向,分析了五种核心机制及评估挑战,强调记忆对代理适应性和应用的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26354 2026-03-10 cs.AI cs.CL cs.LG 85%

Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents

你的代理可能误进化:自我进化大语言模型代理中的新兴风险

Shuai Shao, Qihan Ren, Chen Qian, Boyi Wei, Dadi Guo, Jingyi Yang, Xinhao Song, Linfeng Zhang, Weinan Zhang, Dongrui Liu, Jing Shao

专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了自我进化代理中因自我进化偏离导致的误进化风险,揭示了其广泛存在及对安全对齐的影响,并提出缓解策略。

Comments Published in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07392 2026-03-10 cs.CL 84%

Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams

大语言模型能否跟上?连续知识流的在线适应基准测试

Jiyeon Kim, Hyunji Lee, Dylan Zhou, Sue Hyun Park, Seunghyun Yoon, Trung Bui, Franck Dernoncourt, Sungmin Cha, Minjoon Seo

机构 * KAIST AI(韩国科学技术院人工智能研究所) UNC Chapel Hill(北卡罗来纳大学教堂山分校) Google(谷歌公司) KRAFTON(KRAFTON公司) Adobe Research(Adobe研究院) New York University(纽约大学)

专题命中 长上下文与记忆 :large language model(title);language model(title);分类 cs.CL

AI总结 本文提出OAKS基准测试,用于评估大语言模型在动态连续知识流中的在线适应能力,发现现有模型在持续更新知识时存在显著局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08274 2026-03-10 cs.CL cs.AI 79%

How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms

在文档问答场景中,大型语言模型会有多大的幻觉?一项跨温度、上下文长度和硬件平台的1720亿token研究

JV Roig

机构 * Kamiwaza AI

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究通过大规模评估揭示了大型语言模型在文档问答中幻觉率与温度、上下文长度及硬件平台的关系,发现模型选择和温度设置显著影响准确性与伪造率。

Comments 18 pages, 12 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03036 2026-03-10 cs.CL cs.LG cs.MA 79%

LatentMem: Customizing Latent Memory for Multi-Agent Systems

LatentMem: 为多智能体系统定制潜在记忆

Muxin Fu, Xiangyuan Xue, Yafu Li, Zefeng He, Siyuan Huang, Xiaoye Qu, Yu Cheng, Yang Yang

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 LatentMem通过定制智能体特定的记忆和优化策略,提升了多智能体系统的性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07997 2026-03-10 cs.AI 77%

CMMR-VLN: Vision-and-Language Navigation via Continual Multimodal Memory Retrieval

CMMR-VLN:通过持续多模态记忆检索实现视觉与语言导航

Haozhou Li, Xiangyu Dong, Huiyan Jiang, Yaoming Zhou, Xiaoguang Ma

机构 * Foshan Graduate School of Innovation at Northeastern University(东北大学创新研究生院) Faculty of Robot Science and Engineering at Northeastern University(东北大学机器人科学与工程学院) College of Software at Northeastern University(东北大学软件学院) School of Aeronautic Science and Engineering at Beihang University(北航航空科学与工程学院)

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 CMMR-VLN通过引入持续多模态记忆检索机制,提升视觉与语言导航任务中对先前经验的选择性利用能力,显著提高导航成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09486 2026-03-10 cs.CV cs.AI cs.MM 77%

Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding

视频-EM:面向长视频理解的事件中心型片段记忆

Yun Wang, Long Zhang, Jingren Liu, Jiaqi Yan, Zhanjie Zhang, Jiahao Zheng, Ao Ma, Run Ling, Xun Yang, Dapeng Wu, Xiangyu Chen, Xuelong Li

机构 * City University of Hong Kong(香港城市大学) University of Science and Technology of China(中国科学技术大学) Tianjin University(天津大学) Nanjing University(南京大学) Zhejiang University(浙江大学) The Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所(TeleAI))

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Video-EM通过事件中心型片段记忆框架,将长视频问答转化为事件构建与记忆细化,提升长视频理解的连贯性与可靠性。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07375 2026-03-10 eess.SY cs.SY 75%

Multi-Agentic AI for Conflict-Aware rApp Policy Orchestration in Open RAN

多智能体AI用于冲突感知的rApp策略编排在开放RAN中

Haiyuan Li, Yulei Wu, Dimitra Simeonidou

专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出多智能体AI框架,用于实现开放RAN中冲突感知的rApp策略自动生成与编排,实验表明其在部署准确性和推理成本上有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06918 2026-03-10 cs.RO 75%

T2Nav Algebraic Topology Aware Temporal Graph Memory and Loop Detection for ZeroShot Visual Navigation

T2Nav 基于代数拓扑的时序图记忆与循环检测用于零样本视觉导航

Quang-Anh N. D., Duc Pham, Minh-Anh Nguyen, Tung Doan, Tuan Dang

机构 * International School, Vietnam National University, Hanoi, Vietnam(越南国家大学河内国际学校) Hanoi University of Science, Vietnam National University, Hanoi, Vietnam(越南国家大学河内科学大学) Hanoi University of Science and Technology, Hanoi, Vietnam(河内科学技术大学) Cognitive Robotics Lab, Department of Electrical Engineering and Computer Science, University of Arkansas, Fayetteville, AR, USA(美国阿肯色大学富尔顿分校电气工程与计算机科学系认知机器人实验室) University of Arkansas, Fayetteville, AR, USA(美国阿肯色大学富尔顿分校)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 T2Nav通过整合异构数据和基于图的推理,实现零样本视觉导航,具备高效探索和路径规划能力,适用于视觉相似但空间不同的目标实例导航。

Journal ref EEE International Conference on Robotics & Automation 2026 (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07023 2026-03-10 cs.CL cs.AI 73%

Hit-RAG: Learning to Reason with Long Contexts via Preference Alignment

通过偏好对齐学习长上下文的推理

Junming Liu, Yuqi Li, Shiping Wen, Zhigang Zeng, Tingwen Huang

机构 * Tongji University(同济大学) The City University of New York(纽约城市大学) University of Technology Sydney(悉尼大学) Huazhong University of Science and Technology(华中科技大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 Hit-RAG通过多阶段偏好对齐框架解决长上下文推理中的注意力稀释和幻觉问题,提升模型在长上下文场景下的推理能力。

Comments 21 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23932 2026-03-10 cs.CL 70%

SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving

SwingArena: 用于长上下文GitHub问题解决的竞争性编程竞技场

Wendong Xu, Jing Xiong, Chenyang Zhao, Qiujiang Chen, Haoran Wang, Hui Shen, Zhongwei Wan, Jianbo Dai, Taiqiang Wu, He Xiao, Chaofan Tao, Z. Morley Mao, Ying Sheng, Zhijiang Guo, Hongxia Yang, Bei Yu, Lingpeng Kong, Quanquan Gu, Ngai Wong

机构 * The University of Hong Kong(香港大学) University of California Los Angeles(加州大学洛杉矶分校) Tsinghua University(清华大学) University of Michigan Ann Arbor(密歇根大学安娜堡分校) The Ohio State University(俄亥俄州立大学) University of Edinburgh(爱丁堡大学) The Chinese University of Hong Kong(香港中文大学) The Hong Kong Polytechnic University(香港理工大学) Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) LMSYS Org(LMSYS组织)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 SwingArena是一个用于评估LLM在真实CI驱动软件开发环境中性能的竞争性编程竞技场,通过模拟软件迭代过程,结合检索增强型代码生成模块,实现对长上下文问题的高效处理。

Comments The paper has been accepted as an oral presentation at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07024 2026-03-10 cs.AI 70%

Enhancing Web Agents with a Hierarchical Memory Tree

通过分层记忆树增强网络代理

Yunteng Tan, Zhi Gao, Xinxiao Wu

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出分层记忆树HMT,通过结构化框架分离逻辑规划与操作执行,提升网络代理在跨网站和跨领域任务中的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 推理与问题求解 91 篇

2603.07091 2026-03-10 cs.SE 90%

Exploring the Reasoning Depth of Small Language Models in Software Architecture: A Multidimensional Evaluation Framework Towards Software Engineering 2.0

探索小型语言模型在软件架构中的推理深度:一项面向软件工程2.0的多维评估框架

Ha Vo, Nhut Tran, Khang Vo, Phat T. Tran-Truong, Son Ha

专题命中 推理与问题求解 :language model(title,abstract);small language model(title,abstract);large language model(abstract);prompting(abstract)

AI总结 本研究通过多维评估框架评估小型语言模型在软件架构中的推理深度,发现参数阈值影响推理能力,揭示了中等大小模型在短上下文下的校准机制及语义多样性与幻觉的相关性。

Comments Accepted at ICSA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08398 2026-03-10 cs.CL cs.AI cs.LG 89%

Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective

揭示大型语言模型中的行为可塑性:基于令牌条件的视角

Liyuan Mao, Le Yu, Jing Zhou, Chujie Zheng, Bowen Yu, Chang Gao, Shixuan Liu, An Yang, Weinan Zhang, JunYang Lin

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出ToCoRL框架,通过令牌条件强化学习使LLM在推理时实现稳定的行为适应,从而实现精确的行为控制。

Comments Work done during an internship at the Qwen Team, Alibaba Group

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08077 2026-03-10 cs.IR 89%

Why Large Language Models can Secretly Outperform Embedding Similarity in Information Retrieval

为何大语言模型可以秘密超越嵌入相似性在信息检索中的表现

Matei Benescu, Ivo Pascal de Jong

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文探讨了大语言模型在信息检索中通过推理超越传统嵌入相似性方法的潜力,并指出当前注释数据集难以准确评估其效果。

Comments 13 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07989 2026-03-10 cs.CV 89%

AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Models

AutoTraces:通过多模态大语言模型实现自回归轨迹预测

Teng Wang, Yanting Lu, Ruize Wang

机构 * Southeast University(东南大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 AutoTraces通过多模态大语言模型实现自回归轨迹预测,采用创新的轨迹标记方案和自动链式推理机制,提升长周期预测精度与跨场景泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07825 2026-03-10 cs.CL 88%

Benchmarking Large Language Models for Quebec Insurance: From Closed-Book to Retrieval-Augmented Generation

魁北克保险领域大型语言模型基准测试:从闭卷到检索增强生成

David Beauchemin, Richard Khoury

机构 * Group for Research in Artificial Intelligence of Laval University (GRAIL)(拉瓦尔大学人工智能研究组)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文提出AEPC-QA基准测试,评估51个LLM在魁北克保险领域的闭卷和检索增强生成性能,揭示推理、RAG效果及专业化悖论等关键发现。

Comments Publish at the Advances in Financial AI: Towards Agentic and Responsible Systems Workshop @ ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07728 2026-03-10 cs.AI 88%

A Novel Multi-Agent Architecture to Reduce Hallucinations of Large Language Models in Multi-Step Structural Modeling

一种新型多智能体架构以减少大型语言模型在多步骤结构建模中的幻觉

Ziheng Geng, Jiachen Liu, Ran Cao, Lu Cheng, Dan M. Frangopol, Minghui Cheng

机构 * Department of Civil and Architectural Engineering, University of Miami(迈阿密大学土木与建筑工程系) HBC Engineering Company(HBC工程公司) College of Civil Engineering, Hunan University(湖南大学土木学院) Department of Computer Science, University of Illinois Chicago(伊利诺伊大学芝加哥分校计算机科学系) Department of Civil and Environmental Engineering, Lehigh University(莱斯利大学土木与环境工程系) School of Architecture, University of Miami(迈阿密大学建筑学院)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出了一种新型多智能体架构,通过并行代理协同工作,利用OpenSeesPy实现多步骤结构建模自动化,有效减少幻觉并提高计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏