arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-02-26 至 2026-02-26 共收录 204 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 43 篇

2602.21681 2026-02-26 cs.SE 78%

AkiraRust: Re-thinking LLM-aided Rust Repair Using a Feedback-guided Thinking Switch

AkiraRust:重新思考利用反馈引导的思考开关的LLM辅助Rust修复

Renshuang Jiang, Yichong Wang, Pan Dong, Xiaoxiang Fang, Zhenling Duan, Tinglue Wang, Yuchen Hu, Jie Yu, Zhe Jiang

专题命中 推理与问题求解 :LLM(title,abstract)

AI总结 AkiraRust通过反馈引导的思考开关机制,实现Rust修复的语义正确性和效率提升。

Comments 7 pages, 11 figures, accepted to DAC

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21763 2026-02-26 cs.CL 77%

Improving Implicit Discourse Relation Recognition with Natural Language Explanations from LLMs

通过LLMs生成的自然语言解释提升隐含话语关系识别

Heng Wang, Changxing Wu

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文通过利用LLMs生成的自然语言解释提升隐含话语关系识别的性能和可解释性。

Comments AAAI26'0ral

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14903 2026-02-26 cs.AI 77%

The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics

链式推理在推理中的潜力:对轨迹动态的深入考察

Gregor Bachmann, Yichen Jiang, Seyed Mohsen Moosavi Dezfooli, Moin Nabi

机构 * Apple(苹果公司)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.AI

AI总结 本文研究了链式推理(CoT)在推理中的潜力,通过分析轨迹动态,发现CoT中不同部分对最终答案的贡献,并探讨了CoT的可转移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13477 2026-02-26 cs.AI 77%

OMNI-LEAK: Orchestrator Multi-Agent Network Induced Data Leakage

OMNI-LEAK:协调多智能体网络引发的数据泄露

Akshat Naik, Jay Culligan, Yarin Gal, Philip Torr, Rahaf Aljundi, Alasdair Paren, Adel Bibi

机构 * University of Oxford(牛津大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 OMNI-LEAK攻击通过多智能体协调设置中的单个提示注入泄露敏感数据,揭示了多代理系统在数据安全方面的关键漏洞。

Comments Preprint; corrected typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21317 2026-02-26 cs.LG 77%

Shared Nature, Unique Nurture: PRISM for Pluralistic Reasoning via In-context Structure Modeling

共享本质,独特培养:通过上下文结构建模实现多元推理的PRISM

Guancheng Tu, Shiyang Zhang, Tianyu Zhang, Yi Zhang, Diji Yang

机构 * University of Pennsylvania(宾夕法尼亚大学) Yale University(耶鲁大学) University of California Santa Cruz(加州大学圣克鲁兹分校)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 PRISM通过上下文结构建模实现多元推理,提升创造力和科学发现能力,展现独特认知轨迹的多元AI新范式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19439 2026-02-26 cs.AI cs.LG math.OC 76%

OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents

OptiRepair:基于LLM代理的闭环供应链优化模型诊断与修复

Ruicheng Ao, David Simchi-Levi, Xinshang Wang

机构 * Massachusetts Institute of Technology(麻省理工学院) Alibaba Group(阿里巴巴集团)

专题命中 推理与问题求解 :LLM(title);分类 cs.AI、cs.LG

AI总结 OptiRepair通过训练LLM代理实现供应链优化模型的闭环诊断与修复,显著提升修复效率和合理性,解决求解器交互和运营合理性两大难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21604 2026-02-26 cs.DB 75%

Towards Autonomous Graph Data Analytics with Analytics-Augmented Generation

迈向具有分析增强生成的自主图数据分析

Qiange Wang, Chaoyi Chen, Jingqi Gao, Zihan Wang, Yanfeng Zhang, Ge Yu

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出分析增强生成(AAG)范式,通过知识驱动的任务规划和算法中心的LLM-分析交互,实现端到端的图分析流水线,将自然语言意图转化为自动执行和可解释结果。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01350 2026-02-26 cs.AI 74%

Error Notebook-Guided, Training-Free Part Retrieval in 3D CAD Assemblies via Vision-Language Models

通过视觉-语言模型实现无训练的3DCAD装配体部件检索:基于错误笔记本的引导

Yunqing Liu, Nan Zhang, Zhiming Tan

机构 * Fujitsu Research & Development Center Shanghai, China(富士通研发中心上海)

专题命中 推理与问题求解 :language model(title);分类 cs.AI

AI总结 本研究提出一种无训练的3D CAD部件检索方法,通过结合错误笔记本与RAG技术,显著提升基于VLM的推理性能,实验显示在专有模型和开源模型上均取得显著效果。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05154 2026-02-26 cs.CL cs.AI cs.IR 73%

Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement

通过参数知识强化来抵御RAG中的上下文干扰

Chenyu Lin, Yilin Wen, Du Su, Hexiang Tan, Fei Sun, Muhan Chen, Chenfu Bao, Zhonghou Lyu

机构 * Baidu Inc.(百度公司) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 Knowledgeable-R1通过参数知识强化提升RAG在知识冲突和一般场景中的鲁棒性与推理准确性,反事实场景中表现优于现有基线。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21269 2026-02-26 cs.LG cs.AI stat.ML 73%

Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space

群体正交化策略优化:将群体策略优化视为希尔伯特空间中的正交投影

Wang Zixian

机构 * China Mobile Communications Group Shandong Co., Ltd. Tai’an Branch(中国移动通信集团山东公司泰安分公司)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 GOPO通过正交投影在希尔伯特空间中优化群体策略,实现精确稀疏性和稳定梯度动态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12415 2026-02-26 cs.LG cs.AI 73%

Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space

正交化策略优化:作为希尔伯特空间中的正交投影的策略优化

Wang Zixian

机构 * China Mobile Communications Group Shandong Co., Ltd. Tai’an Branch(中国移动通信集团山东公司泰安分支)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 OPO通过正交投影在希尔伯特空间中实现策略优化,利用守恒定律而非方差减少启发式方法,提升模型在长时程训练和分布外泛化中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10947 2026-02-26 cs.AI cs.LG 73%

Spurious Rewards: Rethinking Training Signals in RLVR

虚假奖励:重新思考强化学习中的训练信号

Rulin Shao, Shuyue Stella Li, Rui Xin, Scott Geng, Yiping Wang, Sewoong Oh, Simon Shaolei Du, Nathan Lambert, Sewon Min, Ranjay Krishna, Yulia Tsvetkov, Hannaneh Hajishirzi, Pang Wei Koh, Luke Zettlemoyer

机构 * University of Washington, Seattle, WA, USA(华盛顿大学) Allen Institute for Artificial Intelligence, Seattle, WA, USA(人工智能研究院) University of California, Berkeley, Berkeley, CA, USA(加州大学伯克利分校)

专题命中 推理与问题求解 :language model(abstract);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 该研究发现,即使使用随机分配的虚假奖励,强化学习方法也能在某些模型上显著提升性能,凸显了验证训练信号的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11684 2026-02-26 cs.CL cs.AI 73%

MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task

MathFimer: 通过填空中间任务扩展推理步骤以增强数学推理

Yuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang, Xin Xu, Mengdi Zhang, Jian Shao, Yueting Zhuang

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 MathFimer通过填空中间任务扩展推理步骤,提升大语言模型的数学推理能力,无需依赖强大外部模型或昂贵计算资源。

Comments ICLR 2026: https://openreview.net/forum?id=14i2wzPPfn

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21706 2026-02-26 cs.CV cs.AI 70%

SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video

SurGo-R1:手术视频中操作区的上下文推理基准测试与建模

Guanyi Qin, Xiaozhen Wang, Zhu Zhuo, Chang Han Low, Yuancan Xiao, Yibing Fu, Haofeng Liu, Kai Wang, Chunjiang Li, Yueming Jin

机构 * National University of Singapore, Singapore(新加坡国立大学) Southern Medical University, China(南方医科大学) Guangzhou Research Translation and Innovation Institute, National University of Singapore, China(广州研究翻译与创新研究院,新加坡国立大学,中国)

专题命中 推理与问题求解 :language model(abstract);RLHF(abstract);分类 cs.AI

AI总结 SurGo-R1通过RLHF优化的多阶段架构,在手术视频中实现了高精度的操作区识别,显著优于通用视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21619 2026-02-26 cs.CL 70%

When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning

更多未必更好:对视觉空间推理中空间与常识信息的系统分析

Muku Akasaka, Soyeon Caren Han

机构 * The University of Melbourne(墨尔本大学)

专题命中 推理与问题求解 :language model(abstract);prompting(abstract);分类 cs.CL

AI总结 本文通过系统分析发现,过多信息未必提升视觉空间推理性能,针对性单空间提示和精确空间定位可提高准确性。

Comments 5 pages, 6 figures, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21371 2026-02-26 cs.LG 70%

Interleaved Head Attention

交错头部注意力

Sai Surya Duvvuri, Chanakya Ekbote, Rachit Bansal, Rishabh Tiwari, Devvrit Khatri, David Brandfonbrener, Paul Liang, Inderjit Dhillon, Manzil Zaheer

机构 * Meta UT Austin(德克萨斯大学) UC Berkeley(伯克利大学) Harvard University(哈佛大学) MIT(麻省理工学院)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 交错头部注意力通过构造伪头实现跨头混合,提升多步推理效率,在多项式任务和顺序敏感任务中参数效率提高,实测在RULER和OpenThoughts上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21857 2026-02-26 cs.AI cs.CL cs.LG 67%

Distill and Align Decomposition for Enhanced Claim Verification

分解与对齐:提升声明验证的Distill和Align分解

Jabez Magomere, Elena Kochkina, Samuel Mensah, Simerjot Kaur, Fernando Acero, Arturo Oncevay, Charese H. Smiley, Xiaomo Liu, Manuela Veloso

专题命中 推理与问题求解 :language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种基于强化学习的分解与对齐方法,通过优化分解质量和验证器对齐,提升声明验证性能,达到71.75%的宏F1,优于现有方法。

Comments EACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06034 2026-02-26 cs.CV 67%

V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval

V-Retrver: 以证据驱动的代理推理用于通用多模态检索

Dongyang Chen, Chaoyang Wang, Dezhao Su, Xi Xiao, Zeyu Zhang, Jing Xiong, Qing Li, Yuzhang Shang, Shichao Kan

机构 * Tsinghua University(清华大学) University of Central Florida(中央佛罗里达大学) Fudan University(复旦大学) The Australian National University(澳大利亚国立大学) The University of Hong Kong(香港大学) Pengcheng Laboratory(鹏城实验室) Central South University(中南大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

AI总结 V-Retrver通过引入证据驱动的代理推理机制,提升多模态检索的准确率和推理可靠性。

Comments Project page: https://github.com/chendy25/V-Retrver

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18292 2026-02-26 cs.LG cs.AI 62%

Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers

解码作为概率单纯形上的优化:从Top-K到Top-P(核)到Best-of-K采样器

Xiaotong Ji, Rasul Tutunov, Matthieu Zimmer, Haitham Bou-Ammar

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于概率单纯形优化的解码框架,统一了Top-K、Top-P等采样方法,并引入Best-of-K采样器提升多样本流程性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21226 2026-02-26 cs.CL cs.AI 62%

IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions

伊斯兰法律基准:评估LLMs在1200年伊斯兰多元法律传统中的知识和推理能力

Ezieddin Elmahjub, Junaid Qadir, Abdullah Mushtaq, Rafay Naeem, Ibrahim Ghaznavi, Waleed Iqbal

机构 * Qatar University(卡塔尔大学) Information Technology University(信息技术大学) Northeastern University(东北大学) Queen Mary University of London(伦敦女王学院)

专题命中 推理与问题求解 :prompting(abstract);分类 cs.CL、cs.AI

AI总结 伊斯兰法律基准评估LLMs在伊斯兰法推理中的能力,揭示其在复杂任务中存在显著缺陷,强调基础知识缺失对AI性能的影响。

Comments This manuscript has been submitted for review to Artificial Intelligence \& Law

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13793 2026-02-26 cs.AI 57%

Med-REFL: Medical Reasoning Enhancement via Self-Corrected Fine-grained Reflection

Med-REFL: 通过自我校正的细粒度反思增强医疗推理

Zongxian Yang, Jiayu Qian, Zegao Peng, Haoyu Zhang, Yu-An Huang, KC Tan, Zhi-An Huang

机构 * City University of Hong Kong (Dongguan), China(香港城市大学(东莞)) Northwestern Polytechnical University, China(西北工业大学) The Hong Kong Polytechnic University, China(香港理工大学)

专题命中 推理与问题求解 :preference optimization(abstract);分类 cs.AI

AI总结 Med-REFL通过自我校正的细粒度反思提升医疗推理,有效解决验证瓶颈,提升模型性能并拓展至其他领域。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21650 2026-02-26 cs.SI cs.AI 57%

PPCR-IM: A System for Multi-layer DAG-based Public Policy Consequence Reasoning and Social Indicator Mapping

PPCR-IM:一个多层DAG基础的公共政策后果推理与社会指标映射系统

Zichen Song, Weijia Li

机构 * Lanzhou University(兰州大学)

专题命中 推理与问题求解 :LLM(abstract);分类 cs.AI

AI总结 PPCR-IM通过多层DAG推理和指标映射,实现公共政策后果的结构化分析与社会影响评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19983 2026-02-26 cs.RO cs.AI 57%

Contextual Safety Reasoning and Grounding for Open-World Robots

面向开放世界机器人的上下文安全推理与 grounding

Zachary Ravichandran, David Snyder, Alexander Robey, Hamed Hassani, Vijay Kumar, George J. Pappas

机构 * University of Pennsylvania(宾夕法尼亚大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理与问题求解 :language model(abstract);分类 cs.AI

AI总结 CORE框架通过视觉语言模型实现在线上下文推理与空间grounding,提供概率安全保证,在开放世界中有效执行上下文适应的安全行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21983 2026-02-26 cs.RO 50%

Humanizing Robot Gaze Shifts: A Framework for Natural Gaze Shifts in Humanoid Robots

使机器人目光转移人性化:一种在人形机器人中实现自然目光转移的框架

Jingchao Wei, Jingkai Qin, Yuxiao Cao, Jingcheng Huang, Xiangrui Zeng, Min Li, Zhouping Yin

机构 * School of Mechanical Science and Engineering, Huazhong University of Science and Technology(机械科学与工程学院,华中科技大学)

专题命中 推理与问题求解 :language model(abstract)

AI总结 本文提出RGS框架,通过结合认知注意力机制与生物模仿运动生成,实现人形机器人自然且上下文合适的人类化目光转移。

Comments submitted to AIM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21952 2026-02-26 cs.CV 50%

MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving

MindDriver: 引入渐进多模态推理用于自动驾驶

Lingjun Zhang, Yujian Yuan, Changjie Wu, Xinyuan Chang, Xin Cai, Shuang Zeng, Linzhe Shi, Sijin Wang, Hang Zhang, Mu Xu

机构 * Amap, Alibaba Group(阿里集团蚂巴公司) The Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学) Xi’an Jiaotong University(西安交通大学)

专题命中 推理与问题求解 :language model(abstract)

AI总结 MindDriver通过渐进多模态推理框架提升自动驾驶系统的推理能力,实现语义到物理空间的转化与轨迹规划,展现优异的性能表现。

Comments CVPR2026; Yujian Yuan and Lingjun Zhang contributed equally with random order

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21552 2026-02-26 cs.CV 50%

Generalizing Visual Geometry Priors to Sparse Gaussian Occupancy Prediction

将视觉几何先验泛化到稀疏高斯占用预测

Changqing Zhou, Yueru Luo, Changhao Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 推理与问题求解 :foundation model(abstract)

AI总结 GPOcc通过利用可泛化的视觉几何先验进行单目占用预测,在速度和精度上均优于现有方法。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20685 2026-02-26 cs.CV 50%

RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space

RAYNOVA:在射线空间中实现尺度-时间自回归世界建模

Yichen Xie, Chensheng Peng, Mazen Abdelfattah, Yihan Hu, Jiezhi Yang, Eric Higgins, Ryan Brigden, Masayoshi Tomizuka, Wei Zhan

机构 * Applied Intuition UC Berkeley(加州大学伯克利分校)

专题命中 推理与问题求解 :foundation model(abstract)

AI总结 RAYNOVA通过双因果自回归框架实现尺度-时间自回归世界建模,在驾驶场景中实现多视角视频生成,具有更高的吞吐量和可控性。

Comments Accepted by CVPR 2026; Project website: https://raynova-ai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21087 2026-02-26 cs.CV 50%

MIRA: Multimodal Iterative Reasoning Agent for Image Editing

MIRA: 多模态迭代推理代理用于图像编辑

Ziyun Zeng, Hang Hua, Jiebo Luo

机构 * University of Rochester(罗切斯特大学) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

专题命中 推理与问题求解 :SFT(abstract)

AI总结 MIRA通过多模态迭代推理代理提升图像编辑的语义一致性和感知质量,结合自研数据集和训练流程,实现与专有系统相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10855 2026-02-26 eess.IV cs.CV 50%

Transformer-based cardiac substructure segmentation from contrast and non-contrast computed tomography for radiotherapy planning

基于变换器的心脏亚结构从对比和非对比计算机断层扫描进行放疗规划

Aneesh Rangnekar, Nikhil Mankuzhy, Jonas Willmann, Chloe Min Seo Choi, Abraham Wu, Maria Thor, Andreas Rimner, Harini Veeraraghavan

专题命中 推理与问题求解 :pretraining(abstract)

AI总结 本研究提出基于变换器的SMIT模型,通过平衡课程学习实现高效训练,准确分割心脏亚结构,提升放疗规划的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 评测与基准 39 篇

2602.21833 2026-02-26 cs.SE 90%

From Restructuring to Stabilization: A Large-Scale Experiment on Iterative Code Readability Refactoring with Large Language Models

从重构到稳定:基于大语言模型的迭代代码可读性重构大规模实验

Norman Peitek, Julia Hess, Sven Apel

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 本文通过大规模实验研究大语言模型在迭代代码重构中的能力,揭示了重构过程的收敛趋势及可读性提升的机制。

详情

展开后加载摘要…

URL PDF HTML 收藏