arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-03-04 至 2026-03-04 共收录 18 信号源:cs.CL, cs.AI, cs.LG

1. 复杂问题求解 18 篇

2506.17871 2026-03-04 cs.CL cs.AI cs.LG 80%

LLM Probability Concentration: How Alignment Shrinks the Generative Horizon

LLM概率集中:对齐如何缩小生成范围

Chenghao Yang, Sida Li, Ari Holtzman

专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现对齐微调通过减少生成多样性,使LLM生成更一致,从而影响复杂推理稳定性。

Comments Codebase: https://github.com/yangalan123/LLMBranchingFactor. V3: Significantly rewrite the whole paper for a clearer structure. Correct problems in the theory parts (Remove emphasis on AEP, discussions on variable LLM generation lengths) and strengthen asymptotic analysis. Add Qwen and OLMo2 experiments. Preliminary SFT v.s. RL comparison to better understand the alignment effects on BF

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03005 2026-03-04 cs.AI 79%

OrchMAS: Orchestrated Reasoning with Multi Collaborative Heterogeneous Scientific Expert Structured Agents

OrchMAS: 基于多协作异构科学专家结构代理的协同推理

Yichao Feng, Haoran Luo, Zhenghong Lin, Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh, Anh Tuan Luu

机构 * Magellan Technology Research Institute (MTRI)(Magellan技术研究 institute) Nanyang Technological University(南洋理工大学)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI

AI总结 OrchMAS通过双层多模型编排框架,实现异构科学专家代理的动态协作,提升复杂科学推理的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02266 2026-03-04 cs.SD cs.AI eess.AS 79%

When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning

当扩展失效时:通过多步骤感知-aware 推理缓解LALMs的音频感知衰减

Ruixiang Mao, Xiangnan Ma, Dan Chen, Ziming Zhu, Yuan Ge, Aokai Hao, Haishu Zhao, Yifu Huo, Qing Yang, Kaiyan Chang, Xiaoqian Liu, Chenglong Wang, Qiaozhi He, Tong Xiao, Jingbo Zhu

机构 * Northeastern University,China(东北大学) NiuTrans Research(NiuTrans研究院)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出MPAR²范式,通过多步骤感知-aware推理缓解LALMs在推理过程中因音频感知衰减导致的性能下降问题。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02438 2026-03-04 cs.CV 78%

ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering

ORCA:用于文档视觉问答的协作代理协同推理

Aymen Lassoued, Mohamed Ali Souibgui, Yousri Kessentini

机构 * Digital Research Center of Sfax, SMARTS Laboratory(突尼斯斯法克斯数字研究中心,SMARTS实验室) École Polytechnique de Tunisie, University of Carthage(突尼斯理工学院,卡塔赫纳大学) Computer Vision Center, Universitat Autònoma de Barcelona(巴塞罗那自治大学计算机视觉中心)

专题命中 复杂问题求解 :reasoning(title,abstract)

AI总结 ORCA通过协作代理协同推理提升文档视觉问答的复杂推理与多步骤工作流处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02939 2026-03-04 cs.AI 77%

ShipTraj-R1: Reinforcing Ship Trajectory Prediction in Large Language Models via Group Relative Policy Optimization

ShipTraj-R1: 通过群体相对策略优化在大型语言模型中强化船舶轨迹预测

Yang Zhan, Yunhao Li, Zhang Chao, Yuxu Lu, Yan Li

机构 * School of Artificial Intelligence, Optics, and Electronics (iOPEN)(人工智能、光学与电子学院) Northwestern Polytechnical University(西北工业大学) Department of Logistics and Maritime Studies(物流与海洋研究系) The Hong Kong Polytechnic University(香港理工大学) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing(测绘遥感信息工程国家重点实验室) Wuhan University(武汉大学)

专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

AI总结 ShipTraj-R1通过群体相对策略优化在大型语言模型中强化船舶轨迹预测,利用动态提示和规则奖励机制提升预测准确性。

Comments Accepted by the 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02640 2026-03-04 cs.CY cs.AI cs.CL cs.MA cs.SI 76%

Credibility Governance: A Social Mechanism for Collective Self-Correction under Weak Truth Signals

可信治理:在弱真相信号下的一种社会机制,用于集体自我校正

Wanying He, Yanxi Lin, Ziheng Zhou, Xue Feng, Min Peng, Qianqian Xie, Zilong Zheng, Yipeng Kang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Tsinghua University(清华大学) University of California, Los Angeles(加州大学洛杉矶分校) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)

专题命中 复杂问题求解 :self-correction(title);分类 cs.CL、cs.AI

AI总结 可信治理通过动态可信度评分和可信度加权背书,提升集体自我校正能力,减少虚假信息影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03236 2026-03-04 cs.CY 71%

Conversational Learning Diagnosis via Reasoning Multi-Turn Interactive Learning

通过推理的多轮交互学习诊断

Fangzhou Yao, Sheng Chang, Weibo Gao, Qi Liu

专题命中 复杂问题求解 :reasoning(title)

AI总结 本文提出ParLD框架,通过多智能体协作实现对话学习中的认知诊断,提升诊断的可靠性和洞察力。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00582 2026-03-04 cs.CL 70%

Super Research: Answering Highly Complex Questions with Large Language Models through Super Deep and Super Wide Research

超研究:通过超深和超宽研究利用大型语言模型解决高度复杂的问题

Yubo Dong, Nianhao You, Yuxuan Hou, Zixun Sun, Yue Zhang, Liang Zhang, Siyuan Zhao, Hehe Fan

机构 * College of Computer Science and Technology(计算机科学与技术学院) Ant Group(蚂蚁集团)

专题命中 复杂问题求解 :reasoning(abstract);planning(abstract);分类 cs.CL

AI总结 通过超深和超宽研究,利用大型语言模型解决高度复杂问题,提供可验证的报告和审计协议,评估模型在复杂研究任务中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03101 2026-03-04 cs.LG cs.AI cs.CV 62%

ALARM: Automated MLLM-Based Anomaly Detection in Complex-EnviRonment Monitoring with Uncertainty Quantification

ALARM: 基于多模态大语言模型的复杂环境监控异常检测与不确定性量化

Congjing Zhang, Feng Lin, Xinyi Zhao, Pei Guo, Wei Li, Lin Chen, Chaoyue Zhao, Shuai Huang

机构 * Department of Industrial and Systems Engineering, University of Washington(华盛顿大学工业与系统工程系) Wyze Labs, Inc.(Wyze实验室)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 ALARM通过结合不确定性量化与多模态大语言模型,实现了复杂环境中的异常检测,展现出在不同领域中的高准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03078 2026-03-04 cs.AI 57%

RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization

RAPO:通过检索增强策略优化扩展LLM代理的探索

Siwei Zhang, Yun Xiong, Xi Chen, Zi'an Jia, Renhong Huang, Jiarong Xu, Jiawei Zhang

机构 * Fudan University(复旦大学) Zhejiang University(浙江大学) UC Davis(加州大学戴维斯分校)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

AI总结 RAPO通过引入检索增强策略优化,扩展LLM代理的探索能力,提升训练效率和探索效果。

Comments Submit to KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03024 2026-03-04 cs.RO cs.AI 57%

MA-CoNav: A Master-Slave Multi-Agent Framework with Hierarchical Collaboration and Dual-Level Reflection for Long-Horizon Embodied VLN

MA-CoNav:一种具有分层协作和双阶段反思的主从多智能体框架用于长周期具身视觉语言导航

Ling Luo, Qianqian Bai

机构 * Southwestern University of Finance(西南财经大学) Department of Informatics,Universitat Hamburg, Hamburg, Germany(汉堡大学信息学院) University of Electronic Science(电子科学大学)

专题命中 复杂问题求解 :planning(abstract);分类 cs.AI

AI总结 MA-CoNav 通过主从分层架构和双阶段反思机制,提升长周期具身视觉语言导航任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16894 2026-03-04 cs.CL 57%

Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment

让LoRA更出色:通过自适应奇异值和专家混合优化对齐提升LoRA

Chenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu, Xiaoye Qu, Wei Wei, Yu Cheng

机构 * School of Computer Science \& Technology, Huazhong University of Science The Chinese University of Hong Kong Zhejiang University

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL

AI总结 GOAT通过自适应奇异值和专家混合优化对齐提升LoRA性能,实现在多个任务上的最优表现。

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02615 2026-03-04 cs.CL 57%

Think, But Don't Overthink: Reproducing Recursive Language Models

思考,但不要过度思考:重现递归语言模型

Daren Wang

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL

AI总结 本研究通过评估不同递归深度对RLM性能的影响,发现更深层次递归会导致模型过度思考,从而降低性能并增加计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10246 2026-03-04 cs.HC cs.CL 57%

Automated Coding of Communications in Collaborative Problem-solving Tasks Using ChatGPT

利用ChatGPT进行协作问题解决任务中通信的自动化编码

Jiangang Hao, Wenju Cui, Patrick Kyllonen, Emily Kerzabi, Lei Liu, Michael Flor

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL

AI总结 本文利用ChatGPT对协作问题解决任务中的通信数据进行自动化编码,探讨不同模型和框架对编码效果的影响,并提出通过优化提示提高编码准确性的方法。

Comments 21 pages, 3 figures, 5 tables. Initially report in the edArXiv:xw6kz

Journal ref Journal of Educational Measurement, (2025). Volume 62, Issue 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15569 2026-03-04 cs.HC 50%

What Are You Doing? Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step Processing

你在做什么?代理LLM车载助手在多步骤处理中的中间反馈影响

Johannes Kirmayr, Raphael Wennmacher, Khanh Huynh, Lukas Stappen, Elisabeth André, Florian Alt

专题命中 复杂问题求解 :reasoning(abstract)

AI总结 研究探讨了代理式LLM车载助手在多步骤处理中中间反馈对用户体验的影响,发现中间反馈能提升信任和效率,同时减少任务负荷,提出适应性反馈策略以平衡透明度与效率。

Comments Accepted at CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04426 2026-03-04 cs.CV 50%

Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation

自适应与自纠正的遮蔽预测用于电影预告片生成

Sidan Zhu, Hongteng Xu, Dixin Luo

机构 * Key Laboratory of Artificial Intelligence Ministry of Education, Shanghai(教育部人工智能重点实验室,上海)

专题命中 复杂问题求解 :self-correction(abstract)

AI总结 SSMP通过双向上下文建模和逐步自纠正方法实现电影预告片生成的最先进结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02590 2026-03-04 cs.CR 50%

Extending the Formalism and Theoretical Foundations of Cryptography to AI

扩展密码学形式化和理论基础以适应人工智能

Federico Villa, F. Betül Durak, Tadayoshi Kohno, Tapdig Maharramli, Franziska Roesner

专题命中 复杂问题求解 :reasoning(abstract)

AI总结 本文提出了一种基于形式化方法的代理访问控制框架,通过统一保密性、完整性和可用性,解决了现有授权机制缺乏共享基础的问题,并证明了模块化分解在代理系统安全性中的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23141 2026-03-04 cs.CV 50%

Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents

Earth-Agent: 解锁地球观测的全貌与潜力

Peilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang, Xinjie Huo, Jinhua Yu, Wanghan Xu, Wenlong Zhang, Lei Bai, Conghui He, Weijia Li

专题命中 复杂问题求解 :reasoning(abstract)

AI总结 Earth-Agent是一种结合RGB和光谱数据的代理框架,通过多模态工具生态系统实现跨模态、多步骤推理,提升地球观测分析的科学性和应用潜力。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏