MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR 2025
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR 2025
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted as a poster paper by ICLR2024. 27 pages, 5 figures, 18 tables. [Source Code](https://github.com/iQua/llmpebase/tree/main/examples/BoTReasoning)
专题命中 复杂问题求解 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS (Spotlight)
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)
Comments Accepted by NDSS Symposium 2025. Please cite this paper as "Yi Yang, Jinghua Liu, Kai Chen, Miaoqian Lin. The Midas Touch: Triggering the Capability of LLMs for RM-API Misuse Detection. In the 32nd Annual Network and Distributed System Security Symposium (NDSS 2025)."
专题命中 复杂问题求解 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICML 2024
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)
专题命中 复杂问题求解 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to NeurIPS 2023 (spotlight). Project website: https://swiftsage.github.io
专题命中 复杂问题求解 :reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 复杂问题求解 :chain-of-thought(abstract);planning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 54 pages, 16 figures
阅读并非推理:弥合视觉-文本压缩中的智能体策略差距
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 该研究针对视觉-文本压缩导致的智能体策略差距,提出CAPS跨模态智能体策略自蒸馏框架,在SearchQA、ALFWorld等数据集上显著提升性能并大幅降低上下文成本。
复杂天然产物的优先策略合成规划
专题命中 复杂问题求解 :planning(title);分类 cs.AI
AI总结 基于大语言模型的智能体框架SynthEx,可规划出传统算法无法实现的复杂天然产物合成路线,其关键步骤获化学家认可,相关路线数据库SynthAtlas已开放。
BioProVLA-Agent:一种经济实惠、基于协议、视觉增强的VLA启用的具身多智能体系统,具备闭环推理能力,用于生物实验室操作
机构 * Key Laboratory of Smart Manufacturing in Energy Chemical Process Ministry of Education, East China University of Science and Technology, Shanghai, CN(能源化工过程智能制造教育部重点实验室,东华大学,上海,中国) ; Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, CN(东华大学计算机科学与工程系,上海,中国) ; Department of Laboratory Medicine, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, CN(复旦大学附属瑞金医院检验医学科,上海,中国) ; School of Information Science and Technology, Shihezi University, Shihezi, CN(石河子大学信息科学与技术学院,石河子,中国)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 本文提出BioProVLA-Agent,通过协议驱动和视觉增强,实现生物实验室操作中的具身多智能体系统,具备闭环推理能力,提升在湿实验室环境中的执行稳定性。
Comments 17 pages, 10 figures
MemSifter: 通过结果驱动的代理推理卸载LLM记忆检索
机构 * Gaoling School of Artificial Intelligence Renmin University of China Beijing China(中国人民大学朝阳学校人工智能 北京 中国) ; Renmin University of China(中国人民大学)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 提出MemSifter框架,用小模型代理替代主LLM进行记忆检索,通过强化学习优化检索结果,在8个基准上达到或超越现有方法。
Comments Code and datasets are available at https://github.com/plageon/MemSifter
思维即压缩:你的推理模型其实是一个上下文压缩器
机构 * Baidu Inc.(百度公司) ; Xi’an Jiaotong University(西安交通大学) ; City University of Hong Kong(香港城市大学) ; Queen Mary University of London(伦敦玛丽女王大学)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 本文提出思维即压缩(TaC)范式,利用推理模型自身的思维痕迹作为压缩上下文,并通过奖励驱动优化(TaC-C)实现可控压缩,在长上下文QA任务上显著优于现有方法。
Comments Under Review
ROS-LLM:一个用于具身AI的ROS框架,具有任务反馈和结构化推理
机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) ; University of Leeds(利兹大学) ; Technical University of Darmstadt(达姆施塔特技术大学) ; East China Normal University(华东师范大学) ; Huawei Technologies(华为技术有限公司) ; ETH Zurich(苏黎世联邦理工学院) ; University College London(伦敦大学学院)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 本文提出一种基于ROS的框架,允许非专家通过自然语言提示和上下文信息编程机器人,结合大语言模型实现任务描述和行为模式控制,实验验证了其在多样化场景中的鲁棒性和扩展性。
Comments This document contains 26 pages and 13 figures
Journal ref Nature Machine Intelligence 8, 313-325 (2026)
自主测试修复的实践极限:基于LLM驱动发现与自我校正的多智能体案例研究
机构 * Independent Researcher(独立研究者)
专题命中 复杂问题求解 :self-correction(title);分类 cs.AI
AI总结 本文通过多智能体系统在企业UI测试中的应用,探讨自主测试的限制,发现受限自主性能提升系统稳定性与可靠性,需结合验证边界和人工监督。
Comments Industrial case study; submitted for review
BRIEF-Pro:基于短到长合成的通用上下文压缩用于快速准确的多跳推理
机构 * University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 复杂问题求解 :reasoning(title);分类 cs.CL
AI总结 BRIEF-Pro通过短到长合成实现通用上下文压缩,提升多跳推理的效率和准确性,适用于多种语言模型。
Comments Accepted by ACL 2026 Findings. Code and data: https://github.com/JasonForJoy/BRIEF
LEAD:突破长horizon推理中的无恢复瓶颈
机构 * EPFL(瑞士联邦理工学院) ; Apple(苹果公司)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 本文提出LEAD方法,通过结合短horizon验证与重叠回放,解决长horizon推理中因非均匀误差分布导致的无恢复瓶颈问题,使o4-mini模型能解决复杂度达n=13的Checkers Jumping问题。
Comments 28 pages, 5 figures, 2 tables. Updated version to reflect the manuscript under review at COLM 2026
BAAI心脏代理:一种智能多模态代理,用于从心脏磁共振成像自动推理和诊断心血管疾病
机构 * Beijing Academy of Artificial Intelligence(北京智源人工智能研究院) ; Department of Radiology, Beijing Anzhen Hospital, Beijing Institute of Heart, Lung & Vascular Diseases, Capital Medical University(首都医科大学附属北京安贞医院放射科,北京市心肺血管疾病研究所) ; Department of MR, the First Affiliated Hospital, Henan Medical University(河南医科大学第一附属医院磁共振科) ; Department of Cardiology, Clinical Center for Coronary Heart Disease, Beijing Institute of Heart, Lung and Blood Vessel Disease, Beijing Anzhen Hospital, Capital Medical University(首都医科大学附属北京安贞医院心内科,冠心病临床中心,北京市心肺血管疾病研究所) ; Department of Cardiac Surgery, Beijing Anzhen Hospital, Institute of Heart, Lung and Vascular Diseases, Capital Medical University(首都医科大学附属北京安贞医院心脏外科,心肺血管疾病研究所)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 本文提出BAAI心脏代理,通过多模态智能系统实现心脏磁共振成像的端到端解读,实现了自动分割、功能量化、组织特征分析和疾病诊断,并在多个心血管疾病数据集上验证了其高准确率和高效性。
不止是‘手段’:通过透明设计的AI数据科学过程支持推理
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 本文探讨了医疗领域中两个AI数据科学系统如何通过透明设计的中间产物促进推理,强调了中间产物在帮助用户分析选择和改进问题中的关键作用。
Comments Accepted to Workshop on Tools for Thought at CHI'26: Understanding, Protecting, and Augmenting Human Cognition with Generative AI - From Vision to Implementation
PRECEPT:通过经验、上下文工程与轨迹探测进行规划韧性 一个具有组合规则学习和帕累托引导提示演化的测试时间适应统一框架
专题命中 复杂问题求解 :planning(title);分类 cs.AI
AI总结 PRECEPT通过经验、上下文工程与轨迹探测,结合组合规则学习和帕累托引导提示演化,提升测试时间适应的鲁棒性和泛化能力。
Comments 50 pages, 14 figures. Code and reproducibility resources: https://github.com/arash-shahmansoori/precept-framework
VERIFY-RL: 在数学推理中用于强化学习的可验证递归分解
机构 * School of Computing and Artificial Intelligence(计算与人工智能学院)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 Verify-RL通过可验证的递归分解提升数学推理强化学习的准确性和稳定性。
Comments 13 pages
ChatCFD: 一种基于大语言模型的端到端CFD自动化代理
机构 * Department of Mechanics and Aerospace Engineering, Southern University of Science and Technology(南方科技大学机械与航空航天工程系) ; School of Astronautics, Beihang University(北航航天学院) ; DP Technology, Beijing(北京DP技术) ; College of Future Technology, Peking University(北京大学未来技术学院) ; National Biomedical Imaging Center, Peking University(北京大学国家生物医学成像中心) ; State Key Laboratory of High-Efficiency Reusable Aerospace Transportation Technology(高性能可重复使用航天运输技术国家重点实验室) ; Key Laboratory of Spacecraft Design Optimization and Dynamic Simulation Technology, Ministry of Education(航天器设计优化与动态仿真技术重点实验室)
专题命中 复杂问题求解 :reasoning(title);分类 cs.CL
AI总结 ChatCFD通过结构化知识和推理,实现端到端CFD自动化,显著提升执行成功率和物理保真度,具备高灵活性和模块化设计。
Comments 19 pages, 8 figures
Journal ref Adv. Intell. Discov. 2025, 2500174
从单 agent 到多 agent 推理:提升 GeneGPT 的基因组 QA 能力
机构 * University of Padua, Italy(帕多瓦大学,意大利) ; Aalto University, Finland(阿alto大学,芬兰)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 GenomAgent通过多 agent框架提升基因组 QA 能力,实现对复杂基因组查询的高效处理,并在多个任务中超越现有系统。
Comments Accepted paper by the 48th European Conference on Information Retrieval (ECIR'26)
ComfySearch: 为ComfyUI工作流实现自主探索与推理
机构 * ComfyUI
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 ComfySearch通过验证引导的工作流构建,有效提升ComfyUI工作流的通过率和解决方案率,适用于复杂创意任务。
评估可回收性中的情境智能:图像推理系统的全面研究
机构 * Harvard College(哈佛学院) ; Stanford University(斯坦福大学) ; Department of Biomedical Informatics(生物医学信息学系) ; Harvard Medical School(哈佛医学院)
专题命中 复杂问题求解 :reasoning(title);分类 cs.AI
AI总结 本研究利用先进视觉-语言模型评估物品可回收性,探讨其在不同场景下的表现,揭示模型在情境理解上的进展与不足。
Comments x
从检索到推理:一个面向网络威胁情报实体识别的框架,具有显式和自适应指令
机构 * Sichuan University(四川大学) ; Renmin University of China(中国人民大学) ; Wuhan University(武汉大学) ; Engineering Research Center of Next-Generation Intelligent Search and Recommendation, Ministry of Education, China(教育部下一代智能搜索与推荐工程研究中心)
专题命中 复杂问题求解 :reasoning(title);分类 cs.CL
AI总结 本文提出TTPrompt框架,通过显式指令和反馈驱动细化,提升网络威胁情报实体识别的性能。