arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-06-09 至 2026-06-09 共收录 109 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 109 篇

2606.08633 2026-06-09 cs.AI cs.LG 新提交 93%

Towards Long-Horizon Vessel Trajectory and Destination Forecasting with Reasoning Large Language Models

面向长时域船舶轨迹与目的地预测的推理型大语言模型

Hongwei Wang, Miao Zhou, Fengde Wang, Yuting Wang, Jiewen Yu, Jun-Yan He, Bohao Qu, Wanbing Zhang, Xiuju Fu, Qing Guo, Zipei Fan, Yingying Xing, Yi Yuan

机构 * Institute of High Performance Computing (IHPC), A*STAR, Singapore(新加坡科技研究局高性能计算研究所) The Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University(同济大学道路与交通工程教育部重点实验室) Meituan Inc., Shenzhen, China(美团(深圳)) Centre for Frontier AI Research (CFAR), A*STAR, Singapore(新加坡科技研究局前沿人工智能研究中心) Nankai University(南开大学) School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);post-training(abstract)

AI总结 提出基于可验证奖励强化学习(RLVR)的Maritime LLM后训练框架,将轨迹转化为语义文本,通过物理有效性约束和层次匹配提升长时域(30天)预测精度,4B模型表现最优。

Comments The IEEE International Conference on Intelligent Transportation Systems (ITSC) 2026, Naples, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06915 2026-06-09 cs.CL cs.AI cs.LG 版本更新 92%

ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning

ThinkBooster: 一种用于LLM推理无缝测试时扩展的统一框架

Vladislav Smirnov, Chieu Nguyen, Sergey Senichev, Minh Ngoc Ta, Ekaterina Fadeeva, Artem Vazhentsev, Daria Galimzianova, Nikolai Rozanov, Viktor Mazanov, Jingwei Ni, Tianyi Wu, Igor Kiselev, Mrinmaya Sachan, Iryna Gurevych, Preslav Nakov, Timothy Baldwin, Artem Shelmanov

机构 * MBZUAI ETH Zürich(苏黎世联邦理工学院) Imperial College London(伦敦帝国理工学院) NUS(国立大学新加坡) Accenture(埃森哲) Innopolis University(因诺普里斯大学) Independent Researcher(独立研究者)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出ThinkBooster框架,通过模块化库、联合评估基准和可部署代理服务,实现LLM推理的测试时计算扩展,在数学和编码任务上验证了性能-计算权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11874 2026-06-09 cs.GT cs.AI cs.DS cs.LO cs.PL 版本更新 92%

Discovering Expert-Level Nash Equilibrium Algorithms with Large Language Models

利用大型语言模型发现专家级纳什均衡算法

Hanyu Li, Dongchen Li, Xiaotie Deng

机构 * CFCS, School of Computer Science, Peking University, Beijing, China(计算机科学系,北京大学,北京,中国) School of Computing and Data Science, The University of Hong Kong, Pokfulam, Hong Kong(计算与数据科学学院,香港大学,薄扶林,香港)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 提出LegoNE框架,将专家证明策略编码为符号语言,自动验证算法的最坏情况保证,结合推理型LLM重新发现并改进了多人博弈的近似纳什均衡算法。

Comments accepted by Nature Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09027 2026-06-09 cs.CL cs.AI 新提交 92%

SafeRun: Enabling Determinism in LLM Planning for Running

SafeRun:在跑步规划中实现LLM的确定性

Meilin Chen, Zepeng Zhai, Jiaxuan Zhao, Yuan Lu

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 针对LLM在跑步规划中因概率性导致安全违规的问题,提出SafeRun框架,通过解耦架构将LLM的软解释与确定性求解器的硬约束分离,实现100%安全评分。

Comments Workshop on Planning in the Era of LLMs (LM4Plan) at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07547 2026-06-09 cs.CL cs.AI cs.SD 新提交 92%

Liberating LLM Capabilities in Full-Duplex Speech Models

在全双工语音模型中释放LLM能力

Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu, Hanyu Liu, Yuan Yao

机构 * Royal Zhang(皇家张)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出Listen-Write-Speak (LWS)三通道范式,使LLM在共享因果注意力上下文中同时监听、书写可见文本并实时口语回应,无需架构修改,实现全双工交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09751 2026-06-09 cs.AI cs.CL cs.LO 版本更新 92%

Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations

基于LLM解释的完备且可靠的神经常识推理

Bradley P. Allen, Prateek Chhikara, Thomas Macaulay Ferguson, Filip Ilievski, Paul Groth

机构 * University of Amsterdam(阿姆斯特丹大学) University of Southern California(南加州大学) Rensselaer Polytechnic Institute(拉特格斯理工学院) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出将LLM直接集成到次协调逻辑的语义解释函数中,实现可靠且完备的神经常识推理,在GPQA和SimpleQA基准上宏F1提升约6个百分点,并成功检测药物安全知识库中的矛盾。

Comments 43 pages, 14 tables, 4 figures. Accepted to the 19th Conference on Neurosymbolic Learning and Reasoning (NeSy 2025); to appear Neurosymbolic Artifical Intelligence Special Issue on NeSy 2025 Extended Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08596 2026-06-09 cs.AI cs.HC 新提交 92%

Distilling LLM Reasoning into an Interpretable Policy Tree for Human-AI Collaboration

将LLM推理蒸馏为可解释的策略树用于人机协作

Beiwen Zhang, Yongheng Liang, Guowei Zou, Haitao Wang, Hejun Wu

机构 * Sun Yat-sen University(中山大学)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出Co-pi-tree方法,通过将大语言模型推理蒸馏为可执行策略树,在Overcooked-AI中平均奖励提升35.4%,同时减少77.7%的LLM查询和97.1%的测试延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09484 2026-06-09 cs.CL 新提交 91%

Detecting Differences Is Not Understanding Structure: Large Language Models Fail at Graph Isomorphism

检测差异不等于理解结构:大型语言模型在图同构任务中失败

Kumar Thushalika, Sukumar Kishanthan, Asela Hevapathige

机构 * University of Ruhuna(鲁胡纳大学) University of Moratuwa(莫拉图瓦大学) University of Melbourne(墨尔本大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.CL

AI总结 本研究通过图同构检测任务揭示LLM的“虚假成功”:虽然LLM在检测同构时准确率接近完美,但面对节点标签置换的相同图时却无法识别,表明其依赖模式而非抽象结构推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02802 2026-06-09 cs.AI 版本更新 90%

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

ChatHealthAI: 将电子健康记录表示与大语言模型对齐以实现基于临床的推理

Bo-Hong Wang, Baicheng Peng, Ruilin Wang, Jun Bai, Ziyang Song, Yue Li

机构 * School of Computer Science, McGill University(麦吉尔大学计算机科学系) Mila - Quebec AI Institute(魁北克人工智能研究所)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);foundation model(abstract)

AI总结 提出ChatHealthAI框架,通过任务感知重采样器将预训练的EHR基础模型的结构化表示与冻结的大语言模型语义空间对齐,实现可解释的临床推理并保持预测性能。

Comments Main paper with appendix, 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01637 2026-06-09 cs.CL cs.AI 版本更新 90%

Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity

误导比纠正更容易:LLM 从众中的有害与有益修正

Jiaming Qu, Lucheng Fu, Yibo Hu

机构 * Amazon(亚马逊) Georgia Institute of Technology(佐治亚理工学院) Illinois Institute of Technology(伊利诺伊理工学院)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 通过控制实验,研究大语言模型在多智能体系统中面对同伴答案时的从众行为,发现同伴一致意见更容易误导原本正确的模型,而权威标签使模型更倾向于选择被认可的答案,且通用推理干预无法可靠地减少有害修正。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09071 2026-06-09 cs.AI 新提交 90%

REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces

REFLECT: 针对LLM智能体轨迹中静默失败的干预支持错误归因

Xiaofeng Lin, Yingxu Wang, Tung Sum Thomas Kwok, Daniel Guo, Sahil Arun Nale, Charles Fleming, Guang Cheng

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出REFLECT方法,通过诊断候选错误步骤、使用诊断特定补丁进行受控重放测试,并利用验证结果作为对比证据来细化归因,在四个基准上取得最高定位准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08661 2026-06-09 cs.CR cs.AI cs.DB 新提交 90%

Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems

数据代理遭受攻击:LLM驱动的分析系统中的漏洞

Kuncan Wang, Ziting Wang, Peizhuo Lv, Haoyang Li, Guoliang Li, Gao Cong, Wei Dong

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡) The Hong Kong Polytechnic University(香港理工大学) Tsinghua University(清华大学)

专题命中 推理与问题求解 :LLM(title,title_cn);分类 cs.AI

AI总结 本研究系统分析了LLM驱动的数据代理的安全漏洞,提出了分层漏洞框架和攻击分类法,并在六个系统上评估了攻击效果,揭示了当前系统的重大安全缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07790 2026-06-09 cs.LG 新提交 90%

Byzantine Cheap Talk: Adversarial Resilience and Topology Effects in LLM Coordination Games

拜占庭廉价谈话:LLM协调博弈中的对抗韧性与拓扑效应

Aya El Mir, Martin Takáč, Salem Lahlou

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 推理与问题求解 :LLM(title,title_cn);分类 cs.LG

AI总结 研究多智能体LLM在协调博弈中面对拜占庭攻击和通信拓扑限制的脆弱性,发现智能体无法集体适应背叛,且显式限制拓扑会破坏合作,而隐式限制则不影响。

Comments Accepted at NETYS 2026 (The International Conference on Networked Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06399 2026-06-09 cs.CL 版本更新 90%

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments

CollabSim: 一种基于CSCW的方法,通过受控多智能体实验研究LLM智能体的协作能力

Jiaju Chen, Bo Sun, Yuxuan Lu, Yun Wang, Dakuo Wang, Bingsheng Yao

机构 * Northeastern University(东北大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出CollabSim框架,结合CSCW理论定义协作能力、控制交互条件并探测智能体内部状态,以系统分析LLM多智能体系统的协作能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09134 2026-06-09 cs.RO cs.AI cs.CL cs.CV cs.GR 新提交 90%

From USD Scenes to Knowledge Graphs: Zero-Shot Ontology Grounding with LLMs

从USD场景到知识图谱:基于LLM的零样本本体接地

Jiangtao Shuai, Zongxiong Chen, Manfred Hauswirth, Sonja Schimmler

机构 * Technical University of Berlin(柏林工业大学) Fraunhofer FOKUS(弗劳恩霍夫开放通信系统研究所)

专题命中 推理与问题求解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 研究利用大语言模型(LLM)零样本地将3D场景对象自动映射到本体类别,无需训练,在厨房场景中达到90-96%准确率,并揭示语义线索是关键。

Comments Accepted to the IEEE ICRA 2026 International Joint Workshop on Ontologies, Semantic Maps and Autonomous Robotics Standardization (J-WOSMARS 2026), Vienna, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11189 2026-06-09 cs.AI cs.LG 版本更新 90%

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP

全局XAI方法能否揭示LLM中的注入行为?SHAP vs 规则提取 vs RuleSHAP

Francesco Sovrano

机构 * Collegium Helveticum at ETH Zurich(苏黎世联邦理工学院霍夫曼学院) Università della Svizzera italiana(瑞士联邦理工学院)

专题命中 推理与问题求解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究通过统计验证的抽象将全局LLM信念映射为数值分数,提出RuleSHAP算法,结合全局SHAP与规则归纳,以更好地捕捉非单变量触发因素,平均MRR@1比RuleFit提升82%。

Comments Accepted for publication at KDD'2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07940 2026-06-09 cs.CR 新提交 89%

SGTO-MAS: Secure Gorilla Troops Optimization for Multi-Agent LLM Systems

SGTO-MAS:面向多智能体大语言模型系统的安全大猩猩部队优化

Saeid Jamshidi

专题命中 推理与问题求解 :LLM(title,summary_cn);large language model(abstract);language model(abstract)

AI总结 提出一种基于大猩猩部队优化的安全感知多智能体LLM协调方法,通过信任建模、风险感知和集体智能的联合优化,在性能、安全性和效率间取得平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16136 2026-06-09 cs.RO 版本更新 89%

Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning

基于思维图的奖励进化:一种用于强化学习的双层语言模型框架

Changwei Yao, Xinzi Liu, Chen Li, Marios Savvides

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Tokyo(东京大学)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);language model(title,abstract);large language model(abstract)

AI总结 本文提出RE-GoT框架,结合LLM与VLM的图思维推理,通过任务分解和视觉反馈迭代优化奖励函数,实验表明在RoboGen和ManiSkill2任务中均优于现有方法。

Journal ref IEEE International Conference on Robotics and Automation (ICRA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07548 2026-06-09 cs.IR cs.AI cs.CL 新提交 89%

Evaluating Advanced Prompting on Gemini Flash for Multi-Hop Biomedical QA

评估 Gemini Flash 上的高级提示工程用于多跳生物医学问答

Ahmed Bajaber, Mohammed Alliheedi

机构 * Saudi Med AI Lab (SMAIL)(沙特医学人工智能实验室(SMAIL)) Prince Sultan University(普森国王大学) Al-Baha University(阿勒巴哈大学)

专题命中 推理与问题求解 :LLM(summary_cn,abstract_cn);prompting(title);large language model(abstract);language model(abstract)

AI总结 本研究通过设计多组件提示(角色扮演、多步思维链示例和格式规则),在 Gemini 2.0 Flash 上实现概念级得分0.720,显著优于基线0.565,并接近下一代模型性能,证明高级提示设计对释放LLM推理能力至关重要。

Comments 8 pages, proceedings of the BioCreative IX Challenge and Workshop (BC9) at IJCAI 2025

Journal ref Proc. BioCreative IX Workshop (BC9), IJCAI 2025, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09131 2026-06-09 cs.AI cs.CL cs.CV cs.LG 新提交 89%

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation

晚期融合足矣:面向视觉饱和的多模态大语言模型的双路径视觉令牌路由

Siyuan Liu, Jinyang Wu

机构 * School of Mechanics and Engineering Science, Peking University(北京大学力学与工程科学学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对多模态大语言模型中视觉令牌在深层饱和的问题,提出双路径视觉令牌路由(DPVR-LF),在饱和点将视觉令牌路由至单层可训练分支,仅最后层融合,以约3%可训练参数保持性能并减少计算。

Comments 18 pages, 4 figures. Submitted to Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07529 2026-06-09 cs.CL cs.AI cs.CV cs.LG cs.MM 新提交 89%

CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models

CAPruner: 概念相邻场景图剪枝器以增强大语言模型的3D空间推理

Shengli Zhou, Xiangchen Wang, Guanhua Chen, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) SpatialTemporal AI(时空人工智能)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出概念相邻场景图剪枝器(CAPruner),通过融合模糊语义相关性和空间邻近性估计关系重要性,在任务特定上下文中选择关键关系,避免关系级标注,显著提升大语言模型在3D视觉语言任务上的空间推理性能。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09193 2026-06-09 cs.CL cs.AI cs.HC 88%

(Ir)rationality and Cognitive Biases in Large Language Models

非理性与大语言模型中的认知偏差

Olivia Macmillan-Scott, Mirco Musolesi

机构 * University College London(伦敦大学) University of Bologna(博洛尼亚大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文通过心理学文献中的任务评估七种语言模型,发现其在非理性表现上与人类相似,但表现形式不同,且存在响应不一致的额外非理性特征。

Journal ref Royal Society Open Science 11(6) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08976 2026-06-09 cs.AI 新提交 88%

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models

RTL-BenchLS:面向大语言模型的RTL推理与生成的大规模基准

Jing Wang, Shang Liu, Wenji Fang, Yuchao Wu, Yugao Zhu, Zhiyao Xie

机构 * Hong Kong University of Science and Technology(香港科技大学)

专题命中 推理与问题求解 :large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出大规模基准RTL-BenchLS,包含超1万个形式验证的Verilog设计,并引入三项自监督推理任务,解决现有基准规模小、任务单一的问题,评估显示当前最佳模型性能较低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08501 2026-06-09 cs.CL 新提交 88%

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models

重回正轨:在扩散大语言模型中对齐奖励与状态以进行推理

Yawen Shao, Jie Xiao, Kai Zhu, Yu Liu, Hongchen Luo, Xueyang Fu, Yang Cao, Wei Zhai, Zheng-Jun Zha

机构 * University of Science and Technology of China(中国科学技术大学) Tongyi Lab(通义实验室) Northeastern University(东北大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 针对扩散大语言模型强化学习中过程奖励与状态轨迹的双重错位问题,提出PAPO框架,通过步骤感知过程奖励和熵引导历史重演实现对齐,在四个基准上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08310 2026-06-09 cs.AI cs.MA 新提交 88%

To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation

核弹还是和平:大语言模型在高风险决策模拟中的(缺失的)伦理推理与行动

John Chen, Sihan Cheng, Can Gurkan, H M Abdul Fattah

机构 * University of Arizona(亚利桑那大学) Northwestern University(西北大学)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 研究LLM在复杂游戏《文明V》中自发升级核授权的现象,通过三种提示干预发现伦理推理未能可靠消除升级,识别出三种失败路径,强调需在复杂决策上下文中测试伦理推理的自发性和行为有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07602 2026-06-09 cs.LG cs.AI 新提交 88%

Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning

面向LEGO空间物理推理的样本高效后训练

Yuhuan Yuan, Zhouliang Yu, Minghao Liu, Weiyang Liu, Ge Lin Kan

机构 * HKUST(GZ)(香港科技大学(广州)) CUHK(香港中文大学) ZODA

专题命中 推理与问题求解 :LLM(summary_cn,abstract);post-training(title);分类 cs.AI、cs.LG

AI总结 针对LLM生成LEGO组装时出现的物理有效但几何语义错位问题,提出基于模型的数据选择方法和样本高效强化学习PVPO,结合体素空间几何奖励,提升结构、语义对齐和物理有效性。

Comments Technical Report V1, 15 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12666 2026-06-09 cs.LG cs.AI 版本更新 88%

RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction

RetroReasoner:一种用于战略 retrosynthesis 预测的推理 LLM

Hanbum Ko, Chanhui Lee, Ye Rin Kim, Rodrigo Hormazabal, Sehui Han, Sungbin Lim, Sungwoong Kim

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系) Department of Statistics, Korea University(韩国大学统计系) Materials Intelligence Lab, LG AI Research(LG人工智能研究实验室)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 RetroReasoner 通过监督微调和强化学习,捕捉化学家基于断键策略的推理过程,提升 retrosynthesis 预测的准确性和多样性。

Comments 35 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08688 2026-06-09 cs.RO cs.CV 新提交 88%

PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback

PhysAgent: 通过轨迹驱动的多智能体反馈实现基于物理的4D合成自动化

Chunji Lv, Jiaxi Ye, Yuchen Jiang, Rexar Lin, Changsheng Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 提出PhysAgent,首个模拟器在环的多智能体框架,通过解耦材料与外力、利用视觉基础模型提取轨迹并借助LLM常识推理,实现自动化、物理可信的4D运动合成,显著提升生成多样性与物理准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01880 2026-06-09 cs.RO 88%

Multimodal Large Language Models for Real-Time Situated Reasoning

多模态大语言模型用于实时情境推理

Giulio Antonio Abbo, Senne Lenaerts, Tony Belpaeme

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract)

AI总结 本文探讨多模态大语言模型如何支持实时情境和价值感知决策,结合GPT-4o与模拟智能扫地机器人平台,展示其在家庭活动、社会规范和用户偏好推理中的能力,以及在清洁、舒适和安全等价值上的细致决策。

Comments Submitted to the interactivity track of the 21st ACM/IEEE International Conference on Human-Robot Interaction on December 2025, accepted January 2026

Journal ref HRI Companion 2026: Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08543 2026-06-09 cs.AI 新提交 87%

PAEC: Position-Aware Entropy Calibration for LLM Reasoning in RLVR

PAEC:面向RLVR中LLM推理的位置感知熵校准

Shumeng Yang, Yisu Liu, Jiayi Zheng, Zhaohui Yang, Linjing Li

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出位置感知熵校准(PAEC),通过局部top-p熵和top-2候选竞争构建软掩码,并施加基于锚点的下界惩罚,防止决策相关位置熵崩溃,提升数学推理性能。

Comments 22 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏