arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-06-23 至 2026-06-23 共收录 778 信号源:cs.CL, cs.AI, cs.LG

1. 推理与问题求解 107 篇

2606.20624 2026-06-23 cs.AI cs.CL cs.LG stat.ML 新提交 90%

In LLM Reasoning, there is Irrationality on top of Value Misalignment

在LLM推理中,价值对齐之上存在非理性

Kejiang Qian, Fengxiang He

机构 * University of Edinburgh(爱丁堡大学)

专题命中 推理与问题求解 :LLM(title,title_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出理性价值风险概念,形式化LLM在推理中即使价值对齐也可能无法最大化对齐价值的问题,并通过实验验证该风险普遍存在且无法被对齐消除。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14445 2026-06-23 cs.CL cs.AI cs.HC cs.LG 版本更新 90%

Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning

Tell Me:基于LLM的心理健康助手,集成RAG、合成对话生成与智能体规划

Trishala Jayesh Ahalpara

机构 * Fujitsu Research of America(富士通美国研究)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出Tell Me系统,利用大语言模型通过RAG实现个性化对话、合成客户-治疗师对话生成以及智能体规划生成自我护理计划,旨在提供可及的心理支持而非替代专业治疗。

Comments 8 pages, 2 figures, 1 Table. Submitted to the Computation and Language (cs.CL) category. Uses the ACL-style template. Code and demo will be released at: https://github.com/trystine/Tell_Me_Mental_Wellbeing_System

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22388 2026-06-23 cs.AI cs.CL 新提交 90%

PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems

PlanBench-XL:评估大规模工具生态系统中LLM工具使用智能体的长时程规划能力

Jiayu Liu, Qihan Lin, Cheng Qian, Rui Wang, Emre Can Acikgoz, Xiaocheng Yang, Jiateng Liu, Zhenhailong Wang, Xiusi Chen, Heng Ji, Dilek Hakkani-Tür

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 推理与问题求解 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 提出PlanBench-XL基准,包含327个零售任务和1665个工具,测试LLM智能体在检索受限工具环境下的长时程规划能力,实验表明GPT-5.4在无阻塞下准确率51.90%,严重阻塞下降至11.36%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20897 2026-06-23 cs.CL cs.AI 新提交 90%

PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality

PeerCheck: 提升大语言模型生成的学术评审至人类水平质量

Zeyuan Chen, Ziqing Yang, Yihan Ma, Michael Backes, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)

专题命中 推理与问题求解 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出PeerCheck框架,分析LLM与人类评审差异,采用思维链和检索增强生成提升评审质量,发现思维链显著改善但RAG存在悖论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03093 2026-06-23 cs.LG cs.CL 版本更新 90%

ATLAS: Verifier-Guided Adaptive Latent Activation Steering for Efficient LLM Reasoning

ATLAS:验证器引导的自适应潜在激活引导用于高效LLM推理

Tuc Nguyen, Thai Le

机构 * Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 提出ATLAS框架,通过轻量级验证器动态调整推理时潜在状态引导策略,实现每步自适应控制,在数学和编码推理任务上提升准确率并减少测试时token使用。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20607 2026-06-23 cs.HC 新提交 90%

LLM4CAD-Editor: An Intent-Aware Large Language Model Framework for Multi-Level Computer-Aided Design Editing

LLM4CAD-Editor:一种面向多层级计算机辅助设计编辑的意图感知大语言模型框架

Yuewan Sun, Zhenghui Sha

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn)

AI总结 提出LLM4CAD-Editor框架,通过结构化领域特定语言LLM4CAD-DSL实现自然语言驱动的CAD编辑,支持功能、操作和参数三级编辑,在参数级编辑上达到96.3%解析准确率和0.935 IoU。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22375 2026-06-23 cs.AI cs.CE cs.IR 新提交 90%

ARIA: A Causal-Aware Framework for Rescuing LLM Reasoning in Trustworthy Materials Discovery

ARIA: 一种用于在可信材料发现中拯救LLM推理的因果感知框架

Yi Cao, Liaoyaqi Wang, Jieneng Chen, Benjamin Van Durme, Alan Yuille, Paulette Clancy

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出ARIA框架,通过因果感知的三级级联(直接因果推理、物理类比迁移、参数回退)解决LLM在材料发现中的上下文隧道问题,提升物理可信性。

Comments Accepted to the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026)

Journal ref Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD '26), August 09--13, 2026, Jeju Island, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21740 2026-06-23 cs.AI 新提交 90%

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents

训练编排器:一种基于监督学习的端到端PDDL规划方法,使用LLM智能体

Rajesh Mangannavar, Zachary Coalson, Pranay Dugar, Prasad Tadepalli

机构 * Oregon State University(俄勒冈州立大学)

专题命中 推理与问题求解 :LLM(title,title_cn);分类 cs.AI

AI总结 提出HALO框架,通过监督学习训练编排器,利用验证器提供的轨迹作为监督信号,在11个PDDL领域上以极低成本实现与前沿LLM相当的成功率。

Comments 22 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20895 2026-06-23 cs.AI 新提交 90%

Neurosymbolic Clinical Trial Matching via LLM-Driven Abduction and Logical Verification

基于LLM驱动溯因与逻辑验证的神经符号临床试验匹配

Baiyang Qu, Leonardo Ranaldi, Xi Wang, Marco Valentino

机构 * University of Leicester(莱斯特大学) University of Edinburgh(爱丁堡大学) University of Sheffield(谢菲尔德大学)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种结合大语言模型与逻辑验证的神经符号框架αNeSy-CTM,通过溯因推理处理噪声和不完整临床文本,在临床试验匹配中相对零样本基线提升30%准确率。

Comments 21 pages (including appendix), 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14935 2026-06-23 cs.AI 新提交 90%

PrologMCP: A Standardized Prolog Tool Interface for LLM Agents

PrologMCP:面向LLM代理的标准化Prolog工具接口

Agnieszka Mensfelt, Adarsh Prabhakaran, Adrian Haret, Vince Trencsenyi, Kostas Stathis

机构 * Royal Holloway, University of London(伦敦大学皇家霍洛威学院)

专题命中 推理与问题求解 :LLM(title,title_cn);language model(abstract);分类 cs.AI

AI总结 提出PrologMCP,一个通过模型上下文协议将Prolog暴露为状态化工具的任务无关开源服务器,使LLM代理能够通过翻译-运行-检查-修复循环稳健地委托演绎推理,在PARARULE-Plus上达到或超越推理型LLM。

Comments Accepted at Joint Workshop on Statistics and Knowledge Integration for Logic, Learning, Ethical Decisions, and LLMs, 18 July 2026, Lisbon v2: Added references to other Prolog MCP servers; fixed typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00012 2026-06-23 cs.LG cs.AI cs.CY cs.IR 版本更新 90%

OGD4All: A Framework for Accessible Interaction with Geospatial Open Government Data Based on Large Language Models

OGD4All: 基于大语言模型的可访问地理空间开放政府数据交互框架

Michael Siebenmann, Javier Argota Sánchez-Vaquerizo, Stefan Arisona, Krystian Samp, Luis Gisler, Dirk Helbing

机构 * Professorship of Computational Social Science, ETH Zurich(计算社会科学教授职位,苏黎世联邦理工学院) Esri R&D Center Zurich(埃斯里苏黎世研发中心) Complexity Science Hub(复杂性科学中心)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.AI、cs.LG

AI总结 提出OGD4All框架,结合语义检索、智能体推理和沙盒执行,实现基于大语言模型的地理空间开放政府数据交互,在430个数据集上达到98%分析正确率和94%召回率,有效减少幻觉风险。

Comments Author Accepted Manuscript (AAM). Proceedings of 2026 IEEE CAI (Granada, Spain). Update manuscript with final DOI. Code & data available at: https://github.com/ethz-coss/ogd4all

Journal ref 2026 IEEE Conference on Artificial Intelligence (CAI), pp. 882-888

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20922 2026-06-23 cs.CR 新提交 89%

Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated Planning

三思而后行:通过隔离规划保护LLM代理免受工具描述投毒攻击

Shanghao Shi, Xiao Wang, Chaoyu Zhang, Hao Li, Wenjing Lou, Thomas Hou, Yevgeniy Vorobeychik, Chongjie Zhang, Ning Zhang

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 提出Tool-Guard系统级防御,通过隔离规划机制将可疑工具调用隔离到受感染列表,阻断投毒描述持续影响,在AgentDojo和ASB基准上显著降低攻击成功率并保持任务效用。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22354 2026-06-23 cs.RO 版本更新 89%

LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning

基于LLM的异构机器人团队通用分层任务规划与执行及事件驱动重规划

Suraj Borate, Bhavish Rai B, Vipul Pardeshi, Madhu Vadali

机构 * Department of Mechanical Engineering, IIT Gandhinagar(印度加尔各直辖区理工学院机械工程系) Sahyadri College of Engineering and Management(萨哈亚德里工程与管理学院) Vishwakarma Institute of Information Technology(维什瓦克arma信息技术学院)

专题命中 推理与问题求解 :LLM(title,title_cn)

AI总结 提出CoMuRoS架构,结合集中式规划与分布式执行,通过LLM解释自然语言目标、分配子任务,并支持事件驱动重规划,在硬件实验中实现自主恢复和协调运输。

Comments full version of this short paper is accepted at Frontiers in Robotics and AI Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22126 2026-06-23 cs.CL 新提交 89%

From Recognition to Understanding: Unlocking Cognitive Time Series Reasoning with LLMs

从识别到理解:利用LLM解锁认知时间序列推理

Xin Qiu, Junlong Tong, Yao Zhang, Yunpu Ma, Wei Zhang, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(东方理工数字孪生研究所) Zhejiang University(浙江大学) Munich Center for Machine Learning, LMU(慕尼黑机器学习中心,慕尼黑大学)

专题命中 推理与问题求解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 针对现有时间序列任务与LLM能力不匹配的问题,提出多模态基准TSCognition和统一框架TSAlign,通过认知推理任务和语义对齐提升LLM的时间序列推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23601 2026-06-23 cs.AI 版本更新 89%

Enhancing Diversity of LLM-Generated Educational Tasks

增强大语言模型生成教育任务的多样性

Manh Hung Nguyen, Sebastian Tschiatschek, Adish Singla

机构 * MPI-SWS(马克斯·普朗克研究所-斯图加特研究所) University of Vienna(维也纳大学)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 针对大语言模型生成教育任务时内容同质化的问题,提出基于发散-收敛思维的两阶段提示框架CreativeDC,在Python编程领域生成的任务多样性提升约1.6倍,同时保持高实用性。

Comments EDM 2026 Poster Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21121 2026-06-23 cs.AI cs.CL 新提交 88%

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models

答案工程:面向大语言模型中协议约束决策的局部轨迹编辑

Victor Lavrenko, Anastasiia Molodnitskaia

机构 * PeaceTech VC Rambam Health Care Campus(拉姆巴姆医疗中心)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 提出答案工程方法,通过在自回归生成时对推理轨迹进行局部规则引导编辑,无需重训练或全局搜索,在突发性感觉神经性听力损失临床基准上将平衡准确率从42.0%提升至80.7%。

Comments 31 pages, 6 figures. Code and data: https://github.com/victorlavrenko/answer-engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23312 2026-06-23 cs.RO 新提交 88%

From Pixels to Concepts: Growing Rich 3D Semantic Scene Graph Forests utilizing Foundation Models

从像素到概念:利用基础模型构建丰富的3D语义场景图森林

David Oberacker, Meike Deitersen, Niklas Spielbauer, Tristan Schnell, Georg Heppner, Arne Roennau

机构 * FZI Research Center for Information Technology(FZI信息技术研究中心) Machine Intelligence and Robotics Lab (MaiRo), Karlsruhe Institute for Technology (KIT)(卡尔斯鲁厄理工学院机器智能与机器人实验室)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);foundation model(title,abstract)

AI总结 提出利用基础模型构建具有开放语义关系的3D场景图森林,通过VLM和LLM推理抽象概念节点与关系,提升机器人场景理解与任务执行能力。

Comments To be published in the Proceedings of the IEEE/RSJ International Conference on Intelligent Robots & Systems (IEEE IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23092 2026-06-23 cs.CL 新提交 88%

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

PIVOTSBench:评估多模态大语言模型中的细粒度人际关系推理

Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu

机构 * Sun Yat-sen University(中山大学) University of Michigan(密歇根大学) Tsinghua University(清华大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 提出PIVOTS基准,基于Social-IQ 2.0和YouTube数据,评估多模态大语言模型在心理学研究基础上预测双向人际关系维度的能力,并包含辅助任务分析关键视觉线索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22909 2026-06-23 cs.DB cs.AI cs.IR 新提交 88%

Graph-Enhanced Large Language Models for Spatial Search

用于空间搜索的图增强大型语言模型

Nicole R. Schneider, Kent O'Sullivan, Hanan Samet

机构 * University of Maryland(马里兰大学) University of Sydney(悉尼大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 针对大语言模型空间推理能力不足的问题,提出图增强方法,将空间数据以图形式存储并与检索增强生成结合,提升复杂空间问题回答能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21832 2026-06-23 cs.AI 新提交 88%

AgentCAT: Simulating Computerized Adaptive Testing via Multi-Agent Large Language Models

AgentCAT:通过多智能体大语言模型模拟计算机自适应测试

Weiyuan Zhou, Haiping Ma, Xiaoshan Yu, Changqian Wang, Shangshang Yang, Xingyi Zhang

机构 * Institute for Clarity in Documentation(文档清晰度研究所) Inria Paris-Rocquencourt(法国国家信息与自动化研究所巴黎-罗康库尔中心) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕默研究实验室) State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, Institute of Physical Science and Information Technology, Anhui University(光电信息获取与防护技术国家重点实验室,物理科学与信息技术学院,安徽大学) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) School of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院) School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China, the Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(认知智能国家重点实验室,中国科学技术大学,人工智能研究院,合肥综合性国家科学中心)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 提出基于大语言模型的多智能体仿真系统AgentCAT,通过考生、选题和监督三个模块模拟动态测试过程,实现能力估计与选题策略的平衡,在真实数据集上验证了有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21378 2026-06-23 cs.LG 新提交 88%

Enhancing Creativity in 3D Generative Design via a TRIZ-Inspired Text-to-CAD Framework

通过TRIZ启发的文本到CAD框架增强3D生成设计的创造力

Dongeon Lee, Leekyo Jeong, Soyoung Yoo, Sunwoong Yang, Namwoo Kang

机构 * Cho Chun Shik Graduate School of Mobility, KAIST, Daejeon, Republic of Korea(韩国科学技术院(KAIST)赵正植移动研究生院,大田,韩国) Samsung Electronics, Suwon, Republic of Korea(三星电子,水原,韩国) Department of Mechanical Engineering, Hanyang University, Ansan, Republic of Korea(汉阳大学机械工程系,安山,韩国) Narnia Labs, Daejeon, Republic of Korea(纳尼亚实验室,大田,韩国)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 提出TRIZ启发的文本到CAD框架,利用LLM生成可编辑CAD模型并系统探索创新设计方案,通过三阶段流程实现结构多样性与质量优化,案例显示质量减少4.0-14.7%。

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20969 2026-06-23 cs.AI cs.SE 新提交 88%

AutoACSL: Synthesizing ACSL Specifications by Integrating LLMs with CPG-Based Static Analysis

AutoACSL: 通过集成LLM与基于CPG的静态分析来综合ACSL规范

Han Zhou, Yu Luo, Dianxiang Xu

机构 * University of Missouri–Kansas City(密苏里大学堪萨斯城分校) University of Central Missouri(中央密苏里大学)

专题命中 推理与问题求解 :LLM(title_cn,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 提出AutoACSL框架,结合大语言模型与代码属性图语义特征,通过反馈驱动循环自动生成可验证的ACSL规范,在604个程序上达到98%生成成功率和96%完全证明率。

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24776 2026-06-23 cs.CL 88%

Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models

开源推理大语言模型的计算-准确率帕累托前沿

Ákos Prucs, Nara Csutora, Mátyás Antal, Márk Marosi

机构 * E-Group Research(E-Group研究) Budapest University of Technology and Economics(布达佩斯技术大学和经济学院)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文研究了开源大语言模型在数学和推理任务中的计算与准确率平衡,发现MoE架构在性能和效率上表现优异,并揭示了计算资源与准确率增长的关系趋势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21724 2026-06-23 cs.CL cs.AI 新提交 88%

Denoising Iterative Self-Correction: Structured Verification Loops for Reliable LLM Reasoning

去噪迭代自校正:用于可靠LLM推理的结构化验证循环

Shen Yin, David Ken, Joel Stremmel

机构 * Thomson Reuters Labs(汤森路透实验室)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出DISC方法,通过将验证输出视为噪声测量,在多次验证-判断-校正循环中逐步减少推理错误,并引入二元判断门控机制防止破坏正确答案,在三个基准上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22570 2026-06-23 cs.CL 新提交 87%

What are Key Factors for Updates in RL for LLM Reasoning?

RL提升LLM推理能力的关键更新因素是什么?

Peidong Wang, Demi Wang, Xufang Luo, Jiahang Xu, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li

机构 * School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Microsoft Research(微软研究院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 通过理论分析RLVR更新,发现离策略程度影响重要性采样比率分布和裁剪行为,提出自适应裁剪策略优化(ACPO),在多种推理基准上优于DAPO和CISPO。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19331 2026-06-23 cs.CL 版本更新 87%

Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation

评估基于大语言模型的议会辩论总结的论证内容

Eoghan Cunningham, Derek Greene, James Cross, Antonio Rago

机构 * School of Computer Science, University College Dublin(都柏林大学计算机科学学院) School of Politics and International Relations, University College Dublin(都柏林大学政治与国际关系学院) Department of Informatics, King’s College London(伦敦国王学院信息学院)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出一种评估议会辩论总结的框架,通过计算论证方法评估总结对辩论论证内容的忠实性,以欧洲议会辩论为例验证方法有效性。

Comments Accepted at KR'26 In The Wild Track. Camera Ready with additional supplementary materials

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21867 2026-06-23 cs.AI cs.CL cs.SC 新提交 87%

ForEx: A Formal Verification Framework for Explainable Reasoning in Logical Fallacy Detection and Annotation

ForEx:逻辑谬误检测与可解释推理的形式验证框架

Pei-Cing Huang, Chienyu Liu, Chan Hsu, Ci-Siang Chen, Pei-Ju Lee, Yihuang Kang

机构 * Department of Information Management(信息管理系) National Sun Yat-Sen University(国立中山大学) National Chung Hsing University(中国科技大学)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出ForEx框架,将LLM生成的解释翻译为Lean4并验证其形式可推导性,通过论证验证矩阵区分标签一致性与形式验证状态,揭示形式可推导性与标签一致性之间的系统性差距。

Comments 2026 IEEE 27th International Conference on Information Reuse and Integration for Data Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04815 2026-06-23 cs.LG cs.AI 版本更新 87%

Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents

边行动边学习:面向在线终身学习智能体的技能增强测试时协同进化框架

Bo Mao, Jie Zhou, Yutao Yang, Xin Li, Xian Wei, Qin Chen, Xingjiao Wu, Liang He

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) Shanghai AI Laboratory(上海人工智能实验室) Software Engineering Institute, East China Normal University(东华大学软件工程学院)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出LifeSkill框架,通过验证器引导的技能学习和在线技能内化,使LLM智能体在测试时持续内化反馈,提升终身学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20888 2026-06-23 cs.CV 新提交 87%

Fine-grained Human Motion Understanding with Language Models

基于语言模型的细粒度人体运动理解

Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew, Jan van Gemert, Xucong Zhang

机构 * Delft University of Technology(代尔夫特理工大学) Honda Research Institute Japan(日本本田研究所)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);language model(title)

AI总结 提出LLM模型\methodname,通过显式时间戳编码和多样化姿态/运动级监督,在多个基准上实现SOTA性能,且仅用2D骨骼输入即可超越3D方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23277 2026-06-23 cs.AI 新提交 86%

GIF: Locally Sound Geometric Information Flow Control for LLMs

GIF: 局部几何信息流控制用于大语言模型

Adam Storek, Nikolaus Holzer, Zhuo Zhang, Suman Jana

机构 * Columbia University(哥伦比亚大学)

专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出GIF框架,利用LLM雅可比矩阵和局部输出几何上界香农互信息,实现可扩展的信息流追踪,在提示注入和隐私泄露任务中达到近完美召回,且计算成本低。

详情

展开后加载摘要…

URL PDF HTML 收藏