arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 93044 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 工具调用 5059 篇

2506.01062 2026-04-10 cs.CL cs.AI cs.LG 67%

SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models

SealQA: 提升搜索增强语言模型在事实性问题中的推理能力

Thinh Pham, Nguyen Nguyen, Pratibha Zunjare, Weiyuan Chen, Yu-Min Tseng, Tu Vu

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 工具调用 :agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 SealQA是一个新的挑战性基准,用于评估搜索增强语言模型在事实性问题上的能力,揭示当前模型在处理冲突、噪声或无用搜索结果时的局限性。

Comments Camera Ready version for ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06571 2026-04-09 cs.CL cs.AI cs.IR cs.LG 67%

LLM-based Schema-Guided Extraction and Validation of Missing-Person Intelligence from Heterogeneous Data Sources

基于大型语言模型的异构数据源中缺失人员情报提取与验证方案

Joshua Castillo, Ravi Mukkamala

机构 * Old Dominion University(欧道明大学)

专题命中 工具调用 :planning(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出Guardian Parser Pack,通过AI驱动的解析与规范化流程,将多源调查文档统一为符合模式的表示,提升缺失人员情报的提取与验证效率,同时展示系统架构及性能评估结果。

Comments 9 pages, 6 figures. Accepted at International Conference on Intelligent Digitization of Systems and Services (IDSS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24265 2026-03-31 cs.IR cs.HC 67%

Beyond the Click: A Framework for Inferring Cognitive Traces in Search

超越点击:一种从行为日志中推断认知轨迹的框架

Saber Zerhoudi, Michael Granitzer

专题命中 工具调用 :agent(abstract);multi-agent(abstract)

AI总结 本文提出一种基于信息觅食理论的框架,通过标注公开数据集生成认知标签,验证认知轨迹在行为特征薄弱时的显著提升效果,并开源工具和代码以支持未来认知感知用户模拟研究。

Journal ref Proceedings of the 48th European Conference on Information Retrieval (ECIR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13734 2026-03-27 cs.CL cs.AI cs.LG 67%

Instruction Following by Principled Boosting Attention of Large Language Models

通过原理性提升大语言模型的注意力来实现指令遵循

Vitoria Guardieiro, Avishree Khare, Adam Stein, Eric Wong

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 工具调用 :tool-use(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出InstABoost方法,通过提升指令相关注意力来增强指令遵循,避免了其他方法的缺陷,提升了指令引导与任务相关上下文的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23152 2026-03-25 cs.RO 67%

PHANTOM Hand

PHANTOM手

Teng Yan, Jiongxu Chen, Qixiang Hua, Yue Yu, Zihang Wang, Yaohua Liu, Bingzhuo Zhong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 工具调用 :tool-use(abstract);planning(abstract)

AI总结 PHANTOM手通过结合精确分析形状与鲁棒合规抓取,实现自由运动规划的亚度精度和连续可预测力输出,提升欠驱动手的负载与灵活性。

Comments 8 pages. Submitted to the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21706 2026-03-24 physics.med-ph 67%

Comprehensive Dosimetric Verification and Positional Sensitivity Analysis in Brachytherapy: A Unified ESAPI Tool for HDR and LDR Treatments

放射治疗中的全面剂量学验证与位置敏感性分析:一种用于 HDR 和 LDR 治疗的统一 ESAPI 工具

J. A. Valgoma

专题命中 工具调用 :planning(abstract);workflow(abstract)

AI总结 本文提出了一种基于 Varian Eclipse Scripting API 的独立软件工具,用于验证 HDR 和 LDR 放射治疗的 QA,通过比较点源和线源模型,分析位置不确定性,并提高临床工作流程的安全性。

Comments 13 pages, 2 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19952 2026-03-23 eess.SY cs.SY 67%

On the Capacity of Future Lane-Free Urban Infrastructure

未来无车道城市基础设施的容量

Patrick Malcolm, Klaus Bogenberger

专题命中 工具调用 :agent(abstract);multi-agent(abstract)

AI总结 本文通过分析与仿真方法探讨了未来无车道城市交通的潜在容量和空间效率,提出OptWULF方法实现了无车道交叉口的高效管理,验证了容量与街道宽度的连续关系及对称需求模式的适应性。

Comments 9 pages, 8 figures, submitted to IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05413 2026-03-18 cs.SD 67%

Building Enterprise Realtime Voice Agents from Scratch: A Technical Tutorial

从零构建企业级实时语音代理:技术教程

Jielin Qiu, Zixiang Chen, Liangwei Yang, Ming Zhu, Zhiwei Liu, Juntao Tan, Wenting Zhao, Rithesh Murthy, Roshan Ram, Akshara Prabhakar, Shelby Heinecke, Caiming Xiong, Silvio Savarese, Huan Wang

机构 * Salesforce AI Research(Salesforce AI研究院)

专题命中 工具调用 :agent(abstract);function calling(abstract)

AI总结 本文介绍如何从零构建企业级实时语音代理,通过Deepgram、vLLM和ElevenLabs实现端到端流程,展示最佳实践与代码实现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15653 2026-03-18 cs.CL cs.AI cs.LG 67%

Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context

递归语言模型与不确定性:自我反思程序搜索在长上下文中的意外效果

Keivan Alizadeh, Parshin Shojaee, Minsik Cho, Mehrdad Farajtabar

机构 * Apple(苹果公司)

专题命中 工具调用 :agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出SRLM框架,通过引入不确定性感知的自我反思,提升长上下文处理能力,实验显示其在多种任务中优于现有基线,且无需递归机制。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02239 2026-03-16 cs.DC 67%

The Role of A-priori Information in Networks of Rational Agents

在理性代理网络中先验信息的作用

Yehuda Afek, Yishay Mansour, Shaked Rafaeli, Moshe Sulamy

专题命中 工具调用 :agent(abstract);tool use(abstract)

AI总结 本文研究了在理性代理网络中,先验信息对均衡的影响,通过分析复制行为和先验分布,得出了在不同分布式计算问题中达到均衡所需的先验知识界限。

Comments This paper is the full version of the DISC 2018 paper. arXiv admin note: substantial text overlap with arXiv:1711.04728

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23109 2026-03-02 cs.RO 67%

Towards Intelligible Human-Robot Interaction: An Active Inference Approach to Occluded Pedestrian Scenarios

迈向可解释的人机交互:一种主动推断方法用于遮挡行人场景

Kai Chen, Yuyao Huang, Guang Chen

机构 * Tongji University(同济大学)

专题命中 工具调用 :agent(abstract);planning(abstract)

AI总结 本文提出基于主动推断的方法,用于解决遮挡行人场景中的安全挑战,通过结合 RBPF 和 CEM 增强的 MPPI 控制器,实现可解释的人机交互。

Comments 14 pages, 6 figures, Proceedings of the 2026 ACM/IEEE International Conference on Human-Robot Interaction (HRI'26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09496 2026-02-11 cs.HC 67%

Jokeasy: Exploring Human-AI Collaboration in Thematic Joke Generation

Jokeasy:探索主题笑话生成中的人机协作

Yate Ge, Lin Tian, Chiqian Xu, Luyao Xu, Meiying Li, Yuanda Hu, Weiwei Guo

专题命中 工具调用 :agent(abstract);workflow(abstract)

AI总结 Jokeasy通过人机协作提升主题笑话生成,结合搜索功能与双角色LLM代理,优化创意流程与素材整合。

Comments Accepted at IASDR 2025. This is the author-accepted version. Correspondence to first author: geyate@gmail.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01129 2026-02-03 cs.CR 67%

SMCP: Secure Model Context Protocol

SMCP: 安全模型上下文协议

Xinyi Hou, Shenao Wang, Yifan Zhang, Ziluo Xue, Yanjie Zhao, Cai Fu, Haoyu Wang

专题命中 工具调用 :workflow(abstract);agentic(abstract)

AI总结 SMCP通过统一身份管理、强认证和细粒度策略执行,提升智能体系统在工具调用中的安全性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01107 2026-02-03 cs.SE cs.AI cs.LG 67%

SPELL: Synthesis of Programmatic Edits using LLMs

通过LLM合成程序性编辑:SPELL

Daniel Ramos, Catarina Gamboa, Inês Lynce, Vasco Manquinho, Ruben Martins, Claire Le Goues

机构 * Carnegie Mellon University(卡内基梅隆大学) Carnegie Mellon University USA(卡内基梅隆大学(美国)) INESC-ID / IST - Universidade de Lisboa(INESC-ID / IST - 莱里斯本大学)

专题命中 工具调用 :agent(abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 SPELL通过LLM提取迁移示例并泛化为可重用的转换脚本,实现自动化API迁移。

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12272 2026-01-21 cs.CV 67%

AgenticPruner: MAC-Constrained Neural Network Compression via LLM-Driven Strategy Search

AgenticPruner: 通过LLM驱动的策略搜索实现MAC约束的神经网络压缩

Shahrzad Esmat, Mahdi Banisharif, Ali Jannesari

机构 * Iowa State University(爱荷华州立大学)

专题命中 工具调用 :agent(abstract);workflow(abstract)

AI总结 AgenticPruner通过LLM驱动策略搜索实现MAC约束的神经网络压缩,通过三个专门代理协调优化,提升收敛成功率并实现精确的MAC预算控制。

Comments 38 pages, 2 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09937 2026-01-16 cs.HC cs.IR 67%

From SERPs to Agents: A Platform for Comparative Studies of Information Interaction

从搜索结果页面到代理:一个用于信息交互比较研究的平台

Saber Zerhoudi, Michael Granitzer

专题命中 工具调用 :agent(abstract);autonomous agent(abstract)

AI总结 UXLab是一个开源平台,用于比较信息交互系统,通过无代码实验设计支持多模态交互研究。

Journal ref Proceedings of the 2026 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05887 2026-01-12 cs.CR 67%

Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense

网络空间AI:一种基于博弈论的AI用于引导攻击与防御

Víctor Mayoral-Vilches, María Sanz-Gómez, Francesco Balassone, Stefan Rass, Lidia Salas-Espejo, Benjamin Jablonski, Luis Javier Navarrete-Lozano, Maite del Mundo de Torres, Cristóbal R. J. Veas Chavez

专题命中 工具调用 :agent(abstract);agentic(abstract)

AI总结 本文提出G-CTR,一种基于博弈论的AI指导层,通过生成摘要引导攻击与防御行为,提升网络安全测试效率与成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03733 2026-01-08 cs.CV cs.AI cs.CL cs.CY cs.LG 67%

RadDiff: Describing Differences in Radiology Image Sets with Natural Language

RadDiff:用自然语言描述放射学图像集的差异

Xiaoxian Shen, Yuhui Zhang, Sahithi Ankireddy, Xiaohan Wang, Maya Varma, Henry Guo, Curtis Langlotz, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 工具调用 :agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 RadDiff通过多模态代理系统实现放射学图像集差异的自然语言描述,结合医学知识和多模态推理,在放射学研究配对中取得高准确率,推动临床影像分析的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10581 2025-12-23 cs.GR 67%

GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search

GraphTracer: LLM代理中基于图的故障追踪以实现鲁棒多轮深度搜索

Heng Zhang, Yuling Shi, Xiaodong Gu, Haochen You, Zijian Zhang, Lubin Gan, Yilei Yuan, Jin Huang

专题命中 工具调用 :agent(abstract);multi-agent(abstract)

AI总结 GraphTracer通过信息流分析和依赖图构建,提升多代理系统在多轮深度搜索中的故障归因准确性与鲁棒性。

Comments This submission has been withdrawn by the authors due to a fundamental error in the methodology that affects the validity of the main results

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09481 2025-12-03 cs.SE cs.AI cs.CL 67%

Evaluating LLMs on Sequential API Call Through Automated Test Generation

通过自动化测试生成评估LLMs的顺序API调用

Yuheng Huang, Jiayang Song, Da Song, Zhenlan Ji, Wenhan Wang, Shuai Wang, Lei Ma

机构 * The University of Tokyo(东京大学) Macau University of Science and Technology(澳门科技大学) Shandong University(山东大学) Hong Kong University of Science and Technology(香港科技大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Alberta(阿尔伯塔大学)

专题命中 工具调用 :tool use(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 本文提出StateGen框架,通过自动化测试生成评估LLMs在顺序API调用中的性能,构建了包含120个测试用例的StateEval基准测试,揭示了当前LLM在API整合方面的改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18335 2025-11-25 cs.CL cs.AI cs.LG 67%

OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas

OmniStruct: 跨多样的模式生成的通用文本到结构生成

James Y. Huang, Wenxuan Zhou, Nan Xu, Fei Wang, Qin Liu, Sheng Zhang, Hoifung Poon, Muhao Chen

机构 * University of Southern California(南加州大学) University of California, Davis(加州大学戴维斯分校) Microsoft Research(微软研究院)

专题命中 工具调用 :function calling(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 OmniStruct提出了一种跨多种模式的通用文本到结构生成方法,通过合成数据训练小型模型,实现与GPT-4o相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13428 2025-11-24 cs.RO 67%

VLM-SFD: VLM-Assisted Siamese Flow Diffusion Framework for Dual-Arm Cooperative Manipulation

VLM-SFD:基于视觉语言模型的双臂协作操作Siamese流扩散框架

Jiaming Chen, Yiyu Jiang, Aoshen Huang, Yang Li, Wei Pan

机构 * Department of Computer Science, The University of Manchester(计算机科学系,曼彻斯特大学) School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)

专题命中 工具调用 :tool use(abstract);planning(abstract)

AI总结 VLM-SFD通过双编码器-解码器架构和视觉语言模型,提升双臂协作操作的模仿学习效率与泛化能力。

Comments Accepted by IEEE RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09148 2025-11-19 cs.CL cs.AI cs.LG 67%

LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls

Kangning Zhang, Wenxiang Jiao, Kounianhua Du, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu

机构 * Shanghai Jiao Tong University(上海交通大学) Xiaohongshu Inc.(小红书公司)

专题命中 工具调用 :tool-use(abstract);分类 cs.AI、cs.CL、cs.LG

Comments The code is accessible at https://github.com/Rednote-DeepExperience/LoopTool. The LoopTool-8B is accessible at https://huggingface.co/zhuiguang-ning/LoopTool-8B

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11235 2025-11-04 cs.RO math.DS math.OC 67%

Ergodic exploration of dynamic distribution

Luka Lanča, Karlo Jakac, Sylvain Calinon, Stefan Ivić

机构 * Faculty of Engineering, University of Rijeka(里耶卡大学工程学院) Idiap Research Institute(伊迪亚普研究 institute)

专题命中 工具调用 :agent(abstract);multi-agent(abstract)

Comments Initial version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00086 2025-11-04 cs.LG cs.AI cs.CL 67%

Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph

Fali Wang, Jihai Chen, Shuhua Yang, Runxue Bao, Tianxiang Zhao, Zhiwei Zhang, Xianfeng Tang, Hui Liu, Qi He, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) University of Pittsburgh(匹兹堡大学) Amazon(亚马逊) Microsoft(微软)

专题命中 工具调用 :agent(abstract);分类 cs.AI、cs.CL、cs.LG

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02878 2025-10-31 cs.LG cs.AI cs.CL 67%

Language Models can Self-Improve at State-Value Estimation for Better Search

Ethan Mendes, Alan Ritter

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 工具调用 :agent(abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19405 2025-10-23 cs.CY 67%

Designing Knowledge Tools: How Students Transition from Using to Creating Generative AI in STEAM classroom

Qian Huang, Nachamma Sockalingam, Thijs Willems, King Wang Poon

专题命中 工具调用 :tool use(abstract);planning(abstract)

Comments to be published in IEEE TALE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16944 2025-10-21 cs.CY cs.SC 67%

Learning Ecology with VERA Using Conceptual Models and Simulations

Spencer Rugaber, Scott Bunin, Andrew Hornback, Sungeun An, Ashok Goel

专题命中 工具调用 :agent(abstract);tool use(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00784 2025-10-20 cs.IR cs.AI cs.CL cs.LG 67%

FIRE: Fact-checking with Iterative Retrieval and Verification

Zhuohan Xie, Rui Xing, Yuxia Wang, Jiahui Geng, Hasan Iqbal, Dhruv Sahnan, Iryna Gurevych, Preslav Nakov

机构 * MBZUAI The University of Melbourne(墨尔本大学)

专题命中 工具调用 :agent(abstract);分类 cs.AI、cs.CL、cs.LG

Comments 4 figures, 8 tables, accepted to Findings of NAACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14688 2025-10-01 cs.CL cs.AI cs.LG 67%

Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations

Mohammed Alkhowaiter, Norah Alshahrani, Saied Alshahrani, Reem I. Masoud, Alaa Alzahrani, Deema Alnuhait, Emad A. Alghamdi, Khalid Almubarak

机构 * Refine AI ASAS AI University of Bisha(比沙大学) University College London(伦敦大学学院) King Salman Global Academy for Arabic(萨勒曼全球阿拉伯学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) King Abdulaziz University(阿卜杜勒阿齐兹大学) HUMAIN

专题命中 工具调用 :function calling(abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏