arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-03-04 至 2026-03-04 共收录 108 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 多智能体 15 篇

2603.01404 2026-03-04 cs.RO 83%

D-GVIO: A Buffer-Driven and Efficient Decentralized GNSS-Visual-Inertial State Estimator for Multi-Agent Systems

D-GVIO: 一种基于缓冲的高效分布式GNSS-视觉-惯性状态估计器用于多智能体系统

Yarong Luo, Wentao Lu, Chi Guo, Ming Li

机构 * School of Robotics, Wuhan University(机器人学院,武汉大学) School of Electronic Information of Wuhan University, Hubei Luojia Laboratory(武汉大学电子信息学院,湖北珞珈实验室)

专题命中 多智能体 :agent(title);multi-agent(title)

AI总结 D-GVIO通过基于缓冲的分布式框架和L-IEKF实现高效鲁棒的GNSS-视觉-惯性状态估计,适用于多智能体系统的协同定位任务。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02711 2026-03-04 cs.AI 83%

A Natural Language Agentic Approach to Study Affective Polarization

一种自然语言代理方法研究情感极化

Stephanie Anneris Malvicini, Ewelina Gajewska, Arda Derbent, Katarzyna Budzynska, Jarosław A. Chudziak, Maria Vanina Martinez

机构 * Warsaw University of Technology (WUT)(华沙技术大学)

专题命中 多智能体 :agentic(title);agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 本文提出一种基于自然语言代理的多代理模型,用于研究社交媒体中的情感极化现象,通过虚拟社区模拟和实验分析,提供新的研究视角和方法。

Comments Accepted at ICAART 2026 (18th International Conference on Agents and Artificial Intelligence). The final published version is available in the conference proceedings (SCITEPRESS)

Journal ref In Proceedings of the 18th International Conference on Agents and Artificial Intelligence (ICAART 2026), Vol. 1, pp. 339-346. SCITEPRESS, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03175 2026-03-04 cs.AI 77%

Saarthi for AGI: Towards Domain-Specific General Intelligence for Formal Verification

Saarthi 用于 AGI:迈向形式验证的领域特定通用智能

Aman Kumar, Deepak Narayan Gadde, Luu Danh Minh, Vaisakh Naduvodi Viswambharan, Keerthan Kopparam Radhakrishna, Sivaram Pothireddypalli

机构 * 1 Infineon Technologies India Private Limited, India 2 Infineon Technologies Dresden AG \& Co. KG, Germany 3 Infineon Technologies Vietnam Company Ltd., Vietnam

专题命中 多智能体 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.AI

AI总结 Saarthi 通过结构化规则书和 RAG 技术提升形式验证的准确性和效率

Comments Published at the DVCon U.S. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12725 2026-03-04 cs.GT econ.TH stat.ML 67%

The Bounds of Algorithmic Collusion; $Q$-learning, Gradient Learning, and the Folk Theorem

算法合谋的边界;$Q$-学习、梯度学习与folk定理

Galit Askenazi-Golan, Domenico Mergoni Cecchelli, Edward Plumb, Clemens Possnig

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文研究了重复游戏中学习动态对算法合谋的影响,首次为多代理$Q$-学习算法提供了收敛性结果。

Comments This is a new version of a previous paper by the title "Reinforcement Learning, Collusion, and the Folk Theorem" by the three (alphabetically) first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02745 2026-03-04 cs.IT cs.AI cs.LG math.IT 62%

Enhancing User Throughput in Multi-panel mmWave Radio Access Networks for Beam-based MU-MIMO Using a DRL Method

通过基于DRL的方法增强多面板毫米波无线电接入网络中用户吞吐量的MU-MIMO

Ramin Hashemi, Vismika Ranasinghe, Teemu Veijalainen, Petteri Kela, Risto Wichman

机构 * Vismika Ranasinghe(维斯米卡·拉纳辛格) Teemu Veijalainen(泰穆·维贾拉西奈恩) Petteri Kela(彼得里·凯拉) Risto Wichman(里索·维希曼)

专题命中 多智能体 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于DRL的波束管理方法,通过优化波束选择提升多面板毫米波网络用户吞吐量并降低延迟。

Comments Accepted to the IEEE International Conference on Communications (ICC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02233 2026-03-04 cs.LG cs.AI 62%

Adaptive Personalized Federated Learning via Multi-task Averaging of Kernel Mean Embeddings

基于核均值嵌入的多任务平均的自适应个性化联邦学习

Jean-Baptiste Fermanian, Batiste Le Bars, Aurélien Bellet

专题命中 多智能体 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于核均值嵌入的多任务平均方法,实现自适应个性化联邦学习,无需预设数据异质性,自动切换全局与局部学习模式,并通过随机傅里叶特征降低通信成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02532 2026-03-04 cs.CV 50%

EIMC: Efficient Instance-aware Multi-modal Collaborative Perception

EIMC: 高效实例感知多模态协作感知

Kang Yang, Peng Wang, Lantao Li, Tianci Bu, Chen Sun, Deying Li, Yongcai Wang

机构 * School of Information, Renmin University of China(中国人民大学信息学院) Sony Research and Development Center China(索尼(中国)研发有限公司) National University of Defense Technology(国防科技大学)

专题命中 多智能体 :agent(abstract)

AI总结 EIMC通过实例感知的多模态协作感知方法,提升自动驾驶安全性,减少带宽使用,实现高效且准确的3D感知。

Comments 9 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 工作流自动化 8 篇

2603.02601 2026-03-04 cs.AI cs.SE 88%

AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows

AgentAssay: 用于非确定性AI代理工作流的令牌高效回归测试

Varun Pratap Bhardwaj

机构 * Independent Researcher(独立研究者)

专题命中 工作流自动化 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.SE

AI总结 AgentAssay通过行为指纹和自适应预算优化,实现非确定性AI代理工作流的高效回归测试,显著降低测试成本并提高检测能力。

Comments Technical Report. 52 pages, 5 figures, 9 theorems, 42 formal definitions. Zenodo DOI: 10.5281/zenodo.18842011

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03018 2026-03-04 cs.AI cs.SE 81%

REGAL: A Registry-Driven Architecture for Deterministic Grounding of Agentic AI in Enterprise Telemetry

REGAL:一种基于注册表的架构,用于企业遥测中代理AI的确定性接地

Yuvraj Agrawal

机构 * Adobe Inc.(Adobe公司)

专题命中 工作流自动化 :agentic(title,abstract);分类 cs.AI、cs.SE

AI总结 REGAL提出一种基于注册表的架构,用于企业遥测中代理AI的确定性接地,通过显式架构方法和语义编译提升确定性计算,解决LLM在私有遥测中的接地问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02586 2026-03-04 cs.AI 79%

LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges

LiveAgentBench: 104个现实挑战上代理系统全面评估基准

Hao Li, Huan Wang, Jinjie Gu, Wenjie Wang, Chenyi Zhuang, Sikang Bian

机构 * Ant Group(蚂蚁集团)

专题命中 工作流自动化 :agentic(title);AI agent(abstract);分类 cs.AI

AI总结 LiveAgentBench通过104个现实挑战全面评估代理系统,采用社会感知驱动的数据生成方法,提供374个任务用于验证和测试,揭示模型实际性能并识别改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02366 2026-03-04 cs.HC cs.AI 70%

PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XR

PlayWrite: 一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统

Esen K. Tütüncü, Qian Zhou, Frederik Brudy, George Fitzmaurice, Fraser Anderson

机构 * Institute of Neurosciences of the University of Barcelona(巴塞罗那大学神经科学研究所) Autodesk Research(Autodesk研究)

专题命中 工作流自动化 :agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 PlayWrite是一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统,通过直接操控虚拟角色和道具,结合多智能体AI管道和大型语言模型,促进高度即兴和游戏化的创作过程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02089 2026-03-04 cond-mat.mtrl-sci physics.chem-ph 50%

High-quality, high-information datasets for universal atomistic machine learning

高质量、高信息量的数据集用于通用原子级机器学习

Cesare Malosso, Filippo Bigi, Paolo Pegolo, Joseph W. Abbott, Philip Loche, Mariana Rossi, Michele Ceriotti, Arslan Mazitov

专题命中 工作流自动化 :workflow(abstract)

AI总结 MAD-1.5数据集通过统一的DFT工作流程和增强的化学空间覆盖,为原子级机器学习提供了高质量、高信息量的训练数据,支持广泛元素的高精度模拟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02569 2026-03-04 cs.HC 50%

An LLM-Assisted Toolkit for Inspectable Multimodal Emotion Data Annotation

一种辅助多模态情绪数据标注的工具包

Zheyuan Kuang, Weiwei Jiang, Nicholas Koemel, Matthew Ahmadi, Emmanuel Stamatakis, Benjamin Tag, Anusha Withana, Zhanna Sarsenbayeva

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出一种LLM辅助的多模态情绪数据标注工具包,通过可检查的工作流实现细粒度标注,提升跨模态一致性检查效率。

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02485 2026-03-04 stat.ME stat.AP 50%

A Decision Analysis Framework for High-fidelity and Low-fidelity Systems with Applications in Manufacturing Processes

面向高保真与低保真系统的决策分析框架及其在制造过程中的应用

Fan Zhang, Qiong Zhang, Madhura Limaye, Dhanashree Shinde, Gang Li, Sai Aditya Pradeep, Srikanth Pilla

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出基于多保真高斯过程的决策分析框架,用于在制造过程中平衡高保真与低保真数据,以实现高效优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02391 2026-03-04 physics.ins-det hep-ex nucl-ex 50%

Internal Charge Amplification in Germanium at 77K and 4K: From Single-Free-Flight Bounds to a Physics-Informed Ionization Model

锗在77K和4K时的内部电荷放大:从单自由飞行界限到物理引导的离子化模型

Dongming Mei, Kunming Dong, Narayan Budhathoki, Shasika Panamaldeniya, Francisco Ponce

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出了一种基于物理引导的离子化模型,用于预测低温下锗材料的临界电场,结合单自由飞行界限与传输特性,为设计稳定且可控的内部电荷放大系统提供指导。

Comments 18 pages, 6 figures, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 软件智能体 3 篇

2510.00857 2026-03-04 cs.CL 70%

ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs

ManagerBench: 评估自主大语言模型中的安全与务实之间的权衡

Adi Simhi, Jonathan Herzig, Martin Tutek, Itay Itzhak, Idan Szpektor, Yonatan Belinkov

机构 * Technion – Israel Institute of Technology(技术学院–以色列理工学院) Google Research(谷歌研究) University of Zagreb(Zagreb大学) Kempner Institute, Harvard University(哈佛大学凯普勒研究所)

专题命中 软件智能体 :autonomous agent(abstract);agentic(abstract);分类 cs.CL

AI总结 ManagerBench评估自主大语言模型在安全与务实权衡中的表现,揭示模型在冲突目标下的决策缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01012 2026-03-04 cs.SE cs.AI 62%

FastCode: Fast and Cost-Efficient Code Understanding and Reasoning

FastCode: 快速且成本有效的代码理解和推理

Zhonghang Li, Zongwei Li, Yuxuan Chen, Han Shi, Jiawei Li, Jierun Chen, Haoli Bai, Chao Huang

机构 * The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 软件智能体 :agentic(abstract);分类 cs.AI、cs.SE

AI总结 FastCode通过结构化勘探机制提升代码推理效率,减少计算成本,优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12967 2026-03-04 cs.HC 50%

Mining Hierarchies with Conviction: Constructing the CS1 Skill Hierarchy with Pairwise Comparisons over Skill Distributions

通过置信度挖掘层次结构:基于技能分布的配对比较构建CS1技能层次结构

Dip Kiran Pradhan Newar, Max Fowler, David H. Smith, Seth Poulsen

专题命中 软件智能体 :planning(abstract)

AI总结 通过置信度度量分析编程技能的先决关系,构建CS1技能层次结构,揭示技能间的依赖性及教学优化策略。

Comments Published in the Taylor & Francis Journal "Computer Science Education"

Journal ref Computer Science Education (2025) 1-22

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与网页智能体 4 篇

2505.13909 2026-03-04 cs.AI cs.CL cs.LG 82%

Efficient Agent Training for Computer Use

高效计算机使用代理的训练

Yanheng He, Jiahe Jin, Pengfei Liu

机构 * Shanghai Jiao Tong University(上海交通大学) SII GAIR

专题命中 GUI与网页智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出PC Agent-E,通过结合人类标注数据与Claude 3.7 Sonnet生成的数据,实现了计算机使用代理的高效训练,取得显著性能提升。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02772 2026-03-04 cs.RO 78%

Agentic Self-Evolutionary Replanning for Embodied Navigation

代理自进化重规划用于具身导航

Guoliang Li, Ruihua Han, Chengyang Li, He Li, Shuai Wang, Wenchao Ding, Hong Zhang, Chengzhong Xu

机构 * Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系) SIAT, Chinese Academy of Sciences(中国科学院上海技术物理研究所) Department of Computer Science, University of Hong Kong(香港大学计算机科学系) Academy for Engineering & Technology, Fudan University(复旦大学工程与技术学院) Department of EEE, Southern University of Science and Technology(南方科技大学电子工程系)

专题命中 GUI与网页智能体 :agentic(title,abstract)

AI总结 SERP通过自进化动作模型和图链式思考重规划,提升具身导航在复杂环境中的鲁棒性和效率。

Comments 8 pages, 10 figures, 4 tables, submitted to IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03533 2026-03-04 cs.CL 57%

Go-Browse: Training Web Agents with Structured Exploration

Go-Browse: 通过结构化探索训练网络代理

Apurva Gandhi, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 GUI与网页智能体 :agent(abstract);分类 cs.CL

AI总结 Go-Browse通过结构化探索方法,利用大规模网络环境数据训练代理,提升任务解决成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02484 2026-03-04 cs.RO math.OC 50%

COLREGs Compliant Collision Avoidance and Grounding Prevention for Autonomous Marine Navigation

自主水运船舶的碰撞规避与搁浅预防:符合国际海上避碰规则

Mayur S. Patil, Nataraj Sudharsan, Veneela Ammula, Jude Tomdio, Jin Wang, Michael Kei, Sivakumar Rathinam, Prabhakar R. Pagilla

机构 * Department of Mechanical Engineering, Texas A&M University, College Station, USA(德克萨斯A&M大学机械工程系) American Bureau of Shipping, Spring, TX, USA(美国船舶局)

专题命中 GUI与网页智能体 :planning(abstract)

AI总结 本文提出了一种基于凸优化的运动规划方法,用于自主水运船舶的碰撞规避、COLREGs合规性和搁浅预防。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 记忆与上下文管理 7 篇

2603.03212 2026-03-04 cs.AI 79%

NeuroSkill(tm): Proactive Real-Time Agentic System Capable of Modeling Human State of Mind

NeuroSkill(tm):一种能够建模人类心理状态的前瞻性实时代理系统

Nataliya Kosmyna, Eugene Hauptmann

专题命中 记忆与上下文管理 :agentic(title,abstract);分类 cs.AI

AI总结 NeuroSkill(tm)通过整合脑机接口信号和SKILL.md描述,实现对人类心理状态的实时建模与多层面互动,采用开源协议并遵循伦理对齐AI标准。

Comments 36 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02626 2026-03-04 cs.AI 79%

See and Remember: A Multimodal Agent for Web Traversal

见与忆:一种用于网页浏览的多模态代理

Xinjun Wang, Shengyao Wang, Aimin Zhou, Hao Hao

机构 * Shanghai Institute of AI for Education(上海人工智能教育研究院) East China Normal University(华东师范大学)

专题命中 记忆与上下文管理 :agent(title,abstract);分类 cs.AI

AI总结 V-GEMS通过视觉 grounding 和显式记忆系统实现精确稳健的网页浏览,实验显示其在导航任务中性能提升28.7%

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02206 2026-03-04 cs.SD 78%

VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures

VoiceAgentRAG: 通过双代理架构解决实时语音代理中RAG的延迟瓶颈

Jielin Qiu, Jianguo Zhang, Zixiang Chen, Liangwei Yang, Ming Zhu, Juntao Tan, Haolin Chen, Wenting Zhao, Rithesh Murthy, Roshan Ram, Akshara Prabhakar, Shelby Heinecke, Caiming Xiong, Silvio Savarese, Huan Wang

机构 * Salesforce AI Research(Salesforce人工智能研究)

专题命中 记忆与上下文管理 :agent(title,abstract)

AI总结 VoiceAgentRAG通过双代理架构实现实时语音代理中RAG的延迟优化,利用预加载缓存提升检索效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03258 2026-03-04 cs.AI 74%

Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals

继承性目标漂移:情境压力可能削弱代理目标

Achyutha Menon, Magnus Saebo, Tyler Crosse, Spencer Gibson, Eyon Jang, Diogo Cruz

机构 * UC San Diego(加州大学圣地亚哥分校) Columbia University(哥伦比亚大学) Georgia Tech(佐治亚理工学院) MATS SPAR

专题命中 记忆与上下文管理 :agentic(title);分类 cs.AI

AI总结 本文研究了语言模型在面对情境压力时目标漂移的继承性问题,发现模型在不同环境下表现出不一致的鲁棒性,强调了对训练后技术改进的必要性。

Comments 22 pages, 7 figures. Accepted at ICLR 2026 Lifelong Agents Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02765 2026-03-04 cs.LG cs.AI 62%

Next Embedding Prediction Makes World Models Stronger

下一步嵌入预测使世界模型更强大

George Bredis, Nikita Balagansky, Daniil Gavrilov, Ruslan Rakhimov

专题命中 记忆与上下文管理 :agent(abstract);分类 cs.AI、cs.LG

AI总结 NE-Dreamer通过时间转换器预测下一步嵌入,无需解码器即可在复杂环境中提升基于模型的强化学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02640 2026-03-04 cs.CY cs.AI cs.CL cs.MA cs.SI 62%

Credibility Governance: A Social Mechanism for Collective Self-Correction under Weak Truth Signals

可信治理:在弱真相信号下的一种社会机制,用于集体自我校正

Wanying He, Yanxi Lin, Ziheng Zhou, Xue Feng, Min Peng, Qianqian Xie, Zilong Zheng, Yipeng Kang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Tsinghua University(清华大学) University of California, Los Angeles(加州大学洛杉矶分校) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)

专题命中 记忆与上下文管理 :agent(abstract);分类 cs.AI、cs.CL

AI总结 可信治理通过动态可信度评分和可信度加权背书,提升集体自我校正能力,减少虚假信息影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06743 2026-03-04 eess.SY cs.SY 50%

Data-Driven Control of Large-Scale Networks with Formal Guarantees: A Small-Gain Free Approach

基于正式保证的数据驱动控制大规模网络:一种无小增益自由的方法

Behrad Samari, Amy Nejati, Abolfazl Lavaei

专题命中 记忆与上下文管理 :agent(abstract)

AI总结 本文提出一种数据驱动方法,通过分而治之策略在未知网络中实现控制,无需传统小增益条件,显著降低样本复杂度。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. Agent评测 18 篇

2512.09882 2026-03-04 cs.AI cs.CR cs.CY 85%

Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing

将AI代理与网络安全专家在真实世界渗透测试中进行比较

Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Jun-shen Ho, Anna Wu, Arnold Tianyi Yang, Neil Perry, Andy Zou, Matt Fredrikson, J. Zico Kolter, Percy Liang, Dan Boneh, Daniel E. Ho

机构 * Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 Agent评测 :AI agent(title,abstract);agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 本文比较了AI代理与网络安全专家在真实世界渗透测试中的表现,发现ARTEMIS在技术深度和提交质量上接近最强人类参与者,但在误报率和GUI任务上存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏