arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-05-25 至 2026-05-25 共收录 151 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 多智能体 32 篇

2601.05427 2026-05-25 cs.GT 88%

Anytime Detection of Strategic Deviations in Multi-Agent Systems

多智能体系统中策略性偏差的任意时刻检测

Etienne Gauthier, Francis Bach, Michael I. Jordan

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 提出基于e值框架的序贯检验方法,通过构造测试超鞅实时检测多智能体系统行为是否偏离均衡基准,统一处理纳什、相关和粗相关均衡,并扩展至随机博弈。

Comments Code at: https://github.com/GauthierE/anytime-detection-deviation

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23099 2026-05-25 cs.MA 88%

SVR-MAD: A Bayesian-Inspired Framework for Posterior-Guided Multi-Agent Debate

SVR-MAD:一种贝叶斯启发的后验引导多智能体辩论框架

Weifan Jiang, Rana Shahout, Minghao Li, Zhenting Qi, Yilun Du, Michael Mitzenmacher, Minlan Yu

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 针对多智能体辩论中上下文快速增长导致可扩展性受限的问题,提出贝叶斯启发的SVR-MAD框架,利用辩论前后信号构建通信图,在降低令牌成本的同时保持或提升准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10743 2026-05-25 eess.SY cs.SY 88%

Scaling and Trade-offs in Multi-agent Autonomous Systems

多智能体自主系统中的规模与权衡

Abram H. Clark, Liraz Mudrik, Colton Kawamura, Nathan C. Redder, João P. Hespanha, Isaac Kaminer

专题命中 多智能体 :agent(title,abstract);multi-agent(title);planning(abstract)

AI总结 通过大规模仿真和量纲分析,揭示自主无人机蜂群在三种典型场景中的标度律,量化智能体数量与平台参数间的权衡,并展示最优路径规划对结果标度律的改善。

Comments 16 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22788 2026-05-25 cs.HC 88%

FACET: Multi-Agent AI Supporting Teachers in Scaling Differentiated Learning for Diverse Students

FACET: 支持教师为多样化学生扩展差异化学习的多智能体AI

Jana Gonnermann-Müller, Jennifer Haase, Nicolas Leins, Moritz Igel, Konstantin Fackeldey, Sebastian Pokutta

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 提出FACET教师端多智能体框架,通过模拟学习者、诊断评估、生成材料和评估四个智能体,在教师参与下支持差异化教学,解决课堂异质性带来的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17076 2026-05-25 cs.LG cs.AI cs.DC cs.MA 87%

S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination

S-Bus: 多智能体LLM状态协调的自动读集重建

Sajjad Khan

机构 * Sajjad Khan

专题命中 多智能体 :agent(title,abstract);multi-agent(title);分类 cs.AI、cs.LG

AI总结 提出S-Bus中间件,通过服务端DeliveryLog机制在提交时从HTTP GET流量中重建每个智能体的读集,实现可观测读隔离(ORI)一致性,防止专用分片拓扑中的结构性竞态条件,并通过形式化证明和实验验证其安全性。

Comments v2: LLM judge validated against human annotator (Zahid Hussain, Mindgigs Peshawar) on PH-3 at strict kappa=0.93 (n=93, 96.8% agreement); over-claim refined to 32% (LLM) / 49% (human). Adds Exp.PG-Comparison Rust-Native and Workload-B chi2=1094.98. 24 pages, 23 tables. Annotation data attached as arXiv ancillary files

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23320 2026-05-25 cs.AI 86%

Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning

人在回路的多智能体呼吸机决策支持与上下文赌博机偏好学习

Sijia Li, Xiaoyu Tan, Qixing Wang, Weiyi Zhao, Chen Zhan, Teqi Hao, Xuemin Wang, Lei Gu, Roland Eils, Xihe Qiu

机构 * Shanghai University of Engineering Science, Shanghai, China Tencent Youtu Lab, Tencent, China Department of Critical Care Medicine, Shanghai Tenth People's Hospital, Tongji University School of Medicine, Shanghai, China Department of Emergency Critical Disease, Songjiang Hospital Affiliated to Shanghai Jiao Tong University School of Medicine, Shanghai, China Max Planck Institute for Heart Lung Research, Bad Nauheim, Germany Fudan University, Shanghai, China BIH at Charit\'e -- Universit\"atsmedizin Berlin, Berlin, Germany

专题命中 多智能体 :agent(title,abstract);multi-agent(title);分类 cs.AI

AI总结 提出VDSS框架,通过人在回路的多智能体协作和上下文赌博机在线偏好适应,实现呼吸机决策支持的个性化与可审计性。

Comments miccai 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17905 2026-05-25 eess.SP 86%

Curriculum-Guided Heterogeneous Multi-Agent Intelligence for Multi-UAV Cooperative ISAC

课程引导的异构多智能体智能用于多无人机协同ISAC

Kang Yan, Luping Xiang, Kang Zheng, Jienan Chen, Jun Liu, Qiang Liu, Kun Yang

专题命中 多智能体 :agent(title,abstract);multi-agent(title)

AI总结 提出一种基于课程学习的异构近端策略优化算法,通过联合轨迹-波束赋形优化解决多无人机与地面基站协同ISAC中的后验克拉美-罗界最小化问题,实现感知性能提升30%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04431 2026-05-25 cs.LG cs.GT 85%

MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems

MaMa: 一种基于博弈论的安全智能体系统设计方法

Jonathan Nöther, Adish Singla, Goran Radanovic

机构 * Max Planck Institute for Software Systems (MPI-SWS)(马克斯·普朗克软件系统研究所)

专题命中 多智能体 :agentic(title,abstract);agent(abstract);multi-agent(abstract);分类 cs.LG

AI总结 提出MaMa算法,通过将安全智能体系统设计形式化为系统设计者与对手之间的Stackelberg博弈,利用基于LLM的对抗搜索自动生成鲁棒安全的系统设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23218 2026-05-25 cs.AI 85%

Foundation Protocol: A Coordination Layer for Agentic Society

Foundation Protocol: 智能体社会的协调层

Bang Liu, Yongfeng Gu, Jiayi Zhang, Zhaoyang Yu, Sirui Hong, Maojia Song, Xiaoqiang Wang, Mingyi Deng, Zijie Zhuang, Ronghao Wang, Mingzhe Cao, Yutong Zhu, Xingjian Li, Yifan Wu, Jianhao Ruan, Yiran Peng, Shuangrui Chen, Jinlin Wang, Yizhang Lin, Dongjie Zhang, Dekun Wu, Chen Ma, Lizi Liao, Han Yu, Jian Pei, Heng Ji, Qiang Yang, Yuyu Luo, Chenglin Wu

机构 * Singapore University of Technology and Design(新加坡科技设计大学) City University of Hong Kong(香港城市大学) Singapore Management University(新加坡管理学院) Nanyang Technological University(南洋理工大学) Duke University(杜克大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Hong Kong Polytechnic University(香港理工大学)

专题命中 多智能体 :agentic(title);agent(abstract);autonomous agent(abstract);multi-agent(abstract)

AI总结 提出Foundation Protocol(FP),一种以图为核心的协调层,用于统一异构实体、支持多方组织和基于事件的协作,并提供经济原语和治理机制,以构建开放、多元、可治理的人机社会。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12787 2026-05-25 cs.AI cs.MA 85%

Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics

Ax-Prover:用于数学和量子物理定理证明的深度推理智能体框架

Benjamin Breen, Marco Del Tredici, Jacob McCarran, Javier Aspuru Mijares, Weichen Winston Yin, Kfir Sulimany, Jacob M. Taylor, Frank H. L. Koppens, Dirk Englund

机构 * Axiomatic_AI(公理人工智能) Massachusetts Institute of Technology (MIT)(麻省理工学院) Institut de Ciències Fotòniques (ICFO)(光子科学研究所) Institució Catalana de Recerca i Estudis Avançats (ICREA)(加泰罗尼亚高级研究与高等学院)

专题命中 多智能体 :agentic(title,abstract);agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 提出多智能体系统Ax-Prover,通过将大语言模型与Lean工具结合,实现跨科学领域的自动化定理证明,并在抽象代数和量子理论新基准上显著超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05364 2026-05-25 cs.SE 85%

Survey of LLM Agent Communication with MCP: A Software Design Pattern Centric Review

LLM智能体通信与MCP综述:以软件设计模式为中心的回顾

Anjana Sarkar, Soumyendu Sarkar

专题命中 多智能体 :agent(title,abstract);agentic(abstract);multi-agent(abstract);分类 cs.SE

AI总结 本文综述了经典软件设计模式如何增强基于LLM的智能体AI系统中通信的可靠性和可扩展性,重点分析了模型上下文协议(MCP),并探讨了中介者、观察者、发布-订阅和代理等模式在MCP兼容框架中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23887 2026-05-25 cs.DB cs.AI cs.CR cs.LG cs.MA 85%

CHRONOS: Temporally-Aware Multi-Agent Coordination for Evolving Data Marketplaces

CHRONOS:面向演化数据市场的时态感知多智能体协调

Joydeep Chandra

机构 * BNRIST, Tsinghua University(北京清华大学智能机器人系统研究院)

专题命中 多智能体 :agent(title);multi-agent(title);分类 cs.AI、cs.LG

AI总结 提出CHRONOS三层架构,通过神经ODE时间衰减、基于变点的Shapley估值和EXP3-IX差分隐私协调,解决时态知识图谱数据市场中索引陈旧、定价偏差和隐私预算过度消耗问题,在四个基准上实现0.937召回率、2.74 qps和4.25总隐私预算。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22842 2026-05-25 cs.CR cs.AI cs.LG 84%

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

归因偏差:当记忆中毒在自主AI系统中看起来像模型失败时

Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder

机构 * Department of Computer Science, University of Texas at El Paso(德克萨斯大学埃尔帕索分校计算机科学系) School of Computing, Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校计算机学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多智能体 :agentic(title);agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

AI总结 本文识别了多智能体AI系统中的“归因偏差”问题,即记忆层攻击导致的行为与模型失败无法区分,并形式化了“语义规范漂移”(SND)作为第三种不当行为路径,提出了反事实组合测试和记忆持久信息流控制等防御方法。

Comments This paper is presently under review at a top-tier security venue

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22897 2026-05-25 cs.LG 77%

From Residuals to Reasons: LLM-Guided Mechanism Inference from Tabular Data

从残差到原因:基于LLM的表格数据机制推断

Mohammad R. Rezaei, Rahul G. Krishnan

机构 * Department of Computer Science(计算机科学系) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 多智能体 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.LG

AI总结 提出多智能体残差上下文学习框架(MARICL),通过让LLM分析基础模型残差并迭代生成修正项,在多个科学基准上提升预测性能并实现机制泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22841 2026-05-25 physics.soc-ph cs.AI cs.CL cs.GT cs.MA econ.GN q-fin.EC 76%

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

联盟内的战略胁迫:格陵兰主权博弈作为人工智能压力测试

Rommin Adl, Peyton Williams

机构 * Grinnell College(格里纳尔学院)

专题命中 多智能体 :agent(abstract,comments);multi-agent(abstract,comments);分类 cs.AI、cs.CL

AI总结 通过多智能体模拟和逆向博弈论,研究美国对格陵兰的领土胁迫行为,揭示前沿大语言模型在联盟博弈中的升级倾向、模型间差异及和平解决路径。

Comments 78 pages, 17 figures, 18 tables. Multi-agent LLM simulation recovering structural utility parameters across 8 frontier models in the Greenland sovereignty crisis. v3: typo pass, fixes phantom action names (REQUEST_MULTILATERAL, INDEPENDENT) and a Blunden date mismatch. v2 added Section V safety findings (legitimacy-laundered escalation, signal decoupling) and Appendix H

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12316 2026-05-25 cs.AI cs.CL cs.CY cs.GT cs.MA 73%

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

GT-HarmBench:通过博弈论视角评估AI安全风险

Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin

机构 * ETH Zürich(苏黎世联邦理工学院) Berea College(贝雷学院) University of Toronto(多伦多大学) Vector Institute(向量研究所) Max Planck Institute for Intelligent Systems, Tübingen, Germany(图宾根德国智能系统马克斯·普朗克研究所)

专题命中 多智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.CL

AI总结 提出GT-HarmBench基准,包含1535个高风险场景,基于博弈论结构(如囚徒困境、猎鹿博弈、斗鸡博弈)评估前沿AI模型在多智能体环境中的安全风险,发现模型在38%的高风险案例中未能选择对社会有益的行为,并验证了博弈论干预可提升18%的社会有益结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23754 2026-05-25 cs.LG 70%

LLM-driven design of physics-constrained constitutive models: two agents are better than one

LLM驱动的物理约束本构模型设计:两个智能体胜过一个

Marius Tacke, Matthias Busch, Kian Abdolazizi, Jonas Eichinger, Kevin Linka, Roland Aydin, Christian Cyron

机构 * Helmholtz-Zentrum Hereon(海德堡中心) Hamburg University of Technology(汉堡技术大学) RWTH Aachen University(亚琛工业大学) Saarland University(萨尔兰州大学) German Center for Artificial Intelligence(德国人工智能中心)

专题命中 多智能体 :agent(abstract);multi-agent(abstract);分类 cs.LG

AI总结 提出多智能体LLM框架,通过Creator生成模型和Inspector审计物理约束,实现自动生成满足物理定律的本构模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23864 2026-05-25 math.OC cs.SY eess.SY 67%

Harnessing Individual Motivation for Collective Efficiency: A Mechanism-Driven Distributed Optimization Method

利用个体动机实现集体效率:一种机制驱动的分布式优化方法

Dongwei Xie, Xuhao Wang, Yujie Tang, Jie Song

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 针对多智能体集体决策中个体自利与全局性能冲突的问题,提出一种机制驱动的分布式优化方法,通过设计影子定价和VCG机制激励个体参与分布式协作,并保证算法收敛性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23356 2026-05-25 eess.SY cs.SY 67%

A Distributed Framework for Data-Driven Safe Coordination in Leader-Follower Networks

一种用于领导者-跟随者网络中数据驱动安全协调的分布式框架

Mirhan Urkmez, Maryam Sharifi, Shahab Heshmati-Alamdari

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文提出分布式数据驱动零化控制屏障函数(3D-ZCBF)框架,通过从输入-状态数据中识别导数边界,在无需显式模型的情况下保证领导者-跟随者多智能体系统的连通性。

Comments Submitted to IEEE TCNS

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23027 2026-05-25 cs.RO 67%

PIMbot: A Self-Adaptive Attack Framework for Adversarial Manipulation of Multi-Robot Reinforcement Learning

PIMbot:一种用于多机器人强化学习对抗性操纵的自适应攻击框架

Zexin Li, Ziliang Zhang, Hyoseung Kim, Cong Liu

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 提出PIMbot框架,通过奖励激励操纵和策略操纵两种互补杠杆,并采用自适应多目标控制器在线平衡,有效操纵多机器人社会困境环境,实验验证了其在仿真和真实嵌入式系统中的有效性。

Comments Extension version of IROS'23

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12337 2026-05-25 econ.TH 67%

Artificial Intelligence in Team Dynamics: Who Gets Replaced and Why?

团队动态中的人工智能:谁被取代以及为什么?

Xienan Cheng, Mustafa Dogan, Pinar Yildirim

专题命中 多智能体 :AI agent(abstract);workflow(abstract)

AI总结 本文通过序贯团队生产模型,研究最优AI部署策略,发现AI会随机取代团队两端工人而非中间工人,且可能未充分利用AI容量,同时提高平均工资并降低工资不平等。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01524 2026-05-25 cs.MA cs.SI math.OC 67%

Cost-Aware Distributed Online Learning with Strict Rejection Behavior against Adversarial Agents

具有严格拒绝对抗智能体行为的成本感知分布式在线学习

Yuhan Suo, Runqi Chai, Senchun Chai, Xudong Zhao, Yuanqing Xia

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 针对物联网多智能体系统中恶意干扰导致的演化不同步问题,提出一种两层自适应调节框架,外层动态调整长期演化率,内层保持鲁棒在线学习,实现低成本的分布式在线学习。

Comments 13 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23263 2026-05-25 cs.RO cs.AI cs.SY eess.SP eess.SY 57%

6G Communication Networks Enabling Embodied Agents: Architecture and Prototype

6G通信网络赋能具身智能体:架构与原型

Lipeng Dai, Luping Xiang, Kun Yang

机构 * State Key Laboratory of Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Institute of Intelligent Networks and Communications (NINE), Nanjing University (Suzhou Campus)(南京大学智能网络与通信研究所(苏州校区))

专题命中 多智能体 :agent(abstract);分类 cs.AI

AI总结 本文提出一种分层通信架构(包括人类意图感知层、基于O-RAN的传输层、智能中间层和具身层),并通过集成触觉设备、工业机械臂、中间平台和5G O-RAN测试床的原型验证了毫秒级延迟和稳定闭环操作。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 工作流自动化 22 篇

2605.23636 2026-05-25 eess.SY cs.SY 88%

RF Instrument Agent (RFIA): Empowering RF Instruments with Natural Language Understanding, Scheduling and Execution of Complex Tasks

RF仪器智能体 (RFIA): 赋予射频仪器自然语言理解、调度与执行复杂任务的能力

Chunhui Li, Wei Fan

专题命中 工作流自动化 :agent(title,summary_cn);planning(abstract);workflow(abstract)

AI总结 提出RFIA框架,通过解耦意图-规划-执行架构和结构化知识库,实现基于自然语言的可靠射频仪器控制,并在商用矢量网络分析仪上验证了多任务自动化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22883 2026-05-25 cs.AI cs.LG cs.PF 84%

Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems

每个成功目标的能量:面向智能体AI系统的目标级能量核算

Deepak Panigrahy, Aakash Tyagi

机构 * Independent Researcher(独立研究者) Texas A\&M University(德克萨斯A&M大学) Texas A\&M University Department of Computer Science(德克萨斯A&M大学计算机科学系)

专题命中 工作流自动化 :agentic(title,abstract);workflow(abstract);分类 cs.AI、cs.LG

AI总结 提出A-LEMS框架,将AI能量核算单位从每次推理能量转换为每个成功目标能量(EpG),并定义编排开销指数(OOI),实验表明智能体工作流平均能耗是线性基线的4.33倍,且能耗主要由编排结构而非推理计算驱动。

Comments 34 pages, 16 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06404 2026-05-25 cs.AI cond-mat.mtrl-sci physics.chem-ph 83%

GENIUS: An Agentic AI Framework for Autonomous Design and Execution of Simulation Protocols

GENIUS: 一种用于自主设计和执行模拟协议的智能AI框架

Mohammad Soleymanibrojeni, Roland Aydin, Diego Guedes-Sobrinho, Alexandre C. Dias, Maurício J. Piotrowski, Wolfgang Wenzel, Celso Ricardo Caldeira Rêgo

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Institute of Nanotechnology(纳米技术研究所) Hamburg University of Technology(汉堡技术大学) Federal University of Paraná(帕拉纳联邦大学) University of Brasília(巴西利亚大学) Federal University of Pelotas(普拉多斯联邦大学) Institute of Physics and International Center of Physics(物理研究所和国际物理中心)

专题命中 工作流自动化 :agentic(title,abstract);workflow(abstract);分类 cs.AI

AI总结 提出GENIUS框架,融合量子ESPRESSO知识图谱与分层大语言模型,通过有限状态错误恢复机器实现从自由文本到验证输入文件的自动生成,在295个基准测试中达到约80%的成功率,并大幅降低推理成本与幻觉。

Journal ref Communications Materials 7, 115 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23043 2026-05-25 cs.CL stat.ML 80%

HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation

HawkesLLM:智能体文本模拟中的语义不确定性传播

Zewei Deng, Tinghan Ye, Liyan Xie

机构 * Department of Industrial and Systems Engineering, University of Minnesota(工业与系统工程系,明尼苏达大学) H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology(H. Milton Stewart工业与系统工程学院,佐治亚理工学院)

专题命中 工作流自动化 :agentic(title,abstract);分类 cs.CL

AI总结 提出HawkesLLM框架,通过多变量Hawkes过程建模时间影响与文本生成分离,解决智能体文本模拟中语义不确定性路径依赖问题,在GDELT新闻级联案例中提升后期语义对齐。

Comments 10 pages, 4 figures, Accepted at the ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23121 2026-05-25 physics.plasm-ph physics.comp-ph 78%

NIMROD-to-IMAS workflow for extended-magnetohydrodynamic data with reusable datasets and implications for IMAS schema development

面向可重用数据集的扩展磁流体动力学数据的NIMROD-to-IMAS工作流及其对IMAS模式开发的启示

Alexei Y. Pankin, Fatima Ebrahimi, Qian Gong, Jacob King, Andreas Kleiner, Jesus Dominguez-Palacios, Norbert Podhorszki, Eric Suchyta

专题命中 工作流自动化 :workflow(title,abstract)

AI总结 本文提出一个将NIMROD代码输入输出转换为ITER IMAS数据字典兼容记录的工作流,解决了跨代码验证/耦合和数据库规模分析的难题,并验证了关键数据的守恒性,同时指出了IMAS框架的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23271 2026-05-25 cs.CV cs.AI 77%

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

EvalVerse:面向专业电影级视频生成的流水线感知与专家校准基准测试

Songlin Yang, Haobin Zhong, Ruilin Zhang, Xiaotong Zhao, Shuai Li, Kai Zheng, Xuyi Yang, Zhe Wang, Zhenchen Tang, Yang Li, Bohai Gu, Zhengwei Peng, Yidan Huang, Mengzhou Luo, Yihang Bo, Dalu Feng, Yujia Zhang, Juntao Ma, Ruiqi Wang, Lvmin Zhang, Yuwei Guo, Frank Guan, Maneesh Agrawala, Hongbo Fu, Alan Zhao, Anyi Rao

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Tencent(腾讯) Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Film Academy(北京电影学院) Stanford University(斯坦福大学) The Chinese University of Hong Kong(香港中文大学) Singapore Institute of Technology(新加坡理工学院)

专题命中 工作流自动化 :agent(abstract);workflow(abstract);agentic(abstract);分类 cs.AI

AI总结 提出EvalVerse框架,通过将电影制作专业知识系统化为评估分类法、构建专家标注数据集并微调视觉语言模型进行链式推理,解决了现有基准忽视电影质量、缺乏领域特异性指标的问题,实现了对生成视频“正确性”与“优良性”的全面评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23204 2026-05-25 cs.AI 77%

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

AutoResearch AI:迈向人工智能驱动的科研自动化以实现科学发现

Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu, Pan Zhou, Yongchao Chen, Ran Xu, Lifang He, Qingsong Wen, Manling Li, Cong Lu, Shuai Li, Pengtao Xie, Yixuan Yuan, Rui Meng, Lei Xing, Lichao Sun, Caiming Xiong, Philip S. Yu, Jianfeng Gao

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学) Tsinghua University(清华大学) Wuhan University(武汉大学) Salesforce Research(Salesforce研究) Squirrel AI Learning(Squirrel AI学习) Northwestern University(西北大学) Independent(独立) Shanghai Jiao Tong University(上海交通大学) University of California San Diego(加州大学圣地亚哥分校) Chinese University of Hong Kong(香港中文大学) University of Illinois Chicago(伊利诺伊大学香槟分校) Stanford University(斯坦福大学) Google Cloud AI Research(谷歌云AI研究) Recursive Superintelligence(递归超级智能) Microsoft Research(微软研究院)

专题命中 工作流自动化 :tool use(abstract);planning(abstract);workflow(abstract);分类 cs.AI

AI总结 本文综述了AI驱动的科研工作流自动化(AutoResearch)的发展,分析了从任务级AI到工作流级研究自动化的转变,并提出了五个评估维度(新颖性、有效性、影响力、可靠性和溯源),指出自主性受领域条件限制。

Comments 49 pages, 12 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏