arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-05-04 至 2026-05-04 共收录 97 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 20 篇

2605.00420 2026-05-04 cs.MA cs.LG q-fin.GN 57%

Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents

前瞻性竞技场:一种去中心化的链上基准,用于评估AI预测代理

Maksym Nechepurenko, Pavel Shuvalov

机构 * Director of Research, Devnull(研发总监,Devnull) Chief Technology Officer, Devnull(首席技术官,Devnull)

专题命中 Agent评测 :AI agent(abstract);分类 cs.LG

AI总结 本文提出Foresight Arena,首个去中心化链上基准,用于评估AI预测代理在真实世界预测市场中的能力,通过概率预测和智能合约评估,结合Brier分数和Alpha分数衡量性能。

Comments 27 pages, 5 figures, 10 tables. Project page: https://foresightarena.xyz/. Code: https://github.com/foresight-arena/contracts

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00400 2026-05-04 cs.IR cs.CL 57%

FollowTable: A Benchmark for Instruction-Following Table Retrieval

FollowTable:一项指令遵循表格检索的基准测试

Rihui Jin, Yuchen Lu, Ting Zhang, Jun Wang, Kuicai Dong, Zhaocheng Du, Dongping Liu, Gang Wang, Yong Liu, Guilin Qi

机构 * Southeast University(东南大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education(新一代人工智能技术及其交叉应用国家重点实验室(东南大学)) Its Interdisciplinary Applications (Southeast University), Ministry of Education(交叉应用(东南大学),教育部)

专题命中 Agent评测 :agentic(abstract);分类 cs.CL

AI总结 本文提出了一项新的表格检索任务IFTR,要求模型同时满足主题相关性和细粒度指令约束。通过引入FollowTable基准测试,评估模型在遵循指令方面的表现,发现现有检索模型在处理表格数据时存在系统性偏差。

Comments SIGIR 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00281 2026-05-04 cs.LG cs.MA math.OC 57%

High-Probability Convergence in Decentralized Stochastic Optimization with Gradient Tracking

去中心化随机优化中带有梯度跟踪的高概率收敛性

Aleksandar Armacki, Haoyuan Cai, Ali H. Sayed

专题命中 Agent评测 :agent(abstract);分类 cs.LG

AI总结 本文研究去中心化随机优化中的高概率收敛性,提出结合梯度跟踪技术的DSGD算法,证明其在非凸和Polyak-Łojasiewicz成本下的最优收敛率,且在放松的子高斯条件下实现高概率收敛。

Comments 49 pages, 4 figures. arXiv admin note: text overlap with arXiv:2510.06141

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00074 2026-05-04 cs.CY cs.AI 57%

Adoption and Use of LLMs at an Academic Medical Center

学术医疗中心中大语言模型的采用与使用

Nigam H. Shah, Nerissa Ambers, Abby Pandya, Timothy Keyes, Juan M. Banda, Srikar Nallan, Carlene Lugtu, Artem A. Trotsyuk, Suhana Bedi, Alyssa Unell, Miguel Fuentes, Francois Grolleau, Sneha S. Jain, Jonathan Chen, Devdutta Dash, Danton Char, Aditya Sharma, Duncan McElfresh, Patrick Scully, Vishanthan Kumar, Clancy Dennis, Connor OBrien, Satchi Mouniswamy, Elvis Jones, Krishna Jasti, Gunavathi Mannika Lakshmanan, Sree Ram Akula, Varun Kumar Singh, Ramesh Rajmanickam, Sudhir Sinha, Vicky Zhou, Xu Wang, Bilal Mawji, Joshua Ge, Wencheng Li, Travis Lyons, Jarrod Helzer, Vikas Kakkar, Ramesh Powar, Darren Batara, Cheryl Cordova, William Frederick, Olivia Tang, Phoebe Morgan, April S. Liang, Stephen P. Ma, Shivam Vedak, Dong-han Yao, Akshay Swaminathan, Mehr Kashyap, Brian Ng, Jamie Hellman, Nikesh Kotecha, Christopher Sharp, Gretchen Brown, Christian Lindmark, Anurang Revri, Michael A. Pfeffer

机构 * Stanford Medicine, Technology and Digital Solutions(斯坦福医学院技术与数字解决方案)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI

AI总结 本文提出ChatEHR系统,通过整合患者多年医疗记录,实现LLM在临床文档中的自动化与交互应用,提升医疗效率与数据准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08597 2026-05-04 cs.GT cs.AI cs.MA econ.TH 57%

Markets with Heterogeneous Agents: Dynamics and Survival of Bayesian vs. No-Regret Learners

异质主体市场:贝叶斯学习者与无遗憾学习者的动态与生存

David Easley, Yoav Kolumbus, Eva Tardos

机构 * Cornell University(康奈尔大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI

AI总结 研究比较贝叶斯学习者与无遗憾学习者在资产市场中的表现,探讨其生存条件及市场主导因素,提出混合策略提升鲁棒性与适应性。

Comments Learning in Markets, Heterogeneous Agents, Regret and Survival, Bayesian Learning, No-Regret Learning, Portfolio Optimization, Kelly Rule, Distribution Shifts, Robust Bayesian Updates

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26858 2026-05-04 physics.soc-ph 50%

A well-motivated model of pedestrian dynamics

一种有说服力的行人动力学模型

Ezel Üsten, Anna Sieben, Mohcine Chraibi, Armin Seyfried

专题命中 Agent评测 :agent(abstract)

AI总结 本文提出基于期望-价值理论的动态动机模型,通过考虑目标接近度、相对位置和个体目标重要性,动态调节行人移动参数,模拟预瓶颈等待场景,揭示人群自我组织的异质性特征。

Comments 35 pages, 12 figures, 2 tables. Manuscript prepared for submission to a Springer journal. Using JuPedSim 1.4.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07041 2026-05-04 cs.DC 50%

AIReSim: A Discrete Event Simulator for Large-scale AI Cluster Reliability Modeling

AIReSim:用于大规模AI集群可靠性建模的离散事件模拟器

Karthik Pattabiraman, Mihir Patel, Fred Lin

专题命中 Agent评测 :planning(abstract)

AI总结 本文提出AIReSim,用于评估大规模AI集群中故障、恢复、调度和修复过程中的设计选择,通过系统性评估参数对整体可靠性的影响,帮助优先改进系统。

Comments To appear in the Industry track of the IEEE/IFIP International Conference on Dependable Systems and Networks (DSN), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏