arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15686 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15686 篇

1606.00424 2026-06-04 econ.GN nlin.AO q-fin.EC 80%

Residential income segregation: A behavioral model of the housing market

住宅收入隔离:住房市场的行为模型

Marco Pangallo, Jean Pierre Nadal, Annick Vignes

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文提出一个基于Agent的模型,研究收入隔离、收入不平等和房价之间的关系,通过分析买家和卖家的行为及价格形成机制,揭示收入分布不均对房价和隔离程度的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.07401 2026-06-04 cs.MA cs.SY eess.SY math.DS 80%

Asynchronous opinion dynamics on the $k$-nearest-neighbors graph

异步意见动力学在k-最近邻图上

Wilbert Samuel Rossi, Paolo Frasca

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 研究异步意见更新机制下的k-最近邻图意见动力学,发现其与传统模型有显著差异,证明当agent数量小于2k时动力学趋于共识。

Comments 17 pages, 4 figures, (to be) presented at the 57th IEEE Conference on Decision and Control, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.05660 2026-06-04 physics.soc-ph econ.GN q-fin.EC 80%

The role of consumer networks in firms' multi-characteristics competition and market-share inequality

消费者网络在企业多特征竞争和市场份额不平等中的作用

Antonios Garas, Athanasios Lapatinas

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文通过多特征空间中的位置分析模型,研究消费者网络对企业竞争和市场份额不平等的影响,采用动态 agent-based 分析,揭示多维产品差异化竞争中的适应性学习机制。

Comments 33 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.06959 2026-06-04 math.OC econ.GN q-fin.EC 80%

Strategic Growth with Recursive Preferences: Decreasing Marginal Impatience

具有递归偏好战略增长:边际不耐受递减

Luis Alcala, Fernando Tohme, Carlos Dabus

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文研究了资本积累两 agent 模型中策略、异质性和增长的相互作用,采用递归效用函数表示递减边际不耐受偏好,分析两种信息结构下的稳态均衡,探讨均衡的存在性和稳定性。

Comments 55 pages, 14 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.0355 2026-06-04 math.DS cs.MA cs.SY eess.SY 80%

On symmetric continuum opinion dynamics

关于对称连续意见动力学

Julien M. Hendrickx, Alex Olshevsky

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文研究了连续agent群体中常见意见动态模型的渐近行为,证明对称交互下意见分布收敛,并探讨更强意义上的收敛性。

Comments 28 pages, 2 figures, 3 files

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.3296 2026-06-04 cond-mat.stat-mech econ.GN q-fin.EC 80%

Endogenous crisis waves: a stochastic model with synchronized collective behavior

内生危机波:具有同步集体行为的随机模型

Stanislao Gualdi, Jean-Philippe Bouchaud, Giulia Cencetti, Marco Tarzia, Francesco Zamponi

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文提出一个简单框架,解释宏观经济Agent基于模型中常见的危机波现象,并验证其在其他物理或生物同步场景中的适用性,通过计算相图确认同步转换的鲁棒性。

Comments 5 pages, 3 figures. This paper is part of the CRISIS project, http://www.crisis-economics.eu

Journal ref Phys. Rev. Lett. 114, 088701 (2015)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02915 2026-06-03 cs.CV 80%

Any2Poster: Any-Source Poster Generation Across Modalities and Domains

Any2Poster: 跨模态和领域的任意源海报生成

Amogh Vinaykumar, Aiden Li, Suozhi Huang, Shilong Liu

机构 * Flower Mound High School(弗洛拉穆恩高中) University College London(伦敦大学学院) Princeton University(普林斯顿大学)

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 提出Any2Poster Bench基准和Any2Poster Agent智能体,实现从多种输入模态和领域生成海报,并通过基于测验和视觉评估的方法验证信息保真度和视觉传达效果。

Comments Project Page: https://github.com/Any2Poster/Any2Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10958 2026-06-02 cs.CV 80%

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

WorldLens:真实世界中驾驶世界模型的全光谱评估

Ao Liang, Lingdong Kong, Tianyi Yan, Hongsi Liu, Wesley Yang, Ziqi Huang, Wei Yin, Jialong Zuo, Yixuan Hu, Dekai Zhu, Dongyue Lu, Youquan Liu, Guangfeng Jiang, Linfeng Li, Xiangtai Li, Long Zhuo, Lai Xing Ng, Benoit R. Cottereau, Changxin Gao, Liang Pan, Wei Tsang Ooi, Ziwei Liu

机构 * WorldBench Team(WorldBench团队) Equal Contributions Project Lead(同等贡献项目负责人) Project Lead(项目负责人) Corresponding Author(通讯作者)

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 提出WorldLens基准,从生成、重建、动作跟随、下游任务和人类偏好五个方面评估生成世界模型在视觉真实性、几何一致性、物理合理性和功能可靠性上的表现,并构建WorldLens-26K数据集和WorldLens-Agent评估模型以实现可扩展的可解释评分。

Comments CVPR 2026 Oral Presentation; 80 pages, 37 figures, 29 tables; Project Page at https://worldbench.github.io/worldlens GitHub at https://github.com/worldbench/WorldLens

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09415 2026-05-19 econ.TH math.OC 80%

Contracting a crowd of heterogeneous agents

对异质群体的合同设计

Guillermo Alonso Alvarez, Erhan Bayraktar, Ibrahim Ekren

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文研究了大规模异质群体的最优合同设计问题,通过线性二次框架分析有限群体和连续极限问题,得到显式的最优合同和均衡努力,并证明了连续合同可以基于大样本评估以获得近似合同,从而在多 agent 交互场景中提供可扩展的近似方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02039 2026-05-19 cs.AI cs.CL cs.DB cs.LG 80%

Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models

在大型语言模型上进行深度数据研究:评估深度数据研究

Wei Liu, Peijie Yu, Michele Orini, Yali Du, Yulan He

专题命中 Agent评测 :agentic(abstract,abstract_cn);agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出深度数据研究(DDR)任务和DDR-Bench基准,评估大型语言模型的探索智能,发现有效探索需要内在策略而非单纯扩展。

Comments 14 pages, 7 tables, 8 figures, accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14723 2026-05-15 cs.AI cs.CL cs.LG 80%

Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model

通过与临床世界模型交互实现患者动态的智能体化

Minghao Wu, Yuting Yan, Zhenyang Cai, Ke Ji, Chuangsen Fang, Ziying Sheng, Xidong Wang, Rongsheng Wang, Hejia Zhang, Shuang Li, Benyou Wang, Hongyuan Zha

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 Agent评测 :agent(abstract);workflow(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出SepsisAgent,通过临床世界模型模拟患者对液体-血管加压药干预的响应,采用提出-模拟-细化流程进行脓毒症治疗推荐,优于传统RL和LLM基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26616 2026-04-30 cs.SI 80%

Impact of Attitude and Bounded Rationality on Collective Behavioral Transitions

态度与有限理性对集体行为转变的影响

Chen Song, Vladimir Cvetkovic, Angela Fontan, Rong Su, Karl H. Johansson

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文基于计划行为理论,提出动态agent模型,探讨态度和有限理性如何影响集体行为转变,并通过调整关键参数控制集体过渡。

Comments This work has been accepted for presentation at the 23rd IFAC World Congress

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09836 2026-04-20 cs.AI cs.CL cs.LG 80%

COMPOSITE-Stem

COMPOSITE-STEM基准:通过70个专家编写的任务评估AI代理的科学推理能力

Kyle Waters, Lucas Nuzzi, Tadhg Looram, Alessandro Tomasiello, Ariel Ghislain Kemogne Kamdoum, Bikun Li, Damien Sileo, Egor Kretov, Francesco Fournier-Facio, Georgios Soloupis, Haile Kassahun, Hew Wolff, Jiaqi Cai, Lianghui Li, Marc Roth, Mohinder Naiya, Naixu Guo, Qicheng Tang, Richard Wheeler, Samuele Sala, Serguei Popov, Steven Dillmann, Yuqi Li

机构 * PortexAI University of Milano-Bicocca(米兰-比科卡大学) University of Calgary(卡尔加里大学) University of Chicago(芝加哥大学) Inria(法国国家信息与自动化技术研究院) Fraunhofer Institute for Individualized Medical Technology IMTE(弗劳恩霍夫个性化医疗技术研究所) University of Cambridge(剑桥大学) Independent(独立) McGill University(麦吉尔大学) Massachusetts Institute of Technology(麻省理工学院) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院) Queen Mary University of London(伦敦大学玛丽女王学院) Dot Ingredients National University of Singapore(新加坡国立大学) Georgia Institute of Technology(佐治亚理工学院) University of Edinburgh(爱丁堡大学) Murdoch University(墨尔本大学) University of Porto(波尔图大学) Stanford University(斯坦福大学) Stony Brook University(史泰津布鲁克大学)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出COMPOSITE-STEM基准,通过70个专家编写任务评估AI代理的科学推理能力,结合精确匹配评分和LLM作为评委的协议,展示AI在科学领域的能力突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05681 2026-04-08 cs.AI cs.CL cs.GT cs.LG cs.MA 80%

LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo

LUDOBENCH:通过基于点的棋盘游戏场景评估LLM的行为决策

Ojas Jain, Dhruv Kumar

机构 * Department of Computer Science and Information Systems, BITS Pilani(比拉理工学院皮拉尼校区计算机科学与信息系统系)

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 LudoBench通过12种不同决策类别中的480个手工设计点场景,评估LLM在Ludo棋盘游戏中的战略推理能力,发现模型在行为上呈现不同特征,揭示了提示敏感性作为关键漏洞。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02081 2026-03-03 cs.DB cs.AI cs.CL cs.LG cs.MA 80%

GenDB: The Next Generation of Query Processing -- Synthesized, Not Engineered

GenDB:查询处理的下一代——合成而非工程

Jiale Lao, Immanuel Trummer

机构 * Cornell University(康奈尔大学)

专题命中 Agent评测 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 GenDB利用大语言模型合成查询执行代码,以替代传统复杂的查询处理引擎,实现更高效和定制化的查询处理系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20446 2026-02-25 cs.CR 80%

Understanding Human-AI Collaboration in Cybersecurity Competitions

理解网络安全竞赛中的人机协作

Tingxuan Tang, Nicolas Janis, Kalyn Asher Montague, Kevin Eykholt, Dhilung Kirat, Youngja Park, Jiyong Jang, Adwait Nadkarni, Yue Xiao

专题命中 Agent评测 :agent(abstract);AI agent(abstract);autonomous agent(abstract);tool use(abstract)

AI总结 研究探讨了网络安全竞赛中人类与AI协作的有效性,发现AI代理在自主引导提示和工具使用时能超越多数人类团队。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02455 2026-02-03 cs.AI cs.CL cs.SE 80%

Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction

Drift-Bench: 通过多轮交互诊断LLM代理在输入故障下的合作破裂

Han Bao, Zheyuan Zhang, Pengcheng Jing, Zhengqing Yuan, Kaiwen Shi, Yanfang Ye

机构 * University of Notre Dame(内布拉斯加大学)

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 Drift-Bench通过多轮交互评估LLM代理在输入故障下的合作破裂,采用角色驱动的用户模拟器和Rise评估协议,揭示澄清效果随用户角色和故障类型的变化,推动代理安全评估。

Comments 65 pages, 40 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03248 2025-11-11 cs.CR 80%

Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework

Junhao Li, Jiahao Chen, Zhou Feng, Chunyi Zhou

专题命中 Agent评测 :agent(abstract);workflow(abstract);agentic(abstract);multi-agent(abstract)

Comments 14 pages, 3 figures; Accepted by MMM 2026; Complete version in progress. Dataset available at https://huggingface.co/datasets/xaddh/multimodal-privacy

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17843 2025-10-22 cs.LG cs.AI cs.SE 80%

GRETEL: A Goal-driven Retrieval and Execution-based Trial Framework for LLM Tool Selection Enhancing

Zongze Wu, Yani Guo, Churong Liang, Runnan Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 Agent评测 :agent(abstract);workflow(abstract);agentic(abstract);分类 cs.AI、cs.LG、cs.SE

Comments 5 pages, 1 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06382 2025-09-09 cs.HC 80%

Context-Adaptive Hearing Aid Fitting Advisor through Multi-turn Multimodal LLM Conversation

Yingke Ding, Zeyu Wang, Xiyuxing Zhang, Hongbin Chen, Zhenan Xu

专题命中 Agent评测 :agent(abstract);workflow(abstract);agentic(abstract);multi-agent(abstract)

Comments Ubicomp Companion 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16416 2025-03-21 cs.AI cs.CL cs.LG 80%

Survey on Evaluation of LLM-based Agents

Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, Michal Shmueli-Scheuer

专题命中 Agent评测 :agent(abstract);tool use(abstract);planning(abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01907 2024-09-04 cs.HC 80%

Focus Agent: LLM-Powered Virtual Focus Group

Taiyu Zhang, Xuesong Zhang, Robbe Cools, Adalberto L. Simeone

专题命中 Agent评测 :agent(title,abstract)

Comments 8 pages, the 24th Intelligent Virtual Agent Conference

Journal ref Taiyu Zhang, Xuesong Zhang, Robbe Cools, and Adalberto Simeone. 2024. Focus Agent: LLM-Powered Virtual Focus Group. In ACM International Conference on Intelligent Virtual Agents (IVA '24), September 16--19, 2024, GLASGOW, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00798 2024-08-13 cs.LG cs.AI cs.CL cs.FL 80%

Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents

Zelong Li, Wenyue Hua, Hao Wang, He Zhu, Yongfeng Zhang

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1410.3334 2014-10-14 cs.MA 80%

DISARM: A Social Distributed Agent Reputation Model based on Defeasible Logic

Kalliopi Kravari, Nick Bassiliades

专题命中 Agent评测 :agent(title,abstract);multi-agent(comments)

Comments Paper under review. Keywords: Semantic Web, Intelligent Multi-agent Systems, Agent Reputation, Defeasible Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17237 2026-08-19 cs.CV cs.AI 新提交 79%

Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement

基于确定性几何与受约束智能体视觉-语言优化的结构平面图到模型转换

Mohammad Talebi-Kalaleh, Qipei Mei

机构 * University of Alberta(阿尔伯塔大学)

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI

AI总结 该研究提出首个无需特定任务检测器训练的结构平面图转模型框架,经100张图基准评估,各构件指标优异,可高效修正绘图错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17177 2026-08-19 cs.SE 新提交 79%

Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation

将AI智能体建立在契约基础上:对规范驱动测试生成的实证评估

Michele Tufano, James McClure, José Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivančić, Livio Dalloro, Pat Rondon

专题命中 Agent评测 :AI agent(title);agent(abstract);分类 cs.SE

AI总结 本研究针对LLM智能体生成测试时易遗漏契约相关边界的问题,提出规范驱动测试生成方法,经Google生产缺陷评估,其缺陷检测率、分支覆盖率均优于基线,测试套件质量也显著更高。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17029 2026-08-19 cs.SE cs.HC 新提交 79%

LadderTeam: Dual-Agent Laddering Elicitation Framework

LadderTeam:双智能体阶梯式引出框架

Manjushree Aithal, Alexander Kotz, James Mitchell

专题命中 Agent评测 :agent(title,abstract);分类 cs.SE

AI总结 LadderTeam是基于双智能体LLM架构的自动化UX线框图访谈框架,采用三种探测策略,经216次模拟访谈验证,实现高收敛率与可执行响应匹配率,且无主题偏离。

Comments 4 pages, 1 figure, 2 tables, Accepted in ACM AI Summit 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16919 2026-08-19 cs.IR cs.AI 新提交 79%

CARA: Cognitive Adaptive Recommendation Agent

CARA:认知自适应推荐智能体

Weijun Gao, Jinyang Dong, Chuanru Ren, Hengxiao Li

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 针对现有推荐方法未明确建模用户偏好转化为决策的局限,提出认知自适应推荐框架CARA,经亚马逊评论实验,其多数指标性能最优,较基线相对提升最高达10.15%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16211 2026-08-18 cs.AI 新提交 79%

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

BaT:构建具备阶段式评分标准的自演化医学研究智能体

Junqi Liu, Yufan He, Yexiao He, Pengfei Guo, Dong Yang, Andriy Myronenko, Can Zhao, Hanrong Ye, Tianhao Qi, Yuyin Zhou, Daguang Xu, Yucheng Tang

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) NVIDIA(英伟达)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 本文提出BaT系统,结合Stage Bank与BiCuRL方法,使医学研究智能体在AutoMedBench-Lite上得分大幅提升,BaT-9B性能优于Claude Opus 4.6。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14943 2026-08-18 cs.AI 新提交 79%

Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid

技能块:智能体应如何加载其技能?预加载、按需工具加载、渐进式披露与混合方式的缓存正确比较

Hironobu Nakasuji

机构 * Microsoft(微软)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 该研究比较四种智能体技能加载方法,发现Hybrid等方法可显著减少token使用,且无质量损失,条件加载在技能大部分无需每轮使用时最有益。

详情

展开后加载摘要…

URL PDF HTML 收藏