arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15662 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15662 篇

1811.07871 2018-11-20 cs.LG cs.AI cs.NE stat.ML 81%

Scalable agent alignment via reward modeling: a research direction

Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.05875 2017-08-22 cs.NI cs.AI cs.MA cs.SE nlin.AO 81%

A novel agent-based simulation framework for sensing in complex adaptive environments

Muaz A. Niazi, Amir Hussain

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.SE

Comments 8 pages

Journal ref IEEE Sensors Journal 11.2 (2011): 404-412

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.04079 2017-01-17 cs.LG cs.AI 81%

Agent-Agnostic Human-in-the-Loop Reinforcement Learning

David Abel, John Salvatier, Andreas Stuhlmüller, Owain Evans

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG

Comments Presented at the NIPS Workshop on the Future of Interactive Learning Machines, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1306.5884 2014-01-03 cs.AI cs.HC cs.LG 81%

Design of an Agent for Answering Back in Smart Phones

Sandeep Venkatesh, Meera V Patil, Nanditha Swamy

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG

Comments This paper has been withdrawn by the author due to a crucial sign erro

详情

展开后加载摘要…

URL PDF HTML 收藏
1207.3760 2012-07-17 cs.NE cs.AI cs.LG nlin.AO 81%

Towards a Self-Organized Agent-Based Simulation Model for Exploration of Human Synaptic Connections

Önder Gürcan, Carole Bernon, Kemal S. Türker

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG

Comments 4 pages, 1 figure, 2nd Computer Science Student Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0605031 2009-12-01 cs.AI cs.MA cs.SE 81%

On the Design of Agent-Based Systems using UML and Extensions

Mihaela Dinsoreanu, Ioan Salomie, Kalman Pusztai

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12338 2026-07-15 cs.AI 新提交 80%

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

多少任务足以做出智能体基准决策?对公共大语言模型智能体基准的重放分析

Wei-Jung Huang

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI;agentic(comments)

AI总结 研究智能体基准测试中多少任务足以做决策,通过重放公共任务级记录,提出部分预算需满足支持完整基准决策、覆盖任务组及控制未解决比较比例等条件,还指出部分评估报告应包含的内容。

Comments KDD 2026 Workshop Agentic AI Evaluation and Trustworthiness

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09908 2026-07-14 cs.CL cs.IR 新提交 80%

RouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation

RouteRec:推荐代理选择与聚合的严格评估

Kaiji Zhou, Vladimir Kalmykov, Yue Feng

机构 * University of Birmingham(伯明翰大学)

专题命中 Agent评测 :agent(title,abstract);分类 cs.CL;AI agent(comments)

AI总结 研究在推荐系统异构代理选择中,RouteRec框架在成本约束下比较请求级硬选与项目级学习聚合,发现在MovieLens-1M数据集上,请求级选择粗糙,项目级聚合更有前景,不同聚合方式有不同效果。

Comments 8 pages, 7 figures. Accepted at AgentSearch 2026 (The First Workshop on Indexing, Retrieval, and Ranking of AI Agents), co-located with SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20634 2026-06-23 cs.AI cs.CY 新提交 80%

DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency

DEMM-Bench:面向智能体运行时治理证据充分性的跨机制基准

Oleg Solozobov

机构 * Independent Researcher (Global)(全球独立研究员)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 提出DEMM-Bench,基于决策证据成熟度模型(DEMM),通过8个证据机制和8个退化条件,衡量智能体运行时记录能否重构决策级属性,解决现有方法过度声称问题。

Comments 41 pages, 8 tables, no figures. Benchmark and dataset paper. Dataset (CC-BY-4.0) and code (Apache-2.0): Zenodo doi:10.5281/zenodo.20426092 and Hugging Face https://huggingface.co/datasets/dev404ai/DEMM-Bench; code at https://github.com/agent-runtime-evidence/decision-evidence-benchmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21404 2026-05-21 cs.LG 80%

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

十二篇LLM代理基准测试论文披露了什么:一项初步审计和开放评分方案

Mahdi Naser Moghadasi, Faezeh Ghaderi

机构 * Research Division, BrightMind AI(BrightMind AI研究部) Texas Tech University(德克萨斯理工大学) University of Texas at Arlington(德克萨斯大学阿灵顿分校)

专题命中 Agent评测 :agent(title,abstract);分类 cs.LG

AI总结 本文通过分析十二篇知名LLM代理基准测试论文,揭示了这些论文在评估方法披露方面的不足,设计了一种开放评分方案以提高透明度和可重复性。

Comments Pilot audit of 12 LLM agent benchmark papers; schema, codebook, and per-paper scoring sheet released. Submission to IEEE Big Data 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19362 2026-05-21 cs.HC cs.AI 80%

Toward User Comprehension Supports for LLM Agent Skill Specifications

向LLM代理技能规范提供用户理解支持

Zikai Alex Wen

机构 * University of Washington, Tacoma School of Engineering \& Technology Tacoma, Washington, USA University of Washington, Tacoma School of Engineering \& Technology

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 研究探讨了技能规范是否有助于用户形成对技能消耗、产生和覆盖范围的有限预期,并通过分析878个网络安全技能的文本线索,发现仅少数规范包含必要的提示,强调应将规范视为面向用户的能劾示范而非仅执行指令的容器。

Comments To appear at ACM CAIS Workshop Agent Skill 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11341 2026-05-13 cs.AI 80%

CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening

CPEMH:一种用于基础模型系统中基于提示的行为评估与保证的代理框架

Giuliano Lorenzoni, Ivens Portugal, Paulo Alencar, Donald Cowan

机构 * University of Waterloo(滑铁卢大学)

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI;agent(comments)

AI总结 本文提出CPEMH框架,用于评估基础模型系统在基于转录数据的心理健康筛查中的提示驱动行为,通过模块化代理设计实现行为保证,案例研究展示了其在对话和临床敏感领域稳定性和审计能力。

Comments 4 pages, 2 figures. Accepted at the AGENT 2026 Workshop (ICSE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15483 2026-03-17 cs.AI 80%

Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis

谈话、评估、诊断:基于用户意识的代理评估与自动错误分析

Penny Chong, Harshavardhan Abichandani, Jiyuan Shen, Atin Ghosh, Min Pyae Moe, Yifan Mai, Daniel Dahlmeier

机构 * SAP Stanford University(斯坦福大学)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 本文提出TED框架,通过用户角色模板、自动评分和错误分析,提升代理评估的全面性与效率,实验显示在模型和用户水平上取得8-10%的性能提升。

Comments Accepted as a conference paper at ICLR 2026. Code and dataset are available in the repository https://github.com/SAP-samples/agent-quality-inspect

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21460 2025-03-28 cs.CL 80%

Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao, Junwei Yang, Yiyang Gu, Bohan Wu, Binqi Chen, Ziyue Qiao, Qingqing Long, Rongcheng Tu, Xiao Luo, Wei Ju, Zhiping Xiao, Yifan Wang, Meng Xiao, Chenwu Liu, Jingyang Yuan, Shichang Zhang, Yiqiao Jin, Fan Zhang, Xian Wu, Hanqing Zhao, Dacheng Tao, Philip S. Yu, Ming Zhang

专题命中 Agent评测 :agent(title,abstract);分类 cs.CL

Comments 329 papers surveyed, resources are at https://github.com/luo-junyu/Awesome-Agent-Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14016 2024-05-31 cs.CL 80%

Towards Uncertainty-Aware Language Agent

Jiuzhou Han, Wray Buntine, Ehsan Shareghi

专题命中 Agent评测 :agent(title,abstract);分类 cs.CL

Comments Our code and data are at https://uala-agent.github.io. (accepted to ACL 2024 Findings). arXiv admin note: text overlap with arXiv:2310.05915

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07570 2022-06-16 cs.MA cs.LG cs.SI stat.ML 80%

Calibrating Agent-based Models to Microdata with Graph Neural Networks

Joel Dyer, Patrick Cannon, J. Doyne Farmer, Sebastian M. Schmon

专题命中 Agent评测 :agent(title,abstract);分类 cs.LG

Comments Accepted for a Spotlight presentation at the ICML 2022 Artificial Intelligence for Agent-based Modelling (AI4ABM) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03147 2022-03-08 cs.AI cs.MA 80%

Automatic Calibration Framework of Agent-Based Models for Dynamic and Heterogeneous Parameters

Dongjun Kim, Tae-Sub Yun, Il-Chul Moon, Jang Won Bae

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI;autonomous agent(comments)

Comments 3 pages, 6 figures, Autonomous Agents and Multiagent Systems (AAMAS 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.09586 2021-08-24 cs.AI 80%

Learning Causal Models of Autonomous Agents using Interventions

Pulkit Verma, Siddharth Srivastava

专题命中 Agent评测 :autonomous agent(title);agent(abstract);分类 cs.AI;planning(comments)

Comments IJCAI 2021 Workshop on Generalization in Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
0910.2874 2010-07-05 cs.AI cs.MA 80%

An Agent Based Classification Model

Feng Gu, Uwe Aickelin, Julie Greensmith

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

Comments 4 pages, 2 figures, 9th European Agent Systems Summer School, Durham, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00945 2026-08-12 cs.HC 版本更新 80%

Sighted by Default: Addressing Implicit Vision Assumptions in Real-Time VLM Assistance for BLV Users

默认以视觉为中心:解决面向盲人和低视力(BLV)用户的实时视觉语言模型(VLM)辅助中的隐含视觉假设问题

Yi Zhao, Siqi Wang, Qiqun Geng, Erxin Yu, Jing Li

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 针对现有面向BLV用户的实时VLM辅助工具存在的以视觉为中心的默认偏见问题,提出VIA-Agent模型,其在保持与Doubao相当成功率的同时,缩短了任务时间并减少了对话轮次,提升了用户信任度。

Comments Accepted to UIST 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05177 2026-08-07 cs.CY 新提交 80%

Using AI-Generated Feedback to Improve Critical Thinking and Writing Proficiency

利用AI生成的反馈提升批判性思维与写作能力

Qi Zhu, Xiaoming Zhai, Yan Zou, Chunlei Gao

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本研究开发了WISE Agent智能体,通过对260名六年级学生的三个月干预,证实其作为认知支架可促进不同水平学生批判性思维的差异化发展,为AI辅助培养批判性思维提供了可行路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04884 2026-08-07 cs.CV 版本更新 80%

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

浑元OCR-1.5:让轻量级OCR视觉语言模型更快更好

Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang, Hao Feng, Yongkun Du, Binghong Wu, Zheng Ruan, Zhiqiong Lu, Liang Wu, Pengyuan Lyu, Huawen Shen, Zibin Lin, Shijing Hu, Jieneng Yang, Hongbing Wen, Guanghua Yu, Hong Liu, Bochao Wang, Can Ma, Han Hu, Chengquan Zhang, Yu Zhou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Large Language Model Department, Tencent(腾讯大语言模型部) Nankai University(南开大学)

专题命中 Agent评测 :agentic(summary_cn,abstract);agent(abstract)

AI总结 研究改进轻量级端到端OCR视觉语言模型。基于HunyuanOCR-1.0架构,用DFlash提升效率,提出Agentic Data Flow增强能力,在多任务表现出色,还将发布模型权重和代码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24821 2026-07-29 cs.MM cs.CV cs.SD 新提交 80%

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

AVE-Compass:迈向音频-视频编辑能力的整体评估

Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 研究针对现有视频编辑基准未充分考虑视听耦合问题,提出AVE-Compass基准及AVE-Agent框架,通过多种评估方式评估编辑能力,能分解复杂指令迭代改进结果,提升了跨模态编辑等方面表现及感知质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24009 2026-07-20 cs.CR cs.AI cs.CL cs.LG 版本更新 80%

Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking

Jailbreak Foundry: 从论文到可运行攻击的可重复基准测试

Zhicheng Fang, Jingjie Zheng, Chenxu Fu, Wei Xu

机构 * Shanghai Qi Zhi Institute(上海启智研究院) Tsinghua University(清华大学) University of Melbourne(墨尔本大学)

专题命中 Agent评测 :agent(abstract);workflow(abstract);multi-agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 JAILBREAK FOUNDRY通过多代理工作流将论文转换为可执行模块,实现可重复基准测试的标准化评估和自动化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10044 2026-06-24 cs.AI cs.CL cs.CY cs.LG 80%

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

脚手架下的安全性:评估条件如何影响测量的安全性

David Gringras

机构 * Harvard University(哈佛大学) MIT(麻省理工学院)

专题命中 Agent评测 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本研究通过62,808次盲法预注册评估,测试了六种前沿模型在四种部署配置下的安全性,发现脚手架架构对安全性影响较小,而格式转换(如选择题与开放式问题)可导致5-20个百分点的测量差异,且模型-脚手架间存在显著异质性,质疑了单一综合安全性分数的实用性。

Comments 74 pages including appendices. 6 frontier models, 62,808 primary observations (~89k total). Pre-registered: OSF DOI 10.17605/OSF.IO/CJW92. Code and data: https://github.com/davidgringras/safety-under-scaffolding

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.00424 2026-06-04 econ.GN nlin.AO q-fin.EC 80%

Residential income segregation: A behavioral model of the housing market

住宅收入隔离:住房市场的行为模型

Marco Pangallo, Jean Pierre Nadal, Annick Vignes

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文提出一个基于Agent的模型,研究收入隔离、收入不平等和房价之间的关系,通过分析买家和卖家的行为及价格形成机制,揭示收入分布不均对房价和隔离程度的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.07401 2026-06-04 cs.MA cs.SY eess.SY math.DS 80%

Asynchronous opinion dynamics on the $k$-nearest-neighbors graph

异步意见动力学在k-最近邻图上

Wilbert Samuel Rossi, Paolo Frasca

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 研究异步意见更新机制下的k-最近邻图意见动力学,发现其与传统模型有显著差异,证明当agent数量小于2k时动力学趋于共识。

Comments 17 pages, 4 figures, (to be) presented at the 57th IEEE Conference on Decision and Control, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.05660 2026-06-04 physics.soc-ph econ.GN q-fin.EC 80%

The role of consumer networks in firms' multi-characteristics competition and market-share inequality

消费者网络在企业多特征竞争和市场份额不平等中的作用

Antonios Garas, Athanasios Lapatinas

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文通过多特征空间中的位置分析模型,研究消费者网络对企业竞争和市场份额不平等的影响,采用动态 agent-based 分析,揭示多维产品差异化竞争中的适应性学习机制。

Comments 33 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.06959 2026-06-04 math.OC econ.GN q-fin.EC 80%

Strategic Growth with Recursive Preferences: Decreasing Marginal Impatience

具有递归偏好战略增长:边际不耐受递减

Luis Alcala, Fernando Tohme, Carlos Dabus

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文研究了资本积累两 agent 模型中策略、异质性和增长的相互作用,采用递归效用函数表示递减边际不耐受偏好,分析两种信息结构下的稳态均衡,探讨均衡的存在性和稳定性。

Comments 55 pages, 14 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.0355 2026-06-04 math.DS cs.MA cs.SY eess.SY 80%

On symmetric continuum opinion dynamics

关于对称连续意见动力学

Julien M. Hendrickx, Alex Olshevsky

专题命中 Agent评测 :agent(summary_cn,abstract)

AI总结 本文研究了连续agent群体中常见意见动态模型的渐近行为,证明对称交互下意见分布收敛,并探讨更强意义上的收敛性。

Comments 28 pages, 2 figures, 3 files

详情

展开后加载摘要…

URL PDF HTML 收藏