arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15662 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15662 篇

2606.15225 2026-06-16 cs.LG cs.AI cs.IR 新提交 82%

Edu-Theater: A Data-Efficient Agent Framework for Scalable Learner Behavior Simulation through Staging Roll-Call

Edu-Theater: 一种通过点名排演实现可扩展学习者行为模拟的数据高效智能体框架

Weibo Gao, Qi Liu, Linan Yue, Zheng Zhang, Yichao Du, Fangzhou Yao, Ao Yu, Zhenya Huang, Shijin Wang

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) Southeast University(东南大学) Alibaba Group(阿里巴巴集团) iFLYTEK Co., Ltd.(科大讯飞股份有限公司)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG

AI总结 提出Edu-Theater框架,通过构建群体水平能力先验和少量诊断查询,利用LLM智能体模拟学习者行为,在减少数据需求的同时提高模拟精度,并增强下游自适应测试等应用。

Comments LLM Agent, Educational Data Mining, Data Synthesis, Human Simulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31278 2026-06-05 cs.AI cs.LG stat.ME 82%

Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation

工业化预测驱动推断:用于可靠生成式AI与智能体系统评估的GLIDE库

Grégoire Martinon, Ibrahim Merad, Mohammed Raki

机构 * University of California, Berkeley(加州大学伯克利分校) Google Research(谷歌研究院)

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI、cs.LG

AI总结 提出GLIDE开源库,统一多种预测驱动推断方法,提供无偏估计与有效置信区间,显著降低人工标注成本。

Comments 8 pages, Accepted to the ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems, Seoul, South Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02494 2026-06-02 cs.SE cs.AI 82%

Monitoring Agentic Systems Before They're Reliable

在代理系统可靠之前对其进行监控

Marisa Ferrara Boston, Glen Hanson, Effi Georgala, JD Hudgens, Heather Frase

机构 * Reins AI USA(Reins AI美国公司) Veraitech USA(Veraitech美国公司)

专题命中 Agent评测 :agentic(title,abstract);分类 cs.AI、cs.SE

AI总结 针对生产环境中代理系统因结构缺陷主导故障的问题,提出一种基于方差信号的三维度三范围监控与分类方法,并通过合成测试验证其有效性。

Comments 9 pages, 2 figures, 3 tables. Accepted to the Workshop on Agentic Software Engineering (AgenticSE), co-located with ACM CAIS 2026 (non-archival)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24358 2026-03-24 cs.SE cs.CL 82%

Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation

通过代理驱动的标注和评估自动评估代码代理

Lingyue Fu, Bolun Zhang, Hao Guan, Yaoming Zhu, Lin Qiu, Weiwen Liu, Xuezhi Cao, Xunliang Cai, Weinan Zhang, Yong Yu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 Agent评测 :agent(title,abstract);分类 cs.CL、cs.SE;autonomous agent(journal_ref)

AI总结 本文提出代理驱动的基准构建流程,引入PRDBench包含50个真实Python项目,通过专门模型提升评估准确性,验证了框架的可扩展性和鲁棒性。

Comments Accepted by AAMAS 2026

Journal ref Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), Paphos, Cyprus, May 25-29, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15969 2026-02-03 cs.LG cs.AI 82%

LinearizeLLM: An Agent-Based Framework for LLM-Driven Exact Linear Reformulation of Nonlinear Optimization Problems

LinearizeLLM: 一个基于代理的框架,用于由LLM驱动的非线性优化问题的精确线性重 formulations

Paul-Niklas Ken Kandora, Simon Caspar Zeller, Aaron Jeremias Elsing, Elena Kuss, Steffen Rebennack

机构 * Institute for Operations Research, Karlsruhe Institute of Technology, Karlsruhe, Germany(运营研究学院,卡尔斯鲁厄技术大学) Institute for Information Systems, Reutlingen University, Reutlingen, Germany(信息系统研究所,鲁特lingen大学)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.LG;workflow(comments)

AI总结 LinearizeLLM通过基于代理的LLM框架实现非线性优化问题的自动精确线性化,显著提升线性化效率和准确性。

Comments This version is a major revision with a new abstract, updated workflow logic and description, an expanded instance set, additional numerical experiments, and corrected bibliography entries

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17580 2025-11-25 cs.MA cs.AI cs.DC cs.SE 82%

A novel strategy for multi-resource load balancing in agent-based systems

基于智能体系统的多资源负载平衡新策略

Leszek Sliwko, Aleksander Zgrzywa

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.SE

AI总结 本文提出了一种基于智能体系统的多资源负载平衡策略,通过智能体的社会行为和适应能力优化复杂企业架构的结构。

Journal ref "A novel strategy for multi-resource load balancing in agent-based systems." International journal of intelligent information and database systems (Print) 3, no. 2 (2009): 180-202

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09645 2025-08-22 cs.CV cs.AI cs.CL 82%

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

Fan Zhang, Shulin Tian, Ziqi Huang, Yu Qiao, Ziwei Liu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.CL

Comments Equal contributions from first three authors. Project page: https://vchitect.github.io/Evaluation-Agent-project Code: https://github.com/Vchitect/Evaluation-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09068 2026-08-11 cs.SE 新提交 81%

Pseudo2CodeQA: A Benchmark for LLM-Based Structured Algorithmic Reasoning in Code Generation

Pseudo2CodeQA:面向基于大语言模型的代码生成中结构化算法推理的基准测试

Shadikur Rahman, Umme Ayman Koana, Syed Muhammad Danish

专题命中 Agent评测 :agentic(summary_cn,abstract);分类 cs.SE

AI总结 本文提出Pseudo2Code基准测试及Pseudo2Code Agentic Framework,实验表明该框架性能优于现有基线,验证了结构化伪代码对代码生成的提升作用。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27084 2026-07-30 cs.CV cs.AI 新提交 81%

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

SciFigQual-Bench:面向全文本上下文的科学图片质量评估基准

Zihan Deng, Chuanzhi Xu, Huiqi Liang, Haoyang Li, Xiaozhen Zhong, Lequan Yu

专题命中 Agent评测 :agent(summary_cn,abstract);分类 cs.AI

AI总结 本文针对现有科学图片质量评估方法的不足,构建了SciFigQual-Bench基准,设计SFQ-Agent框架实现自动化评估,实验显示其在eval1200子集上表现优于主流方案。

Comments † Equal contribution. Affiliations: 1: The University of Hong Kong 2: The University of Sydney 3: University of Electronic Science and Technology of China Corresponding authors: Zihan Deng (zhdeng@hku.hk), Chuanzhi Xu (chuanzhi.xu@sydney.edu.au) Project page: https://frankdengai.github.io/SciFigQual-Bench Source code & dataset: https://github.com/FrankDengAI/SciFigQual-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20498 2026-07-24 cs.AI 新提交 81%

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

AISE-Bench:用于学术知识图谱信息检索的全周期精选基准测试

Fanjin Zhang, Zhengyang Wang, Ruixuan Huang, Kefan Zhang, Amy Xin, Yuanchun Wang, Shu Zhao, Evgeny Kharlamov, Jie Tang, Juanzi Li

机构 * Renmin University of China(中国人民大学) Anhui University(安徽大学) Z-Lab Z.ai Tsinghua University(清华大学) Bosch Center for AI(博世人工智能中心) University of Oslo(奥斯陆大学)

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);planning(abstract);workflow(abstract)

AI总结 研究针对学术知识图谱信息检索基准测试不足的问题,引入AISE-Bench,通过定制工作流程和综合评估协议构建基准测试,为多步API使用的大语言模型智能体提供新测试平台,助力评估与改进。

Comments 9 pages, accepted by KDD 2026

Journal ref Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '26), August 09-13, 2026, Jeju Island, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03015 2026-07-07 cs.AI 新提交 81%

Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents

超越预测:预测市场代理中的信念到交易层

Yishu Wang, Yuxuan Wang, Jiaqi Deng, Hanyang Tang

机构 * Hong Kong University of Science and Technology(香港科学与技术大学) The University of Hong Kong(香港大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 Agent评测 :agent(summary_cn,abstract);分类 cs.AI

AI总结 以预测未来事件为通用人工智能测试平台,提出Raven-Agent这一预测市场自主交易代理,在存档决策集的控制重放中表现出色。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14207 2026-06-30 cs.AI 81%

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

人类恶意的回声在智能体中:为多轮在线骚扰攻击评估LLMs

Trilok Padhi, Pinxian Lu, Abdulkadir Erol, Tanmay Sutar, Gauri Sharma, Mina Sonmez, Munmun De Choudhury, Ugur Kursuncu

专题命中 Agent评测 :agent(abstract);planning(abstract);agentic(abstract);multi-agent(abstract)

AI总结 本文提出一个多轮在线骚扰攻击基准测试,通过合成数据集、多智能体模拟和三种 jailbreak 方法评估LLMs在多轮交互中的安全性,发现jailbreak调优显著提高攻击成功率,揭示了智能体在多轮对话中模仿人类攻击模式的弱点。

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00987 2026-06-02 cs.CV cs.AI 81%

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

多时相指代分割的开源基准与基线

Bingyu Li, Da Zhang, Tao Huo, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究所) China Telecom(中国电信) School of Artificial Intelligence, Optics and Electronics (iOPEN)(人工智能、光学与电子学院) Northwestern Polytechnical University(西北工业大学)

专题命中 Agent评测 :agent(summary_cn,abstract);分类 cs.AI

AI总结 提出多时相指代分割任务,通过自动化数据构建管道CRAFT-Agent生成首个基准MTRefSeg-21K,并设计两阶段训练的变化感知LVLM框架MTRefSeg-R1,实现优于现有基线的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19196 2026-05-20 cs.CL 81%

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

时间到REFLECT:我们能否信任LLM裁判来评估基于证据的研究代理?

Leyao Wang, Yanan He, Peng Chen, Asaf Yehudai, Yixin Liu, Rex Ying, Michal Shmueli-Scheuer, Arman Cohan

机构 * Yale University(耶鲁大学) IBM Research(IBM研究院)

专题命中 Agent评测 :agent(abstract);tool use(abstract);tool-use(abstract);agentic(abstract)

AI总结 本文提出REFLECT基准,用于评估LLM裁判在代理环境中的细粒度失败检测,揭示当前LLM裁判在推理、工具使用和报告质量上的可靠性不足,为构建更可靠的评估流程提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21510 2026-04-24 cs.CL 81%

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving

OptiVerse:一个面向优化问题求解的综合基准

Xinyu Zhang, Boxuan Zhang, Yuchen Wan, Lingling Zhang, YiXing Yao, Bifan Wei, Yaqiang Wu, Jun Liu

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, China(教育部智能网络与网络安全重点实验室) Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, China(陕西省大数据知识工程重点实验室) Lenovo Research(联想研究院)

专题命中 Agent评测 :agent(summary_cn,abstract);分类 cs.CL

AI总结 OptiVerse通过1000个跨领域问题评估LLM在复杂优化任务中的表现,揭示模型在难题上的性能下降,并提出Dual-View Auditor Agent提升准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05674 2026-04-10 cs.CR cs.AI 81%

Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark

迈向高效的攻击性安全LLM代理:超参数调优、LLM作为裁判以及一个轻量级CTF基准

Minghao Shao, Nanda Rani, Kimberly Milner, Haoran Xi, Meet Udeshi, Saksham Aggarwal, Venkata Sai Charan Putrevu, Sandeep Kumar Shukla, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri, Muhammad Shafique

机构 * New York University(纽约大学) New York University Abu Dhabi(纽约大学阿布扎比分校) Indian Institute of Technology Kanpur(印度理工学院坎普尔分校) International Institute of Information Technology Hyderabad(国际信息技术学院海得拉巴分校)

专题命中 Agent评测 :agent(abstract);planning(abstract);agentic(abstract);multi-agent(abstract)

AI总结 本文系统研究了驱动代理成功的关键因素,提出CTFJudge框架和CTF Competency Index指标,分析超参数对性能的影响,并开放CTFTiny基准供研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06696 2026-04-09 cs.AI 81%

AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents

AgentGate: 一种轻量级的结构化路由引擎用于物联网代理

Yujun Cheng, Enfang Cui, Hao Qin, Zhiyuan Liang, Qi Xu

机构 * School of Artificial Intelligence, University of Science and Technology Beijing(北京科技大学人工智能学院) China Telecom Research Institute(中国电信研究院) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);multi-agent(abstract)

AI总结 本文提出AgentGate,一种轻量级结构化路由引擎,用于在资源受限条件下高效且隐私友好的代理系统,通过分阶段决策和结构化接地提升路由效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03131 2026-04-06 cs.CR cs.AI 81%

A Systematic Security Evaluation of OpenClaw and Its Variants

对OpenClaw及其变种的系统性安全评估

Yuhang Wang, Haichang Gao, Zhenxing Niu, Zhaoxiang Liu, Wenjing Zhang, Xiang Wang, Shiguo Lian

机构 * Xidian University(西安电子科技大学) Data Science & Artificial Intelligence Research Institute, China Unicom(中国联通数据科学与人工智能研究院)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);tool use(abstract);planning(abstract)

AI总结 本文系统评估了六个OpenClaw系列代理框架的安全性,发现代理系统存在显著安全漏洞,且其风险高于孤立使用的基础模型。

Comments 39 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13201 2026-03-17 q-bio.GN cs.AI cs.MA 81%

Benchmarking LLM-based agents for single-cell omics analysis

对基于大语言模型的代理在单细胞组学分析中的性能评估

Yang Liu, Lu Zhou, Xiawei Du, Ruikun He, Xuguang Zhang, Rongbo Shen, Yixue Li

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);multi-agent(abstract)

AI总结 本文提出一个评估系统,用于严格评估代理在单细胞组学分析中的能力,发现Grok3-beta在测试框架中表现最佳,多代理框架通过角色分工提升协作与执行效率。

Comments please see clear figures in this version. 6 main figures; 13 supplementary figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06583 2026-03-10 cs.HC cs.CY cs.LG 81%

XInsight: Integrative Stage-Consistent Psychological Counseling Support Agents for Digital Well-Being

XInsight:集成化阶段一致的心理咨询支持代理用于数字福祉

Fei Wang, Jiangnan Yang, Junjie Chen, Yuxin Liu, Kun Li, Yanyan Wei, Dan Guo, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) Anhui University(安徽大学) United Arab Emirates University(阿联酋大学) Intelligent Interconnected Systems Laboratory of Anhui Province (HFUT)(安徽省智能互联系统实验室(HFUT))

专题命中 Agent评测 :agent(abstract);planning(abstract);workflow(abstract);multi-agent(abstract)

AI总结 XInsight是一种集成化多代理框架,通过阶段一致的工作流程和统一循环,提升网络应用中数字福祉的心理支持效果。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12876 2026-02-25 cs.AI 81%

BrowseComp-$V^3$: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents

BrowseComp-$V^3$:一种视觉、垂直和可验证的多模态浏览代理基准测试

Huanyao Zhang, Jiepeng Zhou, Bo Li, Bowen Zhou, Yanzhe Shan, Haishan Lu, Zhiyong Cao, Jiaoyang Chen, Yuqian Han, Zinan Sheng, Zhengwei Tao, Hao Liang, Jialong Wu, Yang Shi, Yuanpeng He, Jiaye Lin, Qintong Zhang, Guochen Yan, Runhao Zhao, Zhengpin Li, Xiaohan Yu, Lang Mei, Chong Chen, Wentao Zhang, Bin Cui

机构 * PKU(北京大学) HKUST(GZ)(香港科技大学) OUC(华侨大学) CASIA(中国科学院自动化研究所) HITSZ(哈尔滨工业大学) THU(清华大学) Huawei Cloud BU(华为云业务部)

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);tool-use(abstract);planning(abstract)

AI总结 BrowseComp-$V^3$提出了一种多模态浏览代理基准测试,通过300个跨模态问题评估深度搜索能力,揭示多模态信息整合的瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05101 2026-01-09 cs.AI 81%

Arabic Prompts with English Tools: A Benchmark

阿拉伯提示与英语工具:一个基准

Konstantin Kubrak, Ahmed El-Moselhy, Ammar Alsulami, Remaz Altuwaim, Hassan Ismail Fawaz, Faisal Alsaby

专题命中 Agent评测 :AI agent(abstract);autonomous agent(abstract);tool-use(abstract);agentic(abstract)

AI总结 本文提出首个阿拉伯语言LLMs工具调用能力评估基准,揭示阿拉伯语交互环境下工具调用准确率下降5-10%的显著差距。

Comments 10 pages, 10 figures, LLMs, Big Data, and Multilinguality for All (LLMs4All) Workshop at IEEE BigData 2025 Conference, Macau, December 10, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05408 2025-12-01 cs.CR cs.AI cs.CY 81%

Frontier AI's Impact on the Cybersecurity Landscape

前沿人工智能对网络安全领域的冲击

Yujin Potter, Wenbo Guo, Zhun Wang, Tianneng Shi, Hongwei Li, Andy Zhang, Patrick Gage Kelley, Kurt Thomas, Dawn Song

机构 * UC Berkeley(加州大学伯克利分校) UC Santa Barbara(加州大学圣芭芭拉分校) Google(谷歌)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);workflow(abstract)

AI总结 本文研究了前沿人工智能在网络安全中的影响,指出人工智能在攻击中的能力已超过防御,呼吁构建新基准、开发防御代理等以缓解风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12349 2025-10-17 cs.AI 81%

SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?

Jianzhu Yao, Kevin Wang, Ryan Hsieh, Haisu Zhou, Tianqing Zou, Zerui Cheng, Zhangyang Wang, Pramod Viswanath

机构 * Princeton University(普林斯顿大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Sentient Foundation(Sentient基金会)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);multi-agent(abstract)

Comments 48 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24651 2025-09-30 cs.AI 81%

"Stop replacing salt with sugar!'': Towards Intuitive Human-Agent Teaching

Nikolaos Kondylidis, Andrea Rafanelli, Ilaria Tiddi, Annette ten Teije, Frank van Harmelen

机构 * Vrije Universiteit Amsterdam(荷兰阿姆斯特丹自由大学) University of Pisa(比萨大学)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI;multi-agent(comments)

Comments 22nd European Conference on Multi-Agent Systems (EUMAS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20998 2025-09-26 cs.AI 81%

CORE: Full-Path Evaluation of LLM Agents Beyond Final State

Panagiotis Michelakis, Yiannis Hadjiyiannis, Dimitrios Stamoulis

机构 * Synkrasis Labs(Synkrasis实验室) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);tool-use(abstract);agentic(abstract)

Comments Accepted: LAW 2025 Workshop NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02494 2025-09-03 cs.AI 81%

GridMind: LLMs-Powered Agents for Power System Analysis and Operations

Hongwei Jin, Kibaek Kim, Jonghwan Kwon

机构 * Argonne National Laboratory(阿贡国家实验室)

专题命中 Agent评测 :agent(abstract);workflow(abstract);agentic(abstract);multi-agent(abstract)

Comments 11 pages, 9 figures, 2 tables. Work under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10385 2025-07-08 cs.CY cond-mat.mtrl-sci cs.AI physics.ins-det 81%

Autonomous Microscopy Experiments through Large Language Model Agents

Indrajeet Mandal, Jitendra Soni, Mohd Zaki, Morten M. Smedskjaer, Katrin Wondraczek, Lothar Wondraczek, Nitya Nand Gosvami, N. M. Anoop Krishnan

专题命中 Agent评测 :agent(abstract);AI agent(abstract);workflow(abstract);agentic(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.10672 2020-01-27 cs.AI 81%

Signaling Friends and Head-Faking Enemies Simultaneously: Balancing Goal Obfuscation and Goal Legibility

Anagha Kulkarni, Siddharth Srivastava, Subbarao Kambhampati

专题命中 Agent评测 :agent(abstract);AI agent(abstract);autonomous agent(abstract);planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30590 2026-06-01 cs.LG cs.AI cs.CL 81%

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

反事实评估揭示临床LLM和智能体的隐藏能力画像

Matt Turk

机构 * Protege Data Lab(Protege数据实验室)

专题命中 Agent评测 :AI agent(abstract,comments);tool use(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 提出因果敏感性评分(CSS),通过沿五个临床维度变异肿瘤病例来评估模型是否按预期方向更新推荐,发现与覆盖度指标排名相反,并揭示所有前沿模型在手术状态干预上的安全盲点。

Comments Accepted to RLEval @ ACM CAIS 2026 (Workshop on Methods and RL Environments for Evaluating AI Agents) and selected for an invited talk based on reviewer ratings. 4-page short paper + appendix

详情

展开后加载摘要…

URL PDF HTML 收藏