arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15603 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15603 篇

2507.06134 2026-02-18 cs.AI 89%

OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety

OpenAgentSafety: 一个全面评估现实世界AI代理安全性的框架

Sanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang, Nouha Dziri, Graham Neubig, Maarten Sap

机构 * Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Allen Institute for Artificial Intelligence(人工智能研究院)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);agentic(abstract);分类 cs.AI

AI总结 OpenAgentSafety提出一个全面评估AI代理安全性的框架,通过真实工具和多任务测试揭示代理在现实世界中的安全漏洞,强调需要加强安全防护。

Comments 26 pages, 10 figures, Accepted at ICLR 2026 and IASEAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02475 2026-02-03 cs.AI 89%

AgentRx: Diagnosing AI Agent Failures from Execution Trajectories

AgentRx: 从执行轨迹诊断AI代理故障

Shraddha Barke, Arnav Goyal, Alind Khare, Avaljot Singh, Suman Nath, Chetan Bansal

机构 * Microsoft Research(微软研究院) Microsoft(微软)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);multi-agent(abstract);分类 cs.AI

AI总结 AgentRx通过自动化诊断框架,从执行轨迹中定位AI代理的关键故障步骤,并提升跨领域的故障归因能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00856 2026-01-29 econ.EM cs.AI 89%

Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks

AI能否掌握计量经济学?来自计量经济学AI代理在专家级任务上的证据

Qiang Chen, Tianyang Han, Jin Li, Ye Luo, Zigan Wang, Yuxiao Wu, Xiaowei Zhang, Tuo Zhou

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);planning(abstract);分类 cs.AI

AI总结 本文提出MetricsAI,一个基于MetaGPT框架的计量经济学AI代理,通过实证分析展示其在专家级任务中超越传统模型的能力,推动社会科学研究的智能化与教育应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01126 2026-01-27 cs.CL 89%

RoboPhD: Self-Improving Text-to-SQL Through Autonomous Agent Evolution

RoboPhD:通过自主代理进化实现自我改进的文本到SQL系统

Andrew Borthwick, Stephen Ash

机构 * Independent Researchers(独立研究者)

专题命中 Agent评测 :agent(title,abstract);autonomous agent(title);AI agent(abstract);agentic(abstract)

AI总结 RoboPhD通过自主进化代理提升文本到SQL性能,无需外部指导,在低成本模型上实现显著改进。

Comments 18 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16685 2026-01-26 cs.AI 89%

AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning

AgentsEval: 通过多智能体推理实现医学影像报告的临床可信评估

Suzhong Fu, Jingqi Dong, Xuan Ding, Rui Sun, Yiming Yang, Shuguang Cui, Zhen Li

机构 * FNii-Shenzhen, The Chinese University of Hong Kong (Shenzhen) School of Science and Engineering(深圳FNii、香港中文大学(深圳)科学与工程学院)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);workflow(abstract);分类 cs.AI

AI总结 AgentsEval通过多智能体推理框架,实现医学影像报告的临床可信评估,提供结构化反馈和稳健的评估结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07553 2025-12-17 cs.AI 89%

COMMA: A Communicative Multimodal Multi-Agent Benchmark

COMMA:一种基于通信的多模态多智能体基准

Timothy Ossowski, Danyal Maqbool, Jixuan Chen, Zefan Cai, Tyler Bradshaw, Junjie Hu

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校) Department of Computer Sciences UC San Diego(计算机科学系加州大学圣地亚哥分校) Department of Radiology University of Wisconsin-Madison(放射学系威斯康星大学麦迪逊分校) Department of Computer Sciences Department of Biostatistics and Medical Informatics University of Wisconsin-Madison(计算机科学系生物统计学与医学信息学系威斯康星大学麦迪逊分校)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);agentic(abstract);分类 cs.AI

AI总结 COMMA基准通过语言通信评估多模态多智能体系统的协作性能,揭示现有模型在智能体协作中的不足。

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05410 2025-11-10 cs.HC cs.SE 89%

Story Arena: A Multi-Agent Environment for Envisioning the Future of Software Engineering

Justin D. Weisz, Michael Muller, Kush R. Varshney

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);AI agent(abstract);分类 cs.SE

Comments 8 pages. Appeared in the 2025 Workshop on The End of Programming (as we know it): Envisioning Radical Re-Conceptualizations of Co-Coding with AI, held in conjunction with the Aarhus 2025 Decennial Conference, August 18-22, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04173 2025-11-10 cs.AI 89%

Open Agent Specification (Agent Spec): A Unified Representation for AI Agents

Soufiane Amini, Yassine Benajiba, Cesare Bernardis, Paul Cayet, Hassan Chafi, Abderrahim Fathan, Louis Faucon, Damien Hilloulin, Sungpack Hong, Ingo Kossyk, Tran Minh Son Le, Rhicheek Patra, Sujith Ravi, Jonas Schweizer, Jyotika Singh, Shailender Singh, Weiyi Sun, Kartik Talamadupula, Jerry Xu

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);agentic(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05347 2025-08-06 cs.CL cs.MA 89%

GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report Evaluation

Zhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng, Huichi Zhou, Zihao Jin, Jiahao Huang, Zhifan Gao, Dominic C Marshall, Yingying Fang, Guang Yang

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);workflow(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04253 2025-06-06 cs.AI cs.HC 89%

HADA: Human-AI Agent Decision Alignment Architecture

Tapio Pitkäranta, Leena Pitkäranta

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Aalto University(阿尔托大学) Department of Industrial Engineering and Management(工业工程与管理系)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);multi-agent(abstract);分类 cs.AI

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11761 2025-04-24 cs.AI 89%

Harnessing Language for Coordination: A Framework and Benchmark for LLM-Driven Multi-Agent Control

Timothée Anne, Noah Syrkis, Meriem Elhosni, Florian Turati, Franck Legendre, Alain Jaquier, Sebastian Risi

机构 * IT University of Copenhagen(哥本哈根技术大学) armasuisse Science+Technology(armasuisse科学与技术)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05584 2023-03-13 cs.RO cs.AI cs.MA 89%

SOCIALGYM 2.0: Simulator for Multi-Agent Social Robot Navigation in Shared Human Spaces

Zayne Sprague, Rohan Chandra, Jarrett Holtz, Joydeep Biswas

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);autonomous agent(abstract);分类 cs.AI

Comments Submitted to RSS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15040 2026-07-31 cs.AI cs.CL 版本更新 89%

Orchard: An Open-Source Agentic Modeling Framework

Orchard:一个开源的智能体建模框架

Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng, Xiao Yu, Rui Yang, Tao Ge, Alessandro Sordoni, Xingdi Yuan, Yelong Shen, Pengcheng He, Tong Zhang, Zhou Yu, Jianfeng Gao

机构 * Microsoft Research(微软研究院) Columbia University(哥伦比亚大学) UIUC(伊利诺伊大学香槟分校)

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);autonomous agent(abstract);tool use(abstract)

AI总结 本文提出Orchard,一个开源的智能体建模框架,通过轻量级环境服务和三种智能体建模食谱,实现了跨领域可重用的智能体数据、训练和评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26645 2026-05-29 cs.AI cs.LG 89%

SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

SciHorizon-DataEVA:面向异构科学数据AI就绪性评估的智能体系统

Dianyu Liu, Chuan Qin, Xi Chen, Xiaohan Li, Wenxi Xu, Yuyang Wang, Xin Chen, Yuanchun Zhou, Hengshu Zhu

机构 * SciHorizon Team, Computer Network Information Center, Chinese Academy of Sciences(科学前沿团队,计算机网络信息中心,中国科学院)

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);planning(abstract);workflow(abstract)

AI总结 提出SciHorizon-DataEVA智能体系统,基于Sci-TQA2原则和层次化多智能体评估方法,实现对异构科学数据的可扩展AI就绪性评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03265 2025-11-10 cs.LG cs.AI 89%

Cognitive Edge Computing: A Comprehensive Survey on Optimizing Large Models and AI Agents for Pervasive Deployment

Xubin Wang, Qing Li, Weijia Jia

专题命中 Agent评测 :AI agent(title,abstract);agent(abstract);tool use(abstract);agentic(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27051 2025-11-03 cs.AI cs.LG 89%

Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement

Aaditya Shukla, Sidney Knowles, Meenakshi Madugula, Dave Farris, Ryan Angilly, Santiago Pombo, Anbang Xu, Lu An, Abhinav Balasubramanian, Tan Yu, Jiaxiang Ren, Rama Akkiraju

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.LG

Comments 20 pages, 5 figures, 5 tables. Presents MAPE-K control loop application to enterprise AI agent improvement with experimental validation on NVIDIA's NVInfo AI system

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09666 2026-08-11 cs.AI 新提交 89%

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

开放评估智能体:视觉生成模型的高效且可提示的评估

Shulin Tian, Ziqi Huang, Fan Zhang, Hongyuan Zhu, Yu Qiao, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S-Lab) Agency for Science, Technology and Research (A*STAR)(科学、技术与研究局(A*STAR)) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 Agent评测 :agent(title,summary_cn);planning(abstract);分类 cs.AI

AI总结 该研究提出Evaluation Agent框架及基于其构建的Open-EA,可高效评估视觉生成模型,将评估时间减至传统方法的10%,还通过EA-CoT-10K和EA-3B减少对专有骨干的依赖,验证了其在多基准及跨家族的有效性。

Comments Journal extension of our ACL 2025 paper (arXiv:2412.09645). 12 pages. Code: https://github.com/Vchitect/Evaluation-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04292 2026-08-06 cs.CV 新提交 89%

Binding Biometrics with AI Agent Identifiers for Delegation of Authority

将生物特征与AI智能体标识符绑定以实现权限委托

Joseph Geo Benjamin, Anil K Jain, Karthik Nandakumar

机构 * Michigan State University(密歇根州立大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);agentic(abstract)

AI总结 本研究提出BIND框架,通过生物特征密码系统将人类生物特征与AI智能体ID及权限范围绑定,实现经认证的权限委托,基于人脸特征的实验验证其在零错误匹配率下96%真实匹配率,支持1024位智能体令牌。

Comments Accepted in IJCB sessions 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00548 2026-08-04 cs.CV 新提交 89%

DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable

DrawAI:用于生成可编辑光栅图像的智能体基准与工作流

Pu Cao, Qingye Kong, Xuedan Yin, Xuekun Zhao, Rupeng Yan, Qing Song, Yao Zhang, Lu Yang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学)

专题命中 Agent评测 :workflow(title,abstract);agentic(title,abstract);agent(abstract)

AI总结 DrawAI提出图像到可编辑重建任务,构建含DrawAI-Bench基准与DrawAI-Flow工作流的智能体方案,经实验验证DrawAI-Flow可提升可编辑结构,不同模型-harness配置的重建质量与成本差异显著。

Comments Project URL: https://drawai.renaissancemind.ai/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31227 2026-07-01 cs.CR 新提交 89%

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

保护AI代理:多层代理红队测试的统一框架

Yong Yang, Xing Zheng, Huiyu Wu, Huangsheng Cheng, Xiaorong Shi, Jing Guo, Bo Yang, Yi Zhou, Xiangfan Wu, Zonghao Ying

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);agentic(abstract)

AI总结 提出AI-Infra-Guard框架,通过分层攻击面(基础设施、协议/工具、代理行为、模型)匹配不同检测范式,覆盖75+组件和1400+漏洞规则,实现代理安全红队测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18935 2026-05-20 econ.EM 89%

The Agentic Economy: Humans, AI Agents, Robots, and the Measurable Transition toward Distributed Economic Action

代理经济:人类、AI代理、机器人以及向分布式经济行动的可测量转型

Davit Gondauri, Mikheil Batiashvili

专题命中 Agent评测 :AI agent(title,abstract);agentic(title,abstract);agent(abstract)

AI总结 本文提出代理经济的概念,并探讨其可测量的先决条件:经济行动在人类、AI代理、工业机器人、可执行协议、计算基础设施和能源系统之间日益分散。研究指出,传统类别如劳动、资本、企业、市场、生产率和信任仍然必要但不完整。方法上采用概念-实证定量诊断设计,利用公开的AI投资、AI采纳、机器人安装和运营库存、数据中心电力需求以及劳动力市场再分配数据。结果表明,AI采纳加速,AI投资信号广泛资本分配,工业机器人代表持续的网络物理行动能力,计算扩展增加数据中心电力压力,劳动力预测更一致于任务再分配而非劳动力消失。本文贡献了一个连接模型/软件代理能力、机器人能力、计算-能源耦合、协议化、可审计信任和人类主权的行动能力框架。结论是代理经济尚未成为全球秩序,但其转型压力足以要求独特的经济词汇、可重复的诊断和未来行业层面的测量。

Comments 52 pages, 5 figures, 16 tables, 8 algorithms. Original conceptual-empirical diagnostic article

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10456 2026-04-14 cs.CV 89%

A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation

一个用于指令驱动电影视频编译的基准和多智能体系统

Peixuan Zhang, Chang Zhou, Ziyuan Zhang, Hualuo Liu, Chunjie Zhang, Jingqi Liu, Xiaohui Zhou, Xi Chen, Shuchen Weng, Si Li, Boxin Shi

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频事业部AI技术中心) Tsinghua University(清华大学) Beijing Academy of Artificial Intelligence(北京智源人工智能研究院) State Key Lab of Multimedia Info. Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室) Nat’l Eng. Research Ctr. of Visual Technology, School of Computer Science, Peking University(北京大学计算机学院国家视觉技术工程研究中心) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);planning(abstract)

AI总结 本文提出CineBench基准和CineAgents系统,解决电影视频编译中的上下文崩溃和时间碎片化问题,通过脚本逆向工程和迭代叙事规划生成更连贯的编译结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09052 2026-03-11 cs.AI cs.CL cs.LG 89%

From Days to Minutes: An Autonomous AI Agent Achieves Reliable Clinical Triage in Remote Patient Monitoring

从天到分钟:一个自主AI代理在远程患者监测中实现可靠的临床分诊

Seunghwan Kim, Tiffany H. Kung, Heena Verma, Dilan Edirisinghe, Kaveh Sedehi, Johanna Alvarez, Diane Shilling, Audra Lisa Doyle, Ajit Chary, William Borden, Ming Jack Po

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 Sentinel通过自主AI代理实现远程患者监测的高效分诊,超越个体医生的灵敏度,提供可扩展且临床可行的解决方案。

Comments 46 pages, 11 figures, Abstract in metadata is shortened to meet arXiv character limits; see PDF for full version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10718 2026-01-19 cs.AI cs.CL cs.IR cs.LG 89%

Japanese AI Agent System on Human Papillomavirus Vaccination: System Design

日本的人乳头瘤病毒疫苗人工智能代理系统:系统设计

Junyu Liu, Siwen Yang, Dexiu Ma, Qian Niu, Zequn Zhang, Momoko Nagai-Tanima, Tomoki Aoyama

机构 * Kyoto University(京都大学) University of Waterloo(滑铁卢大学) Texas Tech University(德克萨斯技术大学) The University of Tokyo(东京大学) USTC(中国科学技术大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本研究设计了一种双用途的人工智能代理系统,通过对话界面提供HPV疫苗信息并生成分析报告,有效应对疫苗犹豫问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15216 2025-12-03 cs.CR cs.AI cs.CL cs.LG 89%

BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems

BountyBench: AI代理攻击者和防御者对现实世界网络安全系统的影响

Andy K. Zhang, Joey Ji, Celeste Menders, Riya Dulepet, Thomas Qin, Ron Y. Wang, Junrong Wu, Kyleen Liao, Jiliang Li, Jinghan Hu, Sara Hong, Nardos Demilew, Shivatmica Murgai, Jason Tran, Nishka Kacheria, Ethan Ho, Denis Liu, Lauren McLane, Olivia Bruvik, Dai-Rong Han, Seungwoo Kim, Akhil Vyas, Cuiyuanxiu Chen, Ryan Li, Weiran Xu, Jonathan Z. Ye, Prerit Choudhary, Siddharth M. Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel E. Ho, Percy Liang

机构 * Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 BountyBench通过评估AI代理在漏洞检测、利用和修补中的表现,揭示AI在网络安全中的影响。

Comments 113 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04555 2025-04-08 cs.DC 89%

SchEdge: A Dynamic, Multi-agent, and Scalable Scheduling Simulator for IoT Edge

Ali Hamedi, Amirali Ghaedi, Amin Soltanbeigi, Athena Abdi

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);workflow(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00881 2025-01-03 cs.MA 89%

Agentic Systems: A Guide to Transforming Industries with Vertical AI Agents

Fouad Bousetouane

专题命中 Agent评测 :AI agent(title,abstract);agentic(title,abstract);agent(abstract)

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00995 2024-07-02 cs.CY cs.SY eess.SY physics.app-ph 89%

Data on the Move: Traffic-Oriented Data Trading Platform Powered by AI Agent with Common Sense

Yi Yu, Shengyue Yao, Tianchen Zhou, Yexuan Fu, Jingru Yu, Ding Wang, Xuhong Wang, Cen Chen, Yilun Lin

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);multi-agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03380 2025-09-04 cs.AI cs.CL 89%

Situating AI Agents in their World: Aspective Agentic AI for Dynamic Partially Observable Information Systems

Peter J. Bentley, Soo Ling Lim, Fuyuki Ishikawa

专题命中 Agent评测 :AI agent(title,abstract);agentic(title,abstract);分类 cs.AI、cs.CL;agent(journal_ref)

Comments 9 pages

Journal ref 7th International Workshop on Agent-Based Modelling of Human Behaviour (ABMHuB'25), ALife 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24112 2026-08-12 cs.AI 版本更新 89%

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

ReMMD: 面向多模态虚假信息检测的现实多语言多图像智能体验证

Chenhao Dang, Dantong Zhu, Jun Yang, Conghui He, Weijia Li

机构 * Shanghai Jiaotong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Central South University(中南大学) China Electronics Technology Group Corporation 15th Research Institute(中国电子科技集团公司第十五研究所)

专题命中 Agent评测 :agent(summary_cn,abstract);agentic(title,abstract);分类 cs.AI

AI总结 提出ReMMD框架,包含多语言多图像基准ReMMDBench和持久记忆验证器ReMMD-Agent,通过原子点分解和可重用证据集实现高效准确的多模态虚假信息检测。

Comments The project is available at https://dang-ai.github.io/ReMMD

详情

展开后加载摘要…

URL PDF HTML 收藏