arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-04-21 至 2026-04-21 共收录 324 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 记忆与上下文管理 20 篇

2508.03793 2026-04-21 cs.CL cs.CR 57%

AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption

AttnTrace: 基于提示注入和知识腐败的上下文归因

Yanting Wang, Runpeng Geng, Ying Chen, Jinyuan Jia

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 记忆与上下文管理 :autonomous agent(abstract);分类 cs.CL

AI总结 本文提出AttnTrace,一种基于LLM对提示生成的注意力权重的上下文追溯方法,通过增强效果和技术改进,提高了准确性和效率,并展示了其在检测长上下文中的提示注入应用。

Comments To appear in IEEE S&P 2026. The code is available at https://github.com/Wang-Yanting/AttnTrace. The demo is available at https://huggingface.co/spaces/SecureLLMSys/AttnTrace

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18223 2026-04-21 cs.CV 50%

Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation

指令作为状态:环境引导和状态条件的语义理解用于具身导航

Zhen Liu, Yuhan Liu, Jinjun Wang, Jianyi Liu, Wei Song, Jingwen Fu

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) The Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) School of Information Science and Technology(信息科学与技术学院) Zhongguancun Academy(中关村学院)

专题命中 记忆与上下文管理 :agent(abstract)

AI总结 本文提出了一种环境引导和状态条件的语义理解方法,通过动态指令-感知纠缠提升具身导航性能,实验表明在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17119 2026-04-21 astro-ph.HE astro-ph.SR physics.plasm-ph 50%

Strong MHD Turbulence and Coherent Structures as Drivers of Cosmic Particle Acceleration

强磁流体湍流与相干结构作为宇宙粒子加速的驱动因素

Loukas Vlahos

专题命中 记忆与上下文管理 :agent(abstract)

AI总结 本文探讨强湍流与相干结构在宇宙粒子加速中的核心作用,强调其在能量转化和粒子加速中的关键地位,提出需整合多尺度等离子体动力学与可观测的高能粒子特征。

Comments 24 pages, 24 figures. Comments and suggestions are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏

2. Agent评测 47 篇

2604.18292 2026-04-21 cs.AI cs.CL 91%

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

Agent-World:通过可扩展环境提升进化通用智能

Guanting Dong, Junting Lu, Junjie Huang, Wanjun Zhong, Longxiang Liu, Shijue Huang, Zhenyu Li, Yang Zhao, Xiaoshuai Song, Xiaoxi Li, Jiajie Jin, Yutao Zhu, Hanbin Wang, Fangyu Lei, Qinyu Luo, Mingyang Chen, Zehui Chen, Jiazhan Feng, Ji-Rong Wen, Zhicheng Dou

机构 * Renmin University of China(中国人民大学) ByteDance Seed(字节跳动种子)

专题命中 Agent评测 :agent(title,title_cn);agentic(abstract);分类 cs.AI、cs.CL

AI总结 本文提出Agent-World,通过可扩展环境实现通用智能进化,展示其在23个挑战性基准测试中超越现有模型和基线。

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25944 2026-04-21 cs.AI 90%

NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving

NuRisk:面向自动驾驶中 agent 级风险评估的视觉问答数据集

Yuan Gao, Mattia Piccinini, Roberto Brusnicki, Yuchen Zhang, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自主车辆系统教授职位,TUM工程与设计学院,慕尼黑技术大学) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所(MIRMI))

专题命中 Agent评测 :agent(title,title_cn);分类 cs.AI

AI总结 NuRisk 数据集通过 2.9K 场景和 1.1M agent 级样本,提供基于鸟瞰图的时序图像及量化风险标注,提升自动驾驶中 agent 行为与情境的时空推理能力。

Comments 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17308 2026-04-21 cs.AI 87%

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

SkillFlow:自主代理终身技能发现与演化的基准测试

Ziao Zhang, Kou Shi, Shiting Huang, Avery Nie, Yu Zeng, Yiming Zhao, Zhen Fang, Qishen Su, Haibo Qiu, Wei Yang, Qingnan Ren, Shun Zou, Wenxuan Huang, Lin Chen, Zehui Chen, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) University of Toronto(多伦多大学) University of Sydney(悉尼大学)

专题命中 Agent评测 :autonomous agent(title,abstract);agent(abstract);workflow(abstract);agentic(abstract)

AI总结 SkillFlow通过166个任务测试自主代理能否发现、修复和维护技能库,揭示终身学习中技能演化的能力差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17133 2026-04-21 cs.AI cs.CR 87%

If Only My CGM Could Speak: A Privacy-Preserving Agent for Question Answering over Continuous Glucose Data

若我的连续葡萄糖监测仪能说话:一个隐私保护的问答代理

Yanjun Cui, Ali Emami, Temiloluwa Prioleau, Nikhil Singh

机构 * Dartmouth College(达特茅斯学院) Emory University(埃默里大学)

专题命中 Agent评测 :agent(title,summary_cn);分类 cs.AI

AI总结 本文提出CGM-Agent框架,通过本地计算实现隐私保护的连续葡萄糖数据问答,评估显示顶级模型在合成和现实查询中准确率分别为94%和88%。

Comments Accepted by ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16729 2026-04-21 cs.CV cs.AI 87%

Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis

基于代理的大型语言模型用于无训练神经放射学图像分析

Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C. Peeken

机构 * Department of Radiation Oncology, TUM University Hospital, Munich, Germany(放射肿瘤科,慕尼黑技术大学医院,德国) Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM)(人工智能在医疗和健康领域的主任,慕尼黑技术大学) Chair for AI for Image-Guided Diagnosis and Therapy, Technical University of Munich (TUM)(人工智能在影像引导诊断和治疗领域的主任,慕尼黑技术大学) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心(MCML),德国慕尼黑) Department of Computing, Imperial College London, London, UK(计算系,伦敦帝国学院,英国伦敦) Deutsches Konsortium für Translationale Krebsforschung (DKTK), Partner Site Munich, Munich, Germany(德国转化癌症研究联盟(DKTK),慕尼黑分部,德国慕尼黑)

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);tool use(abstract);multi-agent(abstract)

AI总结 本文提出一种无训练代理框架,利用外部工具实现脑MRI分析,涵盖预处理、病灶分割和体积分析,并通过公开BraTS数据集评估代理AI在神经放射学任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14691 2026-04-21 cs.AI cs.CL cs.CY 87%

CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations

CAMO:一种用于从微观行为到宏观涌现的LLM代理模拟自动因果发现的框架

Xiangning Yu, Yuwei Guo, Yuqi Hou, Xiao Xue, Qun Ma

机构 * College of Intelligence and Computing, Tianjin University(天津大学智能计算学院) Tianjin Key Laboratory of Healthy Habitat and Smart Technology(天津健康人居环境与智能技术重点实验室) Laboratory of Computation and Analytics of Complex Management Systems, Tianjin University(复杂管理系统计算与分析实验室)

专题命中 Agent评测 :agent(title,abstract);agentic(title);分类 cs.AI、cs.CL

AI总结 CAMO通过分析LLM代理模拟中的微观行为,自动发现导致宏观结果的因果机制,提供可解释的因果链和干预手段,提升对社会涌现的理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16762 2026-04-21 cs.CR cs.AI 86%

CapSeal: Capability-Sealed Secret Mediation for Secure Agent Execution

CapSeal:能力密封的秘密调解用于安全的代理执行

Shutong Jin, Ruiyi Guo, Ray C. C. Cheung

机构 * City University of Hong Kong(香港城市大学) Beijing Foreign Studies University(北京外国语大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(abstract);agentic(abstract);分类 cs.AI

AI总结 CapSeal通过本地可信代理限制秘密访问,解决代理执行中秘密泄露问题,提供非导出动作能力。

Comments 11 pages, 5 figures. Research preprint on secure secret mediation for agent systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17989 2026-04-21 cs.AI 85%

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum

AIT Academy:通过儒家三领域课程培养完整智能体

Jiaqi Li, Lvyang Zhang, Yang Zhao, Wen Lu, Lidong Zhai

机构 * Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学网络安全学院)

专题命中 Agent评测 :agent(title,abstract);AI agent(abstract);tool use(abstract);分类 cs.AI

AI总结 本文提出AIT Academy框架,通过儒家三领域课程培养完整智能体,实验显示多领域视角在安全意识校准中有诊断价值。

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18240 2026-04-21 cs.AI 83%

AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

AJ-Bench:用于环境感知评估的代理作为裁判基准测试

Wentao Shi, Yu Wang, Yuyang Zhao, Yuxin Chen, Fuli Feng, Xueyuan Hao, Xi Su, Qi Gu, Hui Su, Xunliang Cai, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Meituan(美团)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 AJ-Bench通过三个领域155个任务评估代理作为裁判的能力,揭示了基于代理的验证方法在信息获取和过程验证中的性能提升及挑战。

Comments Accepted to ACL 2026 Findings. 43 pages total, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16753 2026-04-21 cs.AI 83%

Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs

何时信任技能:单智能体LLM的延迟评估与认知警惕

Eren Unlu

机构 * Globeholder Paris, France(巴黎Globeholder机构)

专题命中 Agent评测 :agent(title,abstract);autonomous agent(abstract);分类 cs.AI

AI总结 本文提出MESA-S框架,通过延迟评估和认知警惕机制提升单智能体LLM的可靠性,减少供应链漏洞并防止自信膨胀。

Comments 7 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17883 2026-04-21 cs.SE cs.HC cs.LG 82%

Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer

规模化的人工智能编码协作需要一个可治理的共识层

Tianfu Wang, Zhezheng Hao, Yin Wu, Wei Wu, Qiang Lin, Hande Dong, Nicholas Jing Yuan, Hui Xiong

机构 * Tencent(腾讯)

专题命中 Agent评测 :agentic(summary_cn,abstract);分类 cs.LG、cs.SE

AI总结 本文提出Agentic Consensus共识层,通过类型属性图替代代码作为主要工程产物,解决AI辅助开发中代码与聊天历史维度塌陷导致的系统不透明问题,强调通过共识熵衡量协作流程的对齐度和干预距离。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16385 2026-04-21 cs.SE cs.AI 81%

StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability

StressWeb: 一个用于评估网络代理在现实交互变异下的鲁棒性的诊断基准

Haoyue Bai, Dong Wang, Long Chen, Bingguang Hao, Pengyang Shao, Yonghui Yang, Yicheng He, Chenyi Zhuang

机构 * Inclusion AI, Ant Group(Inclusion AI,蚂蚁集团) National University of Singapore(新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.SE

AI总结 本文提出StressWeb基准,通过构建可控的网络环境和引入结构化扰动,评估网络代理在现实交互中的鲁棒性,揭示隐藏的失败模式和鲁棒性差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17557 2026-04-21 cs.LO cs.AI 79%

Causal-Temporal Event Graphs: A Formal Model for Recursive Agent Execution Traces

因果-时间事件图:递归代理执行记录的正式模型

Simon Foldvik

机构 * Independent researcher(独立研究者)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 本文提出因果-时间事件图(CTEG)作为递归代理执行记录的正式模型,通过递归闭包和单调算子的最小固定点,实现执行轨迹的结构化表示与验证。

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17309 2026-04-21 cs.AI 79%

Knows: Agent-Native Structured Research Representations

Knows: 代理原生结构化研究表示

Guangsheng Yu, Xu Wang

机构 * Independent Researcher(独立研究者)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI

AI总结 Knows通过轻量级规范将结构化声明、证据、溯源和可验证关系绑定到研究成果,提升LLM代理处理长文档的能力,实验显示使用侧车可显著提升弱模型准确率,且减少输入token数量。

Comments This paper serves as a technical report/white paper for the Knows.Academy project (https://knows.academy)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17238 2026-04-21 cs.CL 79%

Personalizing Student-Agent Interactions Using Log-Contextualized Retrieval-Augmented Generation (RAG)

基于日志上下文的检索增强生成在个性化学生代理互动中的应用

Clayton Cohn, Surya Rayala, Caitlin Snyder, Joyce Fonteles, Shruti Jain, Naveeduddin Mohammed, Umesh Timalsina, Sarah K. Burriss, Ashwin T S, Namrata Srivastava, Menton Deweese, Angela Eeds, Gautam Biswas

机构 * 1Department of Computer Science, Vanderbilt University, Nashville, USA 2College of Engineering \& Science, University of Detroit Mercy, Detroit, USA 3The School for Science Math, Vanderbilt University, Nashville, USA

专题命中 Agent评测 :agent(title,abstract);分类 cs.CL

AI总结 本文提出日志上下文化检索增强生成(LC-RAG)方法,通过环境日志增强协作对话的检索能力,使协作同伴代理Copa能提供个性化指导,支持学生在C2STEM环境中的批判性思维和知识决策。

Comments Peer reviewed; appeared in the International Conference on Artificial Intelligence in Education (AIED25) Workshop on Epistemics and Decision-Making in AI-Supported Education

Journal ref https://sites.google.com/view/edm-aied-2025/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16342 2026-04-21 cs.HC 78%

SAGE: Sensor-Augmented Grounding Engine for LLM-Powered Sleep Care Agent

SAGE:基于传感器的地面引擎用于LLM驱动的睡眠护理代理

Hansoo Lee, Yoonjae Cho, Sonya S. Kwak, Rafael A. Calvo

专题命中 Agent评测 :agent(title,abstract)

AI总结 SAGE通过整合传感器数据,解决睡眠护理中数据与行动之间的鸿沟问题,提升个性化和信任度。

Comments Accepted to the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26). 6 pages

Journal ref Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14240 2026-04-21 cs.AI 77%

LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

LiveResearchBench: 一个用于真实世界中以用户为中心的深度研究的实时基准

Jiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen, Austin Xu, Zixuan Ke, Frederic Sala, Aws Albarghouthi, Caiming Xiong, Shafiq Joty

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Stanford University(斯坦福大学) Salesforce AI Research(Salesforce AI研究院)

专题命中 Agent评测 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.AI

AI总结 本文提出LiveResearchBench,一个包含100个专家精选任务的实时基准,涵盖日常生活、企业与学术领域,要求进行广泛、动态的实时网络搜索与综合。通过DeepEval评估17个前沿深度研究系统,揭示其优势、失败模式及关键组件。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12261 2026-04-21 cs.CL cs.AI 76%

Infherno: End-to-end Agent-based FHIR Resource Synthesis from Free-form Clinical Notes

Infherno:从非结构化临床笔记到端到端的FHIR资源合成

Johann Frei, Nils Feldhus, Lisa Raithel, Roland Roller, Alexander Meyer, Frank Kramer

机构 * IT-Infrastructure for Translational Medical Research(转化医学研究信息基础设施) BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所) Technische Universität Berlin(柏林技术大学) German Research Center for Artificial Intelligence (DFKI), Berlin(德国人工智能研究中心(DFKI),柏林) IKIM, Charité - Universitätsmedizin Berlin(IKIM,柏林夏里特大学医学中心)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL

AI总结 本文提出Infherno框架,利用LLM代理、代码执行和医疗术语数据库工具,解决从非结构化临床笔记到FHIR资源的端到端合成问题,实现与人类基线的竞争力。

Comments EACL 2026 System Demonstrations | Code: https://github.com/j-frei/Infherno | Demo: https://infherno.misit-augsburg.de

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16493 2026-04-21 cs.DB cs.AI cs.CL cs.LG 75%

NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions

NL2SQLBench: 一个模块化评估和基准测试框架,用于LLM驱动的NL2SQL解决方案

Shizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta, Peng Lu, Beng Chin Ooi

机构 * National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学)

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出NL2SQLBench框架,用于评估LLM驱动的NL2SQL方法,通过三个核心模块分析现有策略并提出新指标,评估十种开源方法,揭示现有方法的准确性和效率问题,为未来创新提供指导。

Comments The paper is accepted by VLDB 2026

Journal ref PVLDB, 19(5): 1001 - 1015, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10981 2026-04-21 cs.AI cs.IR 74%

ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks

ATANT v1.1:针对内存、长上下文和代理记忆基准的连续性评估

Samuel Sameer Tanguturi

机构 * Kenotic Labs(肯otic实验室)

专题命中 Agent评测 :agentic(title);分类 cs.AI

AI总结 本文基于ATANT v1.0,分析了现有内存基准与连续性定义的不匹配,指出各基准仅覆盖少量连续性属性,提出连续性评估的校准对比例。

Comments Companion paper to arXiv:2604.06710 (ATANT v1.0). 12 pages, 1 table, 2 appendices. Related-work extension; does not modify the v1.0 standard

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17019 2026-04-21 cs.AI 74%

Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents

Mini-BEHAVIOR-Gran:揭示指令粒度对语言引导的具身智能体的影响

Sukai Huang, Chenyuan Zhang, Fucai Ke, Zhixi Cai, Gholamreza Haffari, Lizhen Qu, Hamid Rezatofighi

机构 * Faculty of Information Technology, Monash University(信息技术学院,莫纳什大学)

专题命中 Agent评测 :agent(abstract,comments);planning(abstract,comments);分类 cs.AI

AI总结 本文提出Mini-BEHAVIOR-Gran基准,通过多级指令变体研究指令粒度对具身智能体性能的影响,发现粒度与性能呈非单调U型关系,粗粒度性能反弹与浅层 grounding 有关。

Comments 23 pages, Keywords: Language Grounding, Language Granularity, Instruction Following Agent, Width-based Planning Research Area: Multimodality and Language Grounding to Vision, Robotics and Beyond Research Area Keywords: vision language navigation, multimodality, neurosymbolic approaches

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18224 2026-04-21 cs.SE cs.AI 73%

WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models

WebCompass:迈向多模态网络编码评估的代码语言模型

Xinping Lei, Xinyu Che, Junqi Xiong, Chenchen Zhang, Yukai Huang, Chenyu Zhou, Haoyang Huang, Minghao Liu, Letian Zhu, Hongyi Ye, Jinhua Hao, Ken Deng, Zizheng Zhan, Han Li, Dailin Li, Yifan Yao, Ming Sun, Zhaoxiang Zhang, Jiaheng Liu

机构 * Nanjing University(南京大学) Kuaishou Technology(快手科技)

专题命中 Agent评测 :agent(abstract,abstract_cn);分类 cs.AI、cs.SE

AI总结 WebCompass提出一个多模态基准测试,用于评估代码语言模型在网络工程中的能力,涵盖生成、编辑和修复三种任务类型,通过多阶段人机协作流程,发现闭源模型在编辑和修复方面表现更优,但美学仍是开放源模型的主要瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17871 2026-04-21 cs.HC 71%

Design and Evaluation of a Culturally Adapted Multimodal Virtual Agent for PTSD Screening

面向PTSD筛查的跨文化多模态虚拟代理设计与评估

Cengiz Ozel, Waleed Nadeem, Samuel Potter, Yahya Bokhari, Bdour Alwuqaysi, Wejdan Alotaibi, Rahaf Fahad Alnufaie, Sabri Boughorbel, Abdulrhman Aljouie, Rakan Altasan, Ehsan Hoque

专题命中 Agent评测 :agent(title)

AI总结 本文设计并评估了适用于军事医疗场景的跨文化多模态虚拟代理Molhim,通过可配置的对话流程实现特定目的的交互,支持结构化多轮对话和自动会后分析,集成PCL-5量表进行PTSD筛查。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28166 2026-04-21 cs.CR cs.AI 70%

Evaluating Privilege Usage of Agents with Real-World Tools

评估具有现实工具的代理的特权使用

Quan Zhang, Lianhang Fu, Lvsi Lian, Gwihwan Go, Yujue Wang, Chijin Zhou, Yu Jiang, Geguang Pu

机构 * Xinjiang University(新疆大学) Tsinghua University(清华大学)

专题命中 Agent评测 :agent(abstract);tool use(abstract);分类 cs.AI

AI总结 本文提出GrantBox,一个用于评估代理特权使用的安全沙盒,通过集成真实工具和允许LLM代理调用真实特权,评估代理在提示注入攻击下的安全能力,发现LLM在面对复杂攻击时平均攻击成功率高达84.80%。

Comments Accepted to the FSE 2026 Ideas, Visions, and Reflections track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05523 2026-04-21 cs.AI 70%

Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition

Market-Bench: 在经济与贸易竞争中评估大语言模型

Yushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 本文提出Market-Bench,通过经济贸易竞争评估大语言模型在经济相关任务中的能力,发现LLM在竞争中表现差异显著,只有少数能持续获利。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05073 2026-04-21 cs.AI 70%

Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities

大语言模型代理中的不确定性量化:基础、新兴挑战与机遇

Changdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Xuefeng Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Carnegie Mellon University(卡内基梅隆大学) Nanyang Technological University(南洋理工大学) University of Pennsylvania(宾夕法尼亚大学) University of Southern California(南加州大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI

AI总结 本文探讨了大语言模型代理中不确定性量化的基础、挑战及未来方向,提出新的框架并分析了现实场景中的技术难题。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17143 2026-04-21 cs.LG 70%

SeekerGym: A Benchmark for Reliable Information Seeking

SeekerGym:一个评估可靠信息检索的基准

Remy Kim, Minseung Lee, Shuo Li, Osbert Bastani

机构 * University of Pennsylvania(宾夕法尼亚大学) Curation Labs Google DeepMind(谷歌DeepMind)

专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.LG

AI总结 SeekerGym旨在评估AI代理检索信息的完整性,通过文档检索任务测试代理对信息完整性的量化能力,实验显示现有模型在Wikipedia和ML调研论文上分别检索出42.5%和29.2%的段落,仍有改进空间。

详情

展开后加载摘要…

URL PDF HTML 收藏