CommentsWe propose SCOUT, a detector allocation framework that predicts each detector's accuracy and latency on a given input before running it, letting operators control the safety-utility trade-off with a single threshold and route to an LLM judge only when needed
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
SkillSieve:一种用于检测恶意AI代理技能的分层分流框架
Yinghan Hou, Zongyou Yang
机构
*
Department of Earth Science and Engineering(地球科学与工程系)
;
Imperial College London(帝国理工学院伦敦分校)
;
Department of Computer Science(计算机科学系)
;
University College London(伦敦大学学院)
;
Lingban Technology Co., Ltd.(灵伴科技有限公司)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
The LLMbda Calculus: AI Agents, Conversations, and Information Flow
LLMlambda 计算:人工智能代理、对话与信息流
Zac Garby, Andrew D. Gordon, David Sands
机构
*
University of Nottingham, UK(诺丁汉大学)
;
University of Edinburgh, UK(爱丁堡大学)
;
Chalmers University of Technology(查尔姆斯理工大学)
;
The University of Gothenburg, Sweden(哥德堡大学)
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
这是一个陷阱!面向网络代理的任务重定向说服基准
Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki, Piotr Błaszczyk, Will Howard, Lukas Aichberger, Chris Russell, Philip H. S. Torr, Adam Mahdi, Adel Bibi
Comments21 pages, 4 figures, 5 tables. Substantially revised: title, framing and several v1 results changed. Adds a coverage sweep and a separability analysis; corrects the DPO configuration, the density-accuracy correlation and the qualitative examples. Code and data: https://huggingface.co/datasets/overthelex/citation-grounding-eval
PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework
PCS-UQ:基于可预测性-可计算性-稳定性框架的不确定性量化
Abhineet Agarwal, Fange Xiao, Rebecca Barter, Omer Ronen, Boyu Fan, Bin Yu
机构
*
Department of Statistics, University of California, Berkeley(加州大学伯克利分校统计学系)
;
Department of Epidemiology, University of Utah(犹他大学流行病学系)
;
Department of Electrical Engineering and Computer Science, University of California, Berkeley(加州大学伯克利分校电气工程与计算机科学系)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
通过梯度上升实现可解释人格控制与提示工程的桥梁
Harshvardhan Saini, Yiming Tang, Dianbo Liu
机构
*
Department of Computer Science(计算机科学系)
;
Indian Institute of Technology (ISM), Dhanbad(印度理工学院(ISM),丹巴德)
;
National University of Singapore(新加坡国立大学)
;
CIFAR Fellow(CIFAR研究员)
Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers
多模态生成式引擎优化:针对视觉-语言模型排序器的排名操纵
Yixuan Du, Chenxiao Yu, Haoyan Xu, Ziyi Wang, Yue Zhao, Xiyang Hu
机构
*
Georgetown University(乔治城大学)
;
University of Southern California(南加州大学)
;
University of Maryland, College Park(马里兰大学学院公园分校)
;
Arizona State University(亚利桑那州立大学)