Reasoning Up the Instruction Ladder for Controllable Language Models
在指令阶梯上推理以实现可控语言模型
Zishuo Zheng, Vidhisha Balachandran, Chan Young Park, Faeze Brahman, Sachin Kumar
机构
*
Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系)
;
Microsoft Research(微软研究院)
;
Allen Institute for AI(艾伦人工智能研究所)
CommentsThe manuscript was submitted under an inappropriate category. In addition, substantial updates and improvements are currently being made to the document. To avoid confusion and ensure that readers access the most accurate version of the work, we request withdrawal of the current manuscript
When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
当提示控制机器人时:多智能体机器人系统中的提示注入攻击
Neha Nagaraja, Amisha Bagari, Hayretdin Bahsi
机构
*
School of Informatics, Computing, and Cyber Systems, Northern Arizona University(北亚利桑那大学信息学、计算与网络安全学院)
;
Department of Software Science, Tallinn University of Technology(塔林理工大学软件科学系)
What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection
如果提示注入从未消失?探索智能体系统中的跨会话存储提示注入
Yuanbo Xie, Wenlei Zhu, Tianyun Liu, Yingjie Zhang, Suchen Liu, Yulin Li, Liya Su, Tingwen Liu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
;
AI Sec Lab, Beijing Chaitin Technology Co.,Ltd(北京柴坦科技有限公司AI安全实验室)
Commentsv4 (31 pp, up from 27): adds Sec. 2.4 concurrent-work map (13 papers), Sec. 4.3 marked implemented (d034 ships in v0.7.3), Sec. 5.7 composed-stack adaptive-attack partial run, Sec. 5.8 50-doc held-out benchmark (d027 1.000->0.000 F1 in isolation, composed engine 0.815 F1), Sec. 7 architectural patterns. Repro tag: v0.7.3. Zenodo DOI 10.5281/zenodo.19644135
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
通道卫士:安全模型无法组合成安全的多智能体系统
Elias Hossain, Md Mehedi Hasan Nipu, Fatema Tuj Johora Faria, Tasfia Nuzhat Ornee, Maleeha Sheikh
机构
*
College of Engineering and Computer Science, University of Central Florida(工程与计算机科学学院,中央佛罗里达大学)
;
Department of Computer Science and Engineering, North South University(计算机科学与工程系,北南大学)
;
Computer Science and Engineering, Ahsanullah University of Science and Technology(计算机科学与工程,阿沙努拉科学与技术大学)
;
Department of Electrical and Computer Engineering, Purdue University Fort Wayne(电气与计算机工程系,普渡大学弗拉特沃恩分校)
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
谁买单?面向真实世界网络代理的以利益相关者为中心的提示注入基准测试
Zihao Wang, Yiming Li, Yutong Wu, Kangjie Chen, Zheyu Liu, Fok Kar Wai, Pin-Yu Chen, Vrizlynn L. L. Thing, Bo Li, Dacheng Tao, Tianwei Zhang
机构
*
Nanyang Technological University, Singapore(南洋理工大学,新加坡)
;
ST Engineering, Singapore(ST工程,新加坡)
;
IBM Research, USA(IBM研究院,美国)
;
University of Illinois Urbana-Champaign, USA(伊利诺伊大学厄巴纳-香槟分校,美国)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
编译AI:基于LLM的工作流自动化确定性代码生成
Geert Trooskens, Aaron Karlsberg, Anmol Sharma, Lamara De Brouwer, Max Van Puyvelde, Matthew Young, John Thickstun, Gil Alterovitz, Walter A. De Brouwer
机构
*
XY.AI Labs, Palo Alto, CA(XY.AI实验室,帕洛阿尔托)
;
Stanford University School of Medicine, Stanford, CA(斯坦福大学医学院,斯坦福)
;
Cornell University, Department of Computer Science, Ithaca, NY(康奈尔大学计算机科学系,伊萨卡)
;
Brigham and Women’s Hospital / Harvard Medical School, Boston, MA(布莱根妇女医院/哈佛医学院,波士顿)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
AgentMisalignment:衡量基于LLM的代理中失调行为的倾向性
Akshat Naik, Emma Gouné, Patrick Quinn, Guillermo Bosch, Francisco Javier Campos Zabala, Jason Ross Brown, Edward James Young
机构
*
Department of Computer Science(计算机科学系)
;
University of Oxford(牛津大学)
;
Institute of Intelligent Systems and Robotics(智能系统与机器人研究所)
;
Sorbonne Université(索邦大学)
;
The Leverhulme Centre for the Future of Intelligence(未来智能中心)
;
University of Cambridge(剑桥大学)
;
Independent Researcher(独立研究者)
;
Department of Computer Science and Technology(计算机科学与技术系)
;
Department of Engineering(工程系)
CommentsOral at the ICML 2026 Workshop on the Impact of Memorization on Trustworthy Foundation Models; Code available at https://github.com/Graph-COM/KVEraser