The Surface You Test Is Not the Surface That Breaks
测试的表面并非断裂的表面
Shifat E Arman, Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin
机构
*
Department of Robotics and Mechatronics Engineering, University of Dhaka(达卡大学机器人与机电工程系)
;
Department of Computer Science and Engineering, University of Dhaka(达卡大学计算机科学与工程系)
专题命中
提示注入
:prompt injection(abstract);分类 cs.AI
AI总结
本文发现工具增强的LLM代理对提示注入的脆弱性依赖于攻击表面(工具输出 vs 工具描述),提出自适应攻击率并强调评估需报告每个表面的脆弱性。
Comments8 Figures, 8 Tables, Under Review at EMNLP
CommentsWe have discovered a critical error in the normalized entropy calculation that may have substantially inflated nearly all results herein. We have since fixed this error in a new work, but we believe that the new work is sufficiently dissimilar in focus, methods, dataset, and results as to be misleading if presented as a simple replacement. As such, we propose removal and retraction instead
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
当LLM学会一致错误:合成欺骗的线性表示的多模型研究
Vahideh Zolfaghari
机构
*
Algoverse AI Research
;
Medical Sciences Education Research Center, Mashhad University of Medical Sciences(马什哈德大学医学科学教育研究中心)
;
Student Research Committee, Department of Health Information Technology and Management, Medical Informatics, School of Allied Medical Sciences, Shahid Beheshti University of Medical Sciences(谢赫·贝赫什提大学医学科学学院学生研究委员会,健康信息科技与管理系,医学信息学)
What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness
什么使LVLMs更少产生幻觉?揭示影响幻觉鲁棒性的架构因素
Yusheng He, Jizhe Zhou, Xia Du, Zheng Lin, Jun Luo, Jiancheng Lv
机构
*
School of Computer Science, Engineering Research Center of Machine Learning and Industry Intelligence, Sichuan University(计算机科学学院,机器学习与产业智能工程研究中心,四川大学)
;
School of Computer and Information Engineering, Xiamen University of Technology(计算机与信息工程学院,厦门理工大学)
;
Department of Electrical and Computer Engineering, University of Hong Kong(电气与计算机工程系,香港大学)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
Benchmarking Uncertainty and its Disentanglement in multi-label Chest X-Ray Classification
多标签胸部X光分类中的不确定性及其解缠基准测试
Simon Baur, Wojciech Samek, Jackie Ma
机构
*
Fraunhofer Heinrich-Hertz-Institut(弗劳恩霍夫海因里希-赫兹研究所)
;
Technische Universität Berlin(柏林技术大学)
;
The Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林学习与数据基础研究所)
dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment
dashi: 一个用于数据集偏移表征以支持可信AI开发和部署的Python库
David Fernández-Narro, Pablo Ferri, Ángel Sánchez-García, Juan M. García-Gómez, Carlos Sáez
机构
*
Biomedical Data Science Lab, Instituto Universitario de Tecnologías de la Información y Comunicaciones, Universitat Politècnica de Valéncia(生物医学数据科学实验室,信息与通信技术大学,巴塞罗那理工大学)
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
迈向人性化的社交媒体生态系统:面向安全、自主与福祉的AI增强人机交互设计模式
Mohd Ruhul Ameen, Akif Islam
机构
*
College of Engineering(工程学院)
;
Computer Sciences Marshall University Huntington, WV, USA(计算机科学马歇尔大学亨廷顿州威斯康星州)
;
Department of Computer Science(计算机科学系)
;
Engineering University of Rajshahi Rajshahi 6205, Bangladesh(工程 Rajshahi 大学 Rajshahi 6205 巴基斯坦)
Comments6 pages, 5 tables, 7 figures, and 2 algorithm tables. Accepted at International Conference on Signal Processing, Information, Communication and Systems (SPICSCON 2025)
Journal ref2025 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON)
机构
*
Department of Electrical Engineering, Indian Institute of Technology Delhi(印度理工学院德里分校电子工程系)
;
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里分校人工智能学院)
;
Hewlett Packard Enterprise, India(印度惠普企业公司)
;
Department of Computer Science & Engineering, Indian Institute of Technology Kharagpur(印度理工学院Khargapur分校计算机科学与工程系)
;
A.K.Choudhury School of Information Technology, University of Calcutta, India(印度加尔各答大学信息科技学院)
;
Department of Computer Science & Engineering, Indian Institute of Technology Bombay(印度理工学院孟买分校计算机科学与工程系)
;
Department of Computer Science, Ashoka University, India(阿什oka大学计算机科学系)
;
Department of Computer Science & Engineering, Indian Institute of Technology Jodhpur(印度理工学院朱罗普分校计算机科学与工程系)
CommentsAccepted to RLEval @ ACM CAIS 2026 (Workshop on Methods and RL Environments for Evaluating AI Agents) and selected for an invited talk based on reviewer ratings. 4-page short paper + appendix
机构
*
Department of Computer Science and Engineering, College of Engineering, Qatar University(卡塔尔大学计算机科学与工程系)
;
College of Science and Engineering, Hamad Bin Khalifa University (HBKU)(哈马德·本·卡伊夫大学(HBKU)理学院)
;
Primary Health Care Corporation (PHCC)(初级卫生保健公司)
;
Birmingham City University(伯明翰城市大学)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
多目标优化中梯度聚合的统一框架
Zeou Hu, Kelvin Ho, Yaoliang Yu
机构
*
Cheriton School of Computer Science(切尔顿计算机科学学院)
;
University of Waterloo(滑铁卢大学)
;
Vector Institute(向量研究所)
;
The Chinese University of Hong Kong(香港中文大学)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
福利、可改进性与方差:最优基准测试项聚合的主-代理方法
Andreas Haupt, Justin Hartenstein, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
机构
*
Department of Economics & Computer Science(经济与计算机科学系)
;
Institute for Computational and Mathematical Engineering(计算与数学工程研究所)
;
Department of Computer Science(计算机科学系)
;
Department of Aeronautics & Astronautics(航空与航天系)
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
通过项目反应理论诊断LLM作为评判者的可靠性
Junhyuk Choi, Sohhyung Park, Chanhee Cho, Hyeonchu Park, Bugeun Kim
机构
*
Department of Artificial Intelligence, Chung-Ang University, Seoul, Republic of Korea(Chung-Ang 大学人工智能系)
;
Department of Industrial Engineering, Seoul National University, Seoul, Republic of Korea(首尔国立大学工业工程系)