Future Mining: Learning for Safety and Security
未来采矿:安全与安全的学习
专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)
AI总结 本文提出了一种统一的智能安全和安全架构,整合多模态感知、安全联邦学习、强化学习、DTN通信和能量感知传感,以提升采矿环境中的安全性和可靠性。
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
未来采矿:安全与安全的学习
专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)
AI总结 本文提出了一种统一的智能安全和安全架构,整合多模态感知、安全联邦学习、强化学习、DTN通信和能量感知传感,以提升采矿环境中的安全性和可靠性。
可信的代理AI需要确定性的架构边界
专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract)
AI总结 本文提出三位一体防御架构,通过确定性的架构执行解决代理AI的安全问题,强调架构调解对可信AI应用的必要性。
MemAdapter:通过生成子图检索实现跨代理记忆范式的快速对齐
机构 * The University of Manchester(曼彻斯特大学) ; Stanford University(斯坦福大学) ; Imperial College London(伦敦帝国理工学院) ; National Institute of Advanced Industrial Science(国家先进工业科学与技术研究院)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 MemAdapter通过生成子图检索实现跨代理记忆范式的快速对齐,提升记忆检索灵活性并降低对齐成本,实验显示其性能优于现有系统。
流畅但异域:即使区域LLM也缺乏文化契合
机构 * Cornell University(康奈尔大学) ; Microsoft Research(微软研究院)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
AI总结 研究发现,即使区域LLM也缺乏文化契合,需通过社区基础的数据和厚宽评估来构建真正主权的LLM。
Comments Under review
AQAScore: 通过音频问答评估文本到音频生成中的语义对齐
机构 * National Taiwan University(国立台湾大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 AQAScore通过音频问答评估文本到音频生成的语义对齐,利用大型语言模型的推理能力,有效捕捉语义不一致性。
Comments Manuscript in progress
揭示和弥合MLLMs中的功能感知差距:通过PET-Bench实现原子视觉对齐和分层评估
机构 * School of Biomedical Engineering, Southern Medical University(生物医学工程学院,南方医科大学) ; School of Biomedical Engineering, Shanghai Jiaotong University(生物医学工程学院,上海交通大学) ; Department of Electronic Engineering, Chinese University of Hong Kong(电子工程系,中国香港大学) ; Faculty of Dentistry, The University of Hong Kong(牙科学院,香港大学) ; Department of Nuclear Medicine, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine(核医学科,广州中医药大学第二附属医院) ; PET Center, Department of Nuclear Medicine, Guangdong Provincial People’s Hospital, Southern Medical University(PET中心,核医学科,广东省人民医院,南方医科大学) ; Department of Nuclear Medicine, Nanfang Hospital, Southern Medical University(核医学科,南芳医院,南方医科大学) ; Division of Nuclear Medicine and Molecular Imaging, Geneva University Hospitals(核医学与分子影像学部,日内瓦大学医院) ; Departments of Radiology, Physics, and Biomedical Engineering, The University of British Columbia(放射学、物理和生物医学工程系,不列颠哥伦比亚大学) ; Medical Artificial Intelligence Laboratory, Westlake University(医学人工智能实验室,西湖大学)
专题命中 安全评测 :alignment(title,abstract);safety(abstract)
AI总结 本文提出AVA方法,通过原子视觉对齐解决MLLMs在功能成像中的感知差距,提升诊断准确性14.83%。
Comments 9 pages, 6 figures, 6 tables
TeleAI-Safety: 一种全面的LLM jailbreaking基准测试,面向攻击、防御和评估
专题命中 安全评测 :safety(title,abstract);jailbreak(abstract)
AI总结 TeleAI-Safety提出了一种全面的LLM安全评估基准,整合多种攻击、防御和评估方法,用于系统性评估LLM的安全性及优化防御策略。
IS-Bench:评估由视觉语言模型驱动的具身智能体在日常家务任务中的交互安全性
机构 * Xiaoya Lu 1, 2(某大学) ; Zeren Chen 2, 3(某大学) ; Xuhao Hu 2, 4(某大学) ; Yijin Zhou 2(某大学) ; Weichen Zhang 2(某大学) ; Dongrui Liu 2(某大学) ; Lu Sheng 3(某研究所) ; Jing Shao 2(某研究所)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 IS-Bench通过多模态基准评估视觉语言模型驱动的具身智能体在日常家务任务中的交互安全性,揭示现有智能体缺乏安全意识的问题。
评估大型语言模型安全防护措施对对抗攻击的鲁棒性
专题命中 安全评测 :safety(title,abstract);jailbreak(abstract)
AI总结 本研究评估了大型语言模型安全防护措施对对抗攻击的鲁棒性,发现模型在未见过的提示上表现显著下降,且存在新的失败模式,强调泛化能力的重要性。
Comments 21 pages, 9 figures, 6 tables
PropensityBench: 通过代理方法评估大语言模型中的潜在安全风险
机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) ; University of Texas at Austin(德克萨斯大学奥斯汀分校) ; University of Maryland, College Park(马里兰大学帕克分校)
专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG
AI总结 PropensityBench通过模拟危险能力评估大语言模型的潜在安全风险倾向,揭示模型在压力下可能选择高风险行为的倾向。
通过评估LLM作为裁判来评估LLM对齐
机构 * Yale University(耶鲁大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出AlignEval基准,通过评估LLM作为裁判的能力来衡量其对齐人类偏好的效果,结果优于现有自动评估基准。
Comments NeurIPS 2025 Camera Ready
机构 * Department of Computer Science, National Center for Text Mining, The University of Manchester(计算机科学系,文本挖掘国家中心,曼彻斯特大学) ; Department of Computer Science and Engineering, Islamic University of Technology(计算机科学与工程系,伊斯兰技术大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments Accepted at EMNLP 2025 (Main)
Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
机构 * department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) ; School of Vehicle and Mobility at Tsinghua University(清华大学车辆与移动系统学院)
专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)
专题命中 安全评测 :trustworthy(title,abstract);safety(abstract)
Comments 5 pages, 2 figures
机构 * SKLCCSE, Beihang University(北京航空航天大学智能科学与技术研究中心) ; Zhongguancun Laboratory(中关村实验室) ; The University of Sydney(悉尼大学) ; Henan University of Science(河南科技大学)
专题命中 安全评测 :safety(title,abstract);alignment(abstract)
机构 * Australian Broadcasting Corporation(澳大利亚广播公司)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 52 pages, 10 figures, author typo corrected, abstract typo corrected
机构 * University of Cincinnati(辛辛那提大学) ; King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹大学科学与技术)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
机构 * Rochester Institute of Technology(罗切斯特理工学院) ; U.S. Naval Research Laboratory(美国海军研究实验室) ; University of Missouri-Kansas City(密苏里大学-Kansas城分校) ; Adobe(Adobe公司) ; Meituan(美团) ; Sun Yat-sen University(孙中山大学) ; Purdue University(普渡大学) ; UC Davis(加州大学戴维斯分校) ; University of Rochester(罗切斯特大学) ; Rice University(莱斯大学) ; Meta AI
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025
专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)
Comments 18 pages, 7 figures
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG
Comments Preprint. This work has been submitted to the Technical AI Governance Workshop at ICML 2025 for review
机构 * Department of Artificial Intelligence in Healthcare, International Academia of Biomedical Innovation Technology(人工智能医疗系,国际生物医学创新技术学院) ; National Taiwan University(台湾大学) ; Neuro Industry Research, Neuro Industry, Inc.(神经产业研究,神经产业公司) ; Department of Biomedical Engineering, University of Southern California(生物医学工程系,南加州大学) ; Taiwan Artificial Intelligence Association(台湾人工智能协会)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
机构 * stanford(斯坦福大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Journal ref ICLR DMLR Data-centric Machine Learning Research (2024), ICML DataWorld (2025)
机构 * School of Computing and Information System, the University of Melbourne, Australia(墨尔本大学计算与信息学院) ; School of Computing, FSE, Macquarie University, Australia(麦考瑞大学计算学院)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL 2025 (Main Proceedings)
机构 * Oracle AI ; Indian Institute of Information Technology Ranchi(印度拉奇信息与技术学院) ; TD Securities(TD证券) ; Columbia University(哥伦比亚大学) ; Hanyang University(翰阳大学)
专题命中 安全评测 :safety(title);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published in the Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2025), Industry Track, pages 558-582
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Amazon(亚马逊) ; Fidelity Investments(富达投资)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments 21 pages, 11 figures, 7 tables
机构 * Warren E Hyde Middle School(沃伦·E·海德中学) ; University of Trento(特伦托大学)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments Accepted to ACM FAccT 2025. To be presented in Athens, June 2025, and published in the conference proceedings. Preprint version; final version will appear in the ACM Digital Library
机构 * Unicom Data Intelligence(中国联通数据智能研究所) ; Data Science & Artificial Intelligence Research Institute(数据科学与人工智能研究院)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments 21 pages, 13 figures, 4 tables
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
Comments This work has been accepted for publication in AI and Ethics