arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9331 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9331 篇

2508.07308 2025-08-12 cs.CL cs.AI cs.IR cs.LG 67%

HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways

Cristian Cosentino, Annamaria Defilippo, Marco Dossena, Christopher Irwin, Sara Joubbi, Pietro Liò

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13677 2025-08-11 cs.CL cs.AI cs.LG 67%

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results

Andrea Santilli, Adam Golinski, Michael Kirchhof, Federico Danieli, Arno Blaas, Miao Xiong, Luca Zappella, Sinead Williamson

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04350 2025-08-07 cs.CL cs.AI cs.CV cs.LG cs.MA 67%

Chain of Questions: Guiding Multimodal Curiosity in Language Models

Nima Iji, Kia Dashtipour

机构 * Edinburgh Napier University(爱丁堡纳皮尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23535 2025-08-01 cs.LG cs.AI cs.CY 67%

Transparent AI: The Case for Interpretability and Explainability

Dhanesh Ramachandram, Himanshu Joshi, Judy Zhu, Dhari Gandhi, Lucas Hartman, Ananya Raval

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21919 2025-07-31 cs.CL cs.AI cs.CY 67%

Training language models to be warm and empathetic makes them less reliable and more sycophantic

Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher

机构 * Oxford Internet Institute(牛津互联网研究所) University of Oxford(牛津大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22034 2025-07-30 cs.AI cs.CL cs.LG 67%

UserBench: An Interactive Gym Environment for User-Centric Agents

Cheng Qian, Zuxin Liu, Akshara Prabhakar, Zhiwei Liu, Jianguo Zhang, Haolin Chen, Heng Ji, Weiran Yao, Shelby Heinecke, Silvio Savarese, Caiming Xiong, Huan Wang

机构 * Salesforce AI Research(Salesforce AI研究部) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 Pages, 17 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20526 2025-07-29 cs.AI cs.CL cs.CY 67%

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Andy Zou, Maxwell Lin, Eliot Jones, Micha Nowak, Mateusz Dziemian, Nick Winter, Alexander Grattan, Valent Nathanael, Ayla Croft, Xander Davies, Jai Patel, Robert Kirk, Nate Burnikell, Yarin Gal, Dan Hendrycks, J. Zico Kolter, Matt Fredrikson

专题命中 安全评测 :red teaming(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19598 2025-07-29 cs.CL cs.AI cs.CR cs.LG 67%

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?

Muntasir Wahed, Xiaona Zhou, Kiet A. Nguyen, Tianjiao Yu, Nirav Diwan, Gang Wang, Dilek Hakkani-Tür, Ismini Lourentzou

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Winner Defender Team at Amazon Nova AI Challenge 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07817 2025-07-16 cs.CL cs.AI cs.LG 67%

On the Effect of Instruction Tuning Loss on Generalization

Anwoy Chatterjee, H S V N S Kowndinya Renduchintala, Sumit Bhatia, Tanmoy Chakraborty

机构 * Dept. of Electrical Engineering(电气工程系) Indian Institute of Technology Delhi(印度理工学院德里) Media and Data Science Research(媒体与数据科学研究) Adobe Inc., India(Adobe公司,印度)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments To appear in Transactions of the Association for Computational Linguistics (TACL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06672 2025-07-15 cs.CY cs.AI cs.LG q-fin.RM 67%

Insuring Uninsurable Risks from AI: Government as Insurer of Last Resort

Cristian Trout

机构 * Independent Researcher, Cambridge Boston Alignment Initiative(独立研究者,剑桥波士顿对齐计划)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to Generative AI and Law Workshop at the International Conference on Machine Learning (ICML 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06895 2025-07-10 cs.CL cs.AI cs.IR cs.LG 67%

SCoRE: Streamlined Corpus-based Relation Extraction using Multi-Label Contrastive Learning and Bayesian kNN

Luca Mariotti, Veronica Guidetti, Federica Mandreoli

机构 * Department of Physical, Computer and Mathematical Sciences - University of Modena and Reggio Emilia(物理、计算机和数学科学系 - 模纳和雷吉奥艾米利亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20589 2025-07-09 cs.CR cs.AI cs.CL cs.LG 67%

LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis

Saeif Alhazbi, Ahmed Mohamed Hussain, Gabriele Oligeri, Panos Papadimitratos

机构 * College of Science and Engineering (CSE), Hamad Bin Khalifa University (HBKU)(哈马德·本·卡伊夫大学科学与工程学院) Networked Systems Security Group, KTH Royal Institute of Technology -- Stockholm, Sweden(瑞典皇家理工学院网络系统安全组)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18045 2025-06-24 cs.CY cs.AI cs.CL 67%

The Democratic Paradox in Large Language Models' Underestimation of Press Freedom

I. Loaiza, R. Vestrelli, A. Fronzetti Colladon, R. Rigobon

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16831 2025-06-23 cs.SE 67%

Accountability of Robust and Reliable AI-Enabled Systems: A Preliminary Study and Roadmap

Filippo Scaramuzza, Damian A. Tamburri, Willem-Jan van den Heuvel

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments To be published in https://link.springer.com/book/9789819672370

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15293 2025-06-19 cs.HC cs.RO 67%

Designing Intent: A Multimodal Framework for Human-Robot Cooperation in Industrial Workspaces

Francesco Chiossi, Julian Rasch, Robin Welsch, Albrecht Schmidt, Florian Michahelles

机构 * Aalto University(阿alto大学) TU Wien(维也纳技术大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments 9 pages

Journal ref The Future of Human-Robot Synergy in Interactive Environments: The Role of Robots at the Workplace @ CHIWork 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15100 2025-06-19 cs.CR 67%

International Security Applications of Flexible Hardware-Enabled Guarantees

Onni Aarne, James Petrie

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10581 2025-06-19 q-bio.QM cs.AI cs.CL cs.LG eess.SP 67%

Large Language Model-informed ECG Dual Attention Network for Heart Failure Risk Prediction

Chen Chen, Lei Li, Marcel Beetz, Abhirup Banerjee, Ramneek Gupta, Vicente Grau

机构 * Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford(生物医学工程研究所,工程科学系,牛津大学) Imperial College London(伦敦帝国学院) University of Sheffield(谢菲尔德大学) Novo Nordisk Research Centre Oxford (NNRCO)(牛津诺和硕研究中心(NNRCO)) Royal Society(皇家学会)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Under journal revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08210 2025-06-17 cs.CV cs.AI cs.CL cs.LG 67%

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

Andrew Z. Wang, Songwei Ge, Tero Karras, Ming-Yu Liu, Yogesh Balaji

机构 * University of Maryland(马里兰大学) NVIDIA(英伟达)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 28575-28585

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12266 2025-06-17 cs.CL cs.AI cs.HC cs.LG 67%

The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs

Avinash Baidya, Kamalika Das, Xiang Gao

机构 * Intuit AI Research(Intuit AI研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025; 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19794 2025-06-12 cs.CV 67%

MVTamperBench: Evaluating Robustness of Vision-Language Models

Amit Agarwal, Srikant Panda, Angeline Charles, Bhargava Kumar, Hitesh Patel, Priyaranjan Pattnayak, Taki Hasan Rafi, Tejaswini Kumar, Hansa Meghwani, Karan Gupta, Dong-Kyu Chae

机构 * Liverpool John Moores University(利物浦约翰·穆里斯大学) Birla Institute of Technology(比拉理工学院) Christ University(基督大学) Columbia University(哥伦比亚大学) New York University(纽约大学) University of Washington(华盛顿大学) Hanyang University(翰阳大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09420 2025-06-12 cs.AI cs.CL cs.HC cs.LG cs.MA 67%

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy

Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu, Chunyu Miao, Dongyuan Li, Aiwei Liu, Yue Zhou, Yankai Chen, Weizhi Zhang, Yangning Li, Liancheng Fang, Renhe Jiang, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Tokyo(东京大学) Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13151 2025-06-10 cs.LG cs.AI cs.CL 67%

MIB: A Mechanistic Interpretability Benchmark

Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to ICML 2025. Project website at https://mib-bench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04089 2025-06-05 cs.LG cs.AI cs.CL cs.RO 67%

AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment

Anastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev, Aleksandr I. Panov

机构 * LMU(慕尼黑大学) MIPT(莫斯科 Institute of Physics and Technology) AIRI(空气研究所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20322 2025-06-04 cs.CL cs.AI cs.CV cs.IR cs.LG 67%

Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms

Mengru Wang, Ziwen Xu, Shengyu Mao, Shumin Deng, Zhaopeng Tu, Huajun Chen, Ningyu Zhang

机构 * Zhejiang University(浙江大学) Tencent(腾讯) National University of Singapore(新加坡国立大学) NUS-NCS Joint Lab(新加坡国立大学NCS联合实验室)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08367 2025-06-04 cs.AI cs.CL cs.CV cs.LG 67%

MCU: An Evaluation Framework for Open-Ended Game Agents

Xinyue Zheng, Haowei Lin, Kaichen He, Zihao Wang, Zilong Zheng, Yitao Liang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref ICML 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18754 2025-06-02 cs.CL cs.AI cs.LG 67%

Few-Shot Optimization for Sensor Data Using Large Language Models: A Case Study on Fatigue Detection

Elsen Ronando, Sozo Inoue

机构 * Graduate School of Life Science and Systems Engineering, Kyushu Institute of Technology(九州工学院生命科学与系统工程研究生院) Department of Informatics, Universitas 17 Agustus 1945 Surabaya(Surabaya 17 August 1945 大学信息系)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 43 pages, 18 figures. Accepted for publication in MDPI Sensors (2025). Final version before journal publication

Journal ref Sensors 2025, 25, 3324

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15210 2025-06-02 cs.LG cs.AI cs.CL 67%

PairBench: Are Vision-Language Models Reliable at Comparing What They See?

Aarash Feizi, Sai Rajeswar, Adriana Romero-Soriano, Reihaneh Rabbany, Valentina Zantedeschi, Spandana Gella, João Monteiro

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22356 2025-05-29 cs.LG cs.AI cs.CY stat.ML 67%

Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings

Angéline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas Papernot

机构 * ETH Zurich(苏黎世联邦理工学院) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19757 2025-05-27 cs.SE cs.AI cs.CL cs.LG 67%

CIDRe: A Reference-Free Multi-Aspect Criterion for Code Comment Quality Measurement

Maria Dziuba, Valentin Malykh

机构 * MTS AI(MTS人工智能公司) ITMO University(ITMO大学) IITU University(IITU大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19696 2025-05-27 cs.CV cs.MM eess.IV 67%

Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality

Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Nour Aburaed, Alessandro Bruno

机构 * R&D Lab F-Initiatives(F-Initiatives研发实验室) Université d’Orleans(奥尔良大学) Université Sorbonne Paris Nord(巴黎-索邦大学) University of Dubai(迪拜大学) IULM University(IULM大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏