arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2508.16135 2025-08-25 cs.LG cs.AI cs.ET eess.IV 62%

Machine Learning in Micromobility: A Systematic Review of Datasets, Techniques, and Applications

Sen Yan, Chinmaya Kaundanya, Noel E. O'Connor, Suzanne Little, Mingming Liu

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 3 tables, and 4 figures, submitted to IEEE Transactions on Intelligent Vehicles

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15847 2025-08-25 cs.CL cs.LG 62%

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns

Mohammed Abu Baker, Lakshmi Babu-Saheer

机构 * Department of Computer Science, Anglia Ruskin University(计算机科学系,安格利亚 Ruskin 大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments 13 pages. Mechanistic analysis of backdoored LLMs (Qwen2.5-3B). Code: https://github.com/mshahoyi/sa_attn_analysis. Base model: unsloth/Qwen2.5-3B-Instruct-unsloth-bnb-4bit. Finetuned models: https://huggingface.co/collections/mshahoyi/simple-sleeper-agents-68a1df3a7aaff310aa0e5336

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08549 2025-08-20 cs.AI cs.CL 62%

GoAI: Enhancing AI Students' Learning Paths and Idea Generation via Graph of AI Ideas

Xian Gao, Zongyun Zhang, Ting Liu, Yuzhuo Fu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01911 2025-08-19 cs.AI cs.CL cs.HC physics.comp-ph 62%

Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning

Yinggan Xu, Hana Kimlee, Yijia Xiao, Di Luo

机构 * NSF Center for Quantum Network(NSF量子网络中心) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICML 2025 Workshop on MAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06141 2025-08-19 cs.LG cs.AI 62%

Emergent Symbol-like Number Variables in Artificial Neural Networks

Satchel Grant, Noah D. Goodman, James L. McClelland

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Journal ref Transactions on Machine Learning Research (TMLR) 2835-8856 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11582 2025-08-18 cs.CL cs.AI 62%

Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models

Qiguang Chen, Dengyun Peng, Jinhao Liu, HuiKang Su, Jiannan Guan, Libo Qin, Wanxiang Che

机构 * LARG, Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(大型语言模型研究组,社会计算与交互机器人研究中心,哈尔滨工业大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10925 2025-08-18 cs.CL cs.AI 62%

gpt-oss-120b & gpt-oss-20b Model Card

OpenAI, :, Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, Che Chang, Kai Chen, Mark Chen, Enoch Cheung, Aidan Clark, Dan Cook, Marat Dukhan, Casey Dvorak, Kevin Fives, Vlad Fomenko, Timur Garipov, Kristian Georgiev, Mia Glaese, Tarun Gogineni, Adam Goucher, Lukas Gross, Katia Gil Guzman, John Hallman, Jackie Hehir, Johannes Heidecke, Alec Helyar, Haitang Hu, Romain Huet, Jacob Huh, Saachi Jain, Zach Johnson, Chris Koch, Irina Kofman, Dominik Kundel, Jason Kwon, Volodymyr Kyrylov, Elaine Ya Le, Guillaume Leclerc, James Park Lennon, Scott Lessans, Mario Lezcano-Casado, Yuanzhi Li, Zhuohan Li, Ji Lin, Jordan Liss, Lily, Liu, Jiancheng Liu, Kevin Lu, Chris Lu, Zoran Martinovic, Lindsay McCallum, Josh McGrath, Scott McKinney, Aidan McLaughlin, Song Mei, Steve Mostovoy, Tong Mu, Gideon Myles, Alexander Neitz, Alex Nichol, Jakub Pachocki, Alex Paino, Dana Palmie, Ashley Pantuliano, Giambattista Parascandolo, Jongsoo Park, Leher Pathak, Carolina Paz, Ludovic Peran, Dmitry Pimenov, Michelle Pokrass, Elizabeth Proehl, Huida Qiu, Gaby Raila, Filippo Raso, Hongyu Ren, Kimmy Richardson, David Robinson, Bob Rotsted, Hadi Salman, Suvansh Sanjeev, Max Schwarzer, D. Sculley, Harshit Sikchi, Kendal Simon, Karan Singhal, Yang Song, Dane Stuckey, Zhiqing Sun, Philippe Tillet, Sam Toizer, Foivos Tsimpourlas, Nikhil Vyas, Eric Wallace, Xin Wang, Miles Wang, Olivia Watkins, Kevin Weil, Amy Wendling, Kevin Whinnery, Cedric Whitney, Hannah Wong, Lin Yang, Yu Yang, Michihiro Yasunaga, Kristen Ying, Wojciech Zaremba, Wenting Zhan, Cyril Zhang, Brian Zhang, Eddie Zhang, Shengjia Zhao

机构 * OpenAI

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10450 2025-08-18 cs.IR cs.AI cs.CL 62%

TokenRec: Learning to Tokenize ID for LLM-based Generative Recommendation

Haohao Qu, Wenqi Fan, Zihuai Zhao, Qing Li

机构 * Department of Computing, The Hong Kong Polytechnic University(计算机系,香港理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by IEEE TKDE. Codes and data are available at https://github.com/Quhaoh233/TokenRec

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03012 2025-08-14 cs.AI cs.CL cs.CV 62%

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

Pegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny, Matthieu Cord

机构 * ISIR, Sorbonne Université(ISIR,索邦大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments ICCV 2025. The first three authors contributed equally. Project page and code: https://pegah- kh.github.io/projects/lmm-finetuning-analysis-and-steering/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09199 2025-08-14 cs.CV cs.AI cs.CL 62%

$Δ$-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation

Jucheng Hu, Suorong Yang, Dongzhan Zhou

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12833 2025-08-13 cs.CV cs.AI cs.LG 62%

SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback

Elior Benarous, Yilun Du, Heng Yang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07484 2025-08-12 cs.CL cs.AI 62%

ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models

Archchana Sindhujan, Shenbin Qian, Chan Chi Chun Matthew, Constantin Orasan, Diptesh Kanojia

机构 * Institute for People-Centred AI and Centre for Translation Studies, School of Computer Science and Electronic Engineering, University of Surrey(以人为本的人工智能研究所和翻译研究中心,计算机科学与电子工程学院,萨里大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to COLM 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22919 2025-08-12 cs.CL cs.AI 62%

A novel language model for predicting serious adverse event results in clinical trials from their prospective registrations

Qixuan Hu, Xumou Zhang, Jinman Kim, Florence Bourgeois, Adam G. Dunn

机构 * School of Computer Science, Faculty of Engineering, University of Sydney(悉尼大学计算机科学学院、工程学院) Computational Health Informatics Program, Boston Children’s Hospital(波士顿儿童医院计算健康信息学项目) Harvard-MIT Center for Regulatory Science and Department of Pediatrics, Harvard Medical School(哈佛-麻省理工监管科学中心和哈佛医学院儿科部门) Sydney School of Public Health, Faculty of Medicine and Health, University of Sydney(悉尼大学公共卫生学院、医学与健康学院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 4 figures. Updated to include Table 2, Supplementary Table 1, and an additional baseline random forest model

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21931 2025-08-12 cs.IR cs.AI cs.CL cs.MA 62%

ARAG: Agentic Retrieval Augmented Generation for Personalized Recommendation

Reza Yousefi Maragheh, Pratheek Vadla, Priyank Gupta, Kai Zhao, Aysenur Inan, Kehui Yao, Jianpeng Xu, Praveen Kanumala, Jason Cho, Sushant Kumar

机构 * Walmart Global Tech(沃尔玛全球科技)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07480 2025-08-12 eess.SP cs.AI cs.LG 62%

EEG-Language Pretraining for Highly Label-Efficient Clinical Phenotyping

Sam Gijsen, Kerstin Ritter

机构 * Charité – Universitätsmedizin Berlin, Department of Psychiatry and Psychotherapy, Berlin, Germany(柏林查理医院医学大学精神病与心理治疗系) Hertie Institute for AI in Brain Health, University of Tübingen, Germany(图宾根大学健康人工智能研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01196 2025-08-12 cs.CL cs.AI 62%

$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models

Zian Su, Ziyang Huang, Kaiyuan Zhang, Xiangyu Zhang

机构 * Purdue University(普渡大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments COLM 2025. The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20758 2025-08-12 stat.AP cs.AI cs.CL 62%

Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth

Seyed Pouyan Mousavi Davoudi, Amin Gholami Davodi, Alireza Amiri-Margavi, Alireza Shafiee Fard, Mahdi Jafari

机构 * Independent Researcher in AI and Statistics(人工智能与统计学独立研究者) Shahrood University of Technology(沙霍罗德大学) University of Pittsburgh(匹兹堡大学) Duquesne University(杜克森大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17797 2025-08-12 cs.AI cs.GT cs.LG cs.MA 62%

Observation Interference in Partially Observable Assistance Games

Scott Emmons, Caspar Oesterheld, Vincent Conitzer, Stuart Russell

机构 * Center for Human-Compatible AI, University of California, Berkeley(人类兼容人工智能中心,加州大学伯克利分校) Foundations of Cooperative AI Lab, Carnegie Mellon University(协作人工智能实验室,卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03155 2025-08-11 cs.LG cs.AI 62%

Fusing Cross-Domain Knowledge from Multimodal Data to Solve Problems in the Physical World

Yu Zheng

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16658 2025-08-11 cs.CL cs.AI 62%

Contextual Reinforcement in Multimodal Token Compression for Large Language Models

Naderdel Piero, Zacharias Cromwell, Nathaniel Wainwright, Matthias Nethercott

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05616 2025-08-08 cs.LG cs.AI cs.NE cs.RO 62%

TrajEvo: Trajectory Prediction Heuristics Design via LLM-driven Evolution

Zhikai Zhao, Chuanbo Hua, Federico Berto, Kanghoon Lee, Zihan Ma, Jiachen Li, Jinkyoo Park

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2505.04480

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05399 2025-08-08 cs.CV cs.AI cs.LG 62%

UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation

Wonjun Kang, Byeongkeun Ahn, Minjae Lee, Kevin Galim, Seunghyuk Oh, Hyung Il Koo, Nam Ik Cho

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Code is available at https://github.com/furiosa-ai/uncage

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05019 2025-08-08 cs.CV cs.AI cs.LG 62%

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes

Sadia Kamal, Tim Oates, Joy Wan

机构 * Department of Computer Science, University of Maryland, Baltimore County(计算机科学系,马里兰大学巴尔的摩县分校) Department of Dermatology, Johns Hopkins University School of Medicine(皮肤科系,约翰霍普金斯大学医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to IJCAI 2025 Workshops. arXiv admin note: substantial text overlap with arXiv:2506.10328

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14964 2025-08-08 cs.CL cs.LG 62%

Efficient Knowledge Injection in LLMs via Self-Distillation

Kalle Kujanpää, Pekka Marttinen, Harri Valpola, Alexander Ilin

机构 * Aalto University(阿alto大学) Finnish Center for Artificial Intelligence (FCAI)(芬兰人工智能中心) System 2 AI

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04571 2025-08-07 cs.IR cs.CL cs.LG 62%

Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation

Claudio Pomo, Matteo Attimonelli, Danilo Danese, Fedelucio Narducci, Tommaso Di Noia

机构 * Sapienza University of Rome(罗马大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted as Full Research Papers at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04073 2025-08-07 cs.CL cs.LG 62%

Efficient Strategy for Improving Large Language Model (LLM) Capabilities

Julián Camilo Velandia Gutiérrez

机构 * Universidad Nacional de Colombia(哥伦比亚国立大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments Based on master's thesis in Systems and Computer Engineering, Universidad Nacional de Colombia (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17963 2025-08-07 cs.LG cs.AI 62%

Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning Tasks

Xingcheng Xu, Zibo Zhao, Haipeng Zhang, Yanqing Yang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ShanghaiTech University(上海科技大学) University of Hong Kong(香港大学) Fudan University(复旦大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted by the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025), Main Conference

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02823 2025-08-06 cs.HC cs.AI cs.CL cs.SE 62%

NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification

Wenshuo Zhang, Leixian Shen, Shuchang Xu, Jindu Wang, Jian Zhao, Huamin Qu, Linping Yuan

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Waterloo(滑铁卢大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted in UIST 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19151 2025-08-05 cs.RO cs.AI cs.LG cs.MA 62%

ReCoDe: Reinforcement Learning-based Dynamic Constraint Design for Multi-Agent Coordination

Michael Amir, Guang Yang, Zhan Gao, Keisuke Okumura, Heedo Woo, Amanda Prorok

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments To appear in CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00674 2025-08-04 cs.AI cs.HC cs.LG 62%

Context-Aware Visualization for Explainable AI Recommendations in Social Media: A Vision for User-Aligned Explanations

Banan Alkhateeb, Ellis Solaiman

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏