arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2510.05862 2025-11-05 cs.CL cs.AI 62%

Revisiting Long-context Modeling from Context Denoising Perspective

Zecheng Tang, Baibei Ji, Juntao Li, Lijun Wu, Haijia Gui, Min Zhang

机构 * Soochow University(苏州大学) LCM Laboratory(长文实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15090 2025-11-05 cs.CL cs.AI 62%

ExpertLens: Activation steering features are highly interpretable

Masha Fedzechkina, Eleonora Gualdoni, Sinead Williamson, Katherine Metcalf, Skyler Seto, Barry-John Theobald

机构 * Apple(苹果公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14846 2025-11-04 cs.AI cs.CL cs.LO 62%

Where to Search: Measure the Prior-Structured Search Space of LLM Agents

Zhuo-Yang Song

机构 * School of Physics, Peking University(物理学院,北京大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 11 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02184 2025-11-04 stat.ML cs.AI cs.CV cs.LG math.ST stat.TH 62%

Double Descent Meets Out-of-Distribution Detection: Theoretical Insights and Empirical Analysis on the role of model complexity

Mouïn Ben Ammar, David Brellmann, Arturo Mendoza, Antoine Manzanera, Gianni Franchi

机构 * U2IS Lab ENSTA Paris(ENSTA巴黎大学U2IS实验室) Safran Tech(萨弗兰技术)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025 (Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01258 2025-11-04 cs.CL cs.AI 62%

Measuring Algorithmic Partisanship via Zero-Shot Classification and Its Implications on Political Discourse

Nathan Junzi Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00002 2025-11-04 cs.LG cs.AI cs.CV 62%

VRScout: Towards Real-Time, Autonomous Testing of Virtual Reality Games

Yurun Wu, Yousong Sun, Burkhard Wunsche, Jia Wang, Elliott Wen

机构 * School of Computer Science University of Auckland(计算机科学学院 奥克兰大学) School of Advanced Technology Xi'an Jiaotong-Liverpool University(先进科技学院 西交利物浦大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27641 2025-11-03 cs.CL cs.LG cs.SY eess.SY 62%

SpecAttn: Speculating Sparse Attention

Harsh Shah

机构 * Machine Learning Department(机器学习系) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2025 Workshop on Structured Probabilistic Inference & Generative Modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23724 2025-11-03 cs.LG cs.AI 62%

SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA

Minrui Luo, Fuhang Kuang, Yu Wang, Zirui Liu, Tianxing He

机构 * Shanghai Qi Zhi Institute(上海启智研究院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Xiongan AI Institute(雄安人工智能研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25933 2025-10-31 cs.AI cs.HC cs.LG cs.NE 62%

Humains-Junior: A 3.8B Language Model Achieving GPT-4o-Level Factual Accuracy by Directed Exoskeleton Reasoning

Nissan Yaron, Dan Bystritsky, Ben-Etzion Yaron

机构 * Humains AI Research(Humains人工智能研究) Inpris Ltd(Inpris公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22149 2025-10-30 cs.AI cs.LG 62%

When Truthful Representations Flip Under Deceptive Instructions?

Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng, Haotian Yu, Xiaotian Han, Pan Li

机构 * Case Western Reserve University(凯斯西储大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24085 2025-10-29 cs.AI cs.LG 62%

Modeling Electric Vehicle Car-Following Behavior: Classical vs Machine Learning Approach

Md. Shihab Uddin, Md Nazmus Shakib, Rahul Bhadani

机构 * Electrical and Computer Engineering, The University of Alabama in Huntsville, Huntsville, AL, USA(电气与计算机工程系,阿拉巴马大学亨茨维尔分校)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15870 2025-10-29 cs.CV cs.AI cs.CL 62%

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, Yuming Lou, Dong Yang, Zhijian Liu, Yukang Chen, Ambrish Dantrey, Ehsan Jahangiri, Sreyan Ghosh, Daguang Xu, Ehsan Hosseini-Asl, Danial Mohseni Taheri, Vidya Murali, Sifei Liu, Yao Lu, Oluwatobi Olabiyi, Yu-Chiang Frank Wang, Rafael Valle, Bryan Catanzaro, Andrew Tao, Song Han, Jan Kautz, Hongxu Yin, Pavlo Molchanov

机构 * NVIDIA

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Technical Report. Code: https://github.com/NVlabs/OmniVinci

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21623 2025-10-27 cs.CL cs.AI 62%

The Universal Landscape of Human Reasoning

Qiguang Chen, Jinhao Liu, Libo Qin, Yimeng Zhang, Yihao Liang, Shangxu Ren, Chengyu Luan, Dengyun Peng, Hanjing Li, Jiannan Guan, Zheng Yan, Jiaqi Wang, Mengkang Hu, Yantao Du, Zhi Chen, Xie Chen, Wanxiang Che

机构 * Harbin Institute of Technology(哈尔滨工业大学) Central South University(中南大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Princeton University(普林斯顿大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) ByteDance Seed (China)(字节跳动种子(中国)) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21359 2025-10-27 cs.CL cs.AI 62%

Influence Guided Context Selection for Effective Retrieval-Augmented Generation

Jiale Deng, Yanyan Shen, Ziyuan Pei, Youmin Chen, Linpeng Huang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20377 2025-10-24 cs.AI cs.CL 62%

IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation

Tianyi Zhang, Florian Mai, Lucie Flek

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉玛尔机器学习与人工智能研究所)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16722 2025-10-24 cs.CL cs.AI 62%

Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification

Himanshu Beniwal, Youngwoo Kim, Maarten Sap, Soham Dan, Thomas Hartvigsen

机构 * Indian Institute of Technology Gandhinagar(印度古吉拉特邦理工学院加尔文加尔)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at MELT Workshop @ COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01472 2025-10-22 cs.CL cs.AI 62%

FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language Model

Jinwei Hu, Zhenglin Huang, Xiangyu Yin, Wenjie Ruan, Guangliang Cheng, Yi Dong, Xiaowei Huang

机构 * School of Computer Science and Informatics, University of Liverpool, UK(计算机科学与信息学学院,利物浦大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025 with minor revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03313 2025-10-21 cs.LG cs.CL 62%

LLM as GNN: Graph Vocabulary Learning for Text-Attributed Graph Foundation Models

Xi Zhu, Haochen Xue, Ziwei Zhao, Wujiang Xu, Jingyuan Huang, Minghao Guo, Qifan Wang, Kaixiong Zhou, Imran Razzak, Yongfeng Zhang

机构 * Rutgers University(罗格斯大学) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) University of Science and Technology of China(中国科学技术大学) Meta AI North Carolina State University(北卡罗来纳州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11704 2025-10-21 cs.CL cs.AI 62%

Adapting Chat Language Models Using Only Target Unlabeled Language Data

Atsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio, Nikolaos Aletras

机构 * University of Sheffield(谢菲尔德大学) Hitachi, Ltd.(日立株式会社) University of Exeter(埃克塞特大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10807 2025-10-21 cs.LG cs.AI stat.ML 62%

HardNet: Hard-Constrained Neural Networks with Universal Approximation Guarantees

Youngjae Min, Navid Azizan

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16068 2025-10-21 cs.CY cs.AI 62%

Co-Designing Interdisciplinary Design Projects with AI

Wei Ting Liow, Sumbul Khan, Lay Kee Ang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

Comments to be published in IEEE TALE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15996 2025-10-21 cs.LG cs.AI 62%

Using Kolmogorov-Smirnov Distance for Measuring Distribution Shift in Machine Learning

Ozan K. Tonguz, Federico Taschin

机构 * Carnegie Mellon University, USA(卡内基梅隆大学,美国) KTH Royal Institute of Technology, Sweden(皇家理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14620 2025-10-17 cs.CL cs.AI 62%

Code-driven Number Sequence Calculation: Enhancing the inductive Reasoning Abilities of Large Language Models

Kedi Chen, Zhikai Lei, Xu Guo, Xuecheng Wu, Siyuan Zeng, Jianghao Yin, Yinqi Zhang, Qin Chen, Jie Zhou, Liang He, Qipeng Guo, Kai Chen, Wei Zhang

机构 * East China Normal University(东华大学) Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Xi’an Jiaotong University(西安交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14304 2025-10-16 cs.CL cs.LG 62%

Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study

Rakesh Paul, Anusha Kamath, Kanishk Singla, Raviraj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan, Niranjan Wartikar

机构 * NVIDIA

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13002 2025-10-16 cs.AI cs.LG 62%

From Narratives to Probabilistic Reasoning: Predicting and Interpreting Drivers' Hazardous Actions in Crashes Using Large Language Model

Boyou Chen, Gerui Xu, Zifei Wang, Huizhong Guo, Ananna Ahmed, Zhaonan Sun, Zhen Hu, Kaihan Zhang, Shan Bao

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12245 2025-10-15 cs.LG cs.AI 62%

MoRA: On-the-fly Molecule-aware Low-Rank Adaptation Framework for LLM-based Multi-Modal Molecular Assistant

Tao Yin, Xiaohong Zhang, Jiacheng Zhang, Li Huang, Zhibin Zhang, Yuansong Zeng, Jin Xie, Meng Yan

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12063 2025-10-15 cs.AI cs.CL 62%

ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization

Sunzhu Li, Zhiyu Lin, Shuling Yang, Jiale Zhao, Wei Chen

机构 * Li Auto Inc.(Li Auto公司) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11953 2025-10-15 cs.LG cs.AI 62%

Sculpting Latent Spaces With MMD: Disentanglement With Programmable Priors

Quentin Fruytier, Akshay Malhotra, Shahab Hamidi-Rad, Aditya Sant, Aryan Mokhtari, Sujay Sanghavi

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) InterDigital Communications, AI Lab(InterDigital通讯人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03597 2025-10-15 cs.GR cs.AI cs.LG 62%

Neon: Negative Extrapolation From Self-Training Improves Image Generation

Sina Alemohammad, Zhangyang Wang, Richard G. Baraniuk

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18966 2025-10-15 cs.LG cs.AI q-bio.BM 62%

Protein Design with Dynamic Protein Vocabulary

Nuowei Liu, Jiahao Kuang, Yanting Liu, Tao Ji, Changzhi Sun, Man Lan, Yuanbin Wu

机构 * School of Computer Science and Technology, East China Normal University(东华师范大学计算机科学与技术学院) College of Foreign Languages and Literatures, Fudan University(复旦大学外国语言文学学院) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏