arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2510.11693 2025-10-14 cs.CL cs.AI cs.CV 62%

Scaling Language-Centric Omnimodal Representation Learning

Chenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu, Mahani Aljunied, Yu Rong

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23703 2025-10-14 cs.AI cs.CL 62%

Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability

Ruida Wang, Yuxin Li, Yi R. Fung, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10398 2025-10-14 cs.CL cs.AI 62%

STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models

Geunyeong Jeong, Juoh Sun, Seonghee Lee, Harksoo Kim

机构 * Konkuk University(韩国康康大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18361 2025-10-14 q-bio.NC cs.AI cs.LG cs.RO 62%

Task-Optimized Convolutional Recurrent Networks Align with Tactile Processing in the Rodent Brain

Trinity Chung, Yuchen Shen, Nathan C. L. Kong, Aran Nayebi

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Machine Learning Department, Carnegie Mellon University(卡内基梅隆大学机器学习系) Department of Psychology, University of Pennsylvania(宾夕法尼亚大学心理学系) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures, 7 tables, NeurIPS 2025 Camera Ready Version (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15607 2025-10-14 cs.CL cs.AI 62%

From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning

David Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi, Iryna Gurevych, Mrinmaya Sachan

机构 * Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) ETH AI Center(ETH人工智能中心) Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany(达姆施塔特技术大学计算机科学系、应用网络安全国家研究中心ATHENE)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Main as an oral presentation. David Dinucu-Jianu and Jakub Macina contributed equally. Code available: https://github.com/eth-lre/PedagogicalRL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09762 2025-10-14 cs.LG cs.AI 62%

PatentVision: A multimodal method for drafting patent applications

Ruo Yang, Sai Krishna Reddy Mudhiganti, Manali Sharma

机构 * Samsung Semiconductor, Inc.(三星半导体公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12963 2025-10-14 cs.AI cs.LG 62%

Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills

Changsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia, Dennis Wei, Parikshit Ram, Nathalie Baracaldo, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19528 2025-10-13 cs.CV cs.AI cs.GR cs.LG 62%

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

Xianfeng Tan, Yuhan Li, Wenxiang Shang, Yubo Wu, Jian Wang, Xuanhong Chen, Yi Zhang, Ran Lin, Bingbing Ni

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Group(阿里巴巴集团)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accept by ICCV 2025 (Highlight). Project website: https://colorful-liyu.github.io/RAGDiffusion-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07513 2025-10-10 cs.LG cs.AI cs.CV cs.DB 62%

MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis

Qinghua Liu, Sam Heshmati, Zheda Mai, Zubin Abraham, John Paparrizos, Liu Ren

机构 * The Ohio State University(俄亥俄州立大学) Bosch Research North America(博世北美研究)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23206 2025-10-10 cs.CL cs.AI 62%

PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness

Huacan Chai, Zijie Cao, Maolin Ran, Yingxuan Yang, Jianghao Lin, Xin Peng, Hairui Wang, Renjie Ding, Ziyu Wan, Muning Wen, Weiwen Liu, Weinan Zhang, Fei Huang, Ying Wen

机构 * Shanghai Jiao Tong University(上海交通大学) LongShine AI Research(LongShine人工智能研究院) Shanghai Innovation Institute(上海创新研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07091 2025-10-09 cs.AI cs.CL 62%

The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas

Baixuan Xu, Tianshi Zheng, Zhaowei Wang, Hong Ting Tsang, Weiqi Wang, Tianqing Fang, Yangqiu Song

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17967 2025-10-09 cs.LG cs.AI 62%

FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models

Ionut-Vlad Modoranu, Mher Safaryan, Erik Schultheis, Max Ryabinin, Artem Chumachenko, Dan Alistarh

机构 * ISTA(因斯托克科学与技术研究院) Together AI

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05197 2025-10-08 cs.AI cs.LG stat.AP stat.ML 62%

Efficient Prediction of Pass@k Scaling in Large Language Models

Joshua Kazdan, Rylan Schaeffer, Youssef Allouah, Colin Sullivan, Kyssen Yu, Noam Levi, Sanmi Koyejo

机构 * Stanford University(斯坦福大学) University of Toronto(多伦多大学) EPFL(苏黎世联邦理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03845 2025-10-07 cs.AI cs.GT cs.LG stat.ML 62%

The Hidden Game Problem

Gon Buzaglo, Noah Golowich, Elad Hazan

机构 * Princeton University(普林斯顿大学) Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02425 2025-10-06 cs.CL cs.CV cs.LG 62%

Words That Make Language Models Perceive

Sophie L. Wang, Phillip Isola, Brian Cheung

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18216 2025-10-03 cs.AI cs.LG 62%

nDNA -- the Semantic Helix of Artificial Cognition

Amitava Das

机构 * BITS Pilani, Goa, India(印度戈阿学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00631 2025-10-03 cs.LG cs.AI physics.ao-ph 62%

Forecasting the Ionosphere from Sparse GNSS Data with Temporal-Fusion Transformers

Giacomo Acciarini, Simone Mestici, Halil Kelebek, Linnea Wolniewicz, Michael Vergalla, Madhulika Guhathakurta, Umaa Rebbapragada, Bala Poduval, Atılım Güneş Baydin, Frank Soboczenski

机构 * Advanced Concepts Team European Space Agency(欧洲航天局高级概念团队) Department of Physics Università degli Studi di Roma Sapienza(罗马大学物理系) Department of Engineering Science University of Oxford(牛津大学工程科学系) Department of Information and Computer Science University of Hawai’i at Mānoa(夏威夷大学信息与计算机科学系) Free Flight Research Lab(自由飞行研究实验室) NASA Headquarters(美国国家航空航天局总部) NASA Jet Propulsion Laboratory(美国国家航空航天局喷气推进实验室) University of New Hampshire(新罕布什尔大学) Department of Computer Science University of Oxford, UK(牛津大学计算机科学系) Department of Computer Science University of York & King’s College London(约克大学计算机科学系及伦敦国王学院计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01004 2025-10-02 cs.CV cs.AI cs.LG 62%

TextCAM: Explaining Class Activation Map with Text

Qiming Zhao, Xingjian Li, Xiaoyu Cao, Xiaolong Wu, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05978 2025-10-02 eess.IV cs.CL cs.CV cs.LG 62%

Imagining Alternatives: Towards High-Resolution 3D Counterfactual Medical Image Generation via Language Guidance

Mohamed Mohamed, Brennan Nichyporuk, Douglas L. Arnold, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to the 2025 MICCAI ELAMI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08521 2025-10-02 cs.LG cs.AI 62%

Mitigating Domain Shift in Federated Learning via Intra- and Inter-Domain Prototypes

Huy Q. Le, Ye Lin Tun, Yu Qiao, Minh N. H. Nguyen, Keon Oh Kim, Eui-Nam Huh, Choong Seon Hong

机构 * Department of Computer Science and Engineering, School of Computing, Kyung Hee University(计算机科学与工程系,计算学院,庆熙大学) Department of Artificial Intelligence, School of Computing, Kyung Hee University(人工智能系,计算学院,庆熙大学) Digital Science and Technology Institute, The University of Danang—Vietnam-Korea University of Information and Communication Technology(数字科学与技术研究所,丹那大学—越南-韩国信息与通信技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26625 2025-10-01 cs.LG cs.AI cs.CV cs.MM 62%

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos

机构 * Meta Superintelligence Labs(Meta 超智能实验室) University of Oxford(牛津大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Project page: https://junlinhan.github.io/projects/lsbs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26239 2025-10-01 cs.LG cs.AI stat.ML 62%

Sandbagging in a Simple Survival Bandit Problem

Joel Dyer, Daniel Jarne Ornia, Nicholas Bishop, Anisoara Calinescu, Michael Wooldridge

机构 * University of Oxford(牛津大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Forthcoming in the "Reliable ML from Unreliable Data Workshop" at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25220 2025-10-01 cs.CL cs.LG 62%

Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI

Eduard Kapelko

机构 * Eduard Kapelko(独立研究者)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments Code is available at: https://www.kaggle.com/code/kapedalex/cycleablationpublic/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24216 2025-09-30 cs.CL cs.CY 62%

MoVa: Towards Generalizable Classification of Human Morals and Values

Ziyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen, Jing Yao, Xiaoyuan Yi, Xing Xie, Chenhao Tan, Lexing Xie

机构 * The Australian National University(澳大利亚国立大学) University of Chicago(芝加哥大学) University of Pennsylvania(宾夕法尼亚大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

Comments 9 pages, 10 figures and tables, EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11425 2025-09-30 cs.SD cs.AI cs.CL eess.AS 62%

FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs

Md Mubtasim Ahasan, Rafat Hasan Khan, Tasnim Mohiuddin, Aman Chadha, Tariq Iqbal, M Ashraful Amin, Amin Ahsan Ali, Md Mofijul Islam, A K M Mahbubur Rahman

机构 * Center for Computational & Data Sciences, Independent University, Bangladesh(计算与数据科学中心,独立大学,孟加拉国) Amazon GenAI(亚马逊生成人工智能) Qatar Computing Research Institute(卡塔尔计算研究所) University of Virginia(弗吉尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21625 2025-09-29 cs.SD cs.AI cs.LG eess.AS 62%

Guiding Audio Editing with Audio Language Model

Zitong Lan, Yiduo Hao, Mingmin Zhao

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20208 2025-09-25 cs.CL cs.AI cs.DB 62%

Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs

Parker Glenn, Alfy Samuel, Daben Liu

机构 * Capital One

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20051 2025-09-25 cs.LG cs.AI 62%

One Filters All: A Generalist Filter for State Estimation

Shiqi Liu, Wenhan Cao, Chang Liu, Zeyu He, Tianyi Zhang, Shengbo Eben Li

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院) College of Engineering, Peking University(北京大学工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19789 2025-09-25 cs.LG cs.AI cs.RO 62%

RDAR: Reward-Driven Agent Relevance Estimation for Autonomous Driving

Carlo Bosio, Greg Woelki, Noureldin Hendy, Nicholas Roy, Byungsoo Kim

机构 * UC Berkeley(伯克利大学) Zoox Inc.(Zoox公司)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18439 2025-09-24 cs.CL cs.AI 62%

Developing an AI framework to automatically detect shared decision-making in patient-doctor conversations

Oscar J. Ponce-Ponte, David Toro-Tobon, Luis F. Figueroa, Michael Gionfriddo, Megan Branda, Victor M. Montori, Saturnino Luz, Juan P. Brito

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 53 pages, 1 figure, 4 tables, 5 supplementary figures, 13 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏