arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2509.18113 2025-09-24 cs.CL cs.LG 62%

Dynamic Prompt Fusion for Multi-Task and Cross-Domain Adaptation in LLMs

Xin Hu, Yue Kang, Guanzi Yao, Tianze Kang, Mengjie Wang, Heyao Liu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12734 2025-09-24 cs.CL cs.AI 62%

Pandora: A Code-Driven Large Language Model Agent for Unified Reasoning Across Diverse Structured Knowledge

Yongrui Chen, Junhao He, Linbo Fu, Shenyu Zhang, Rihui Jin, Xinbang Dai, Jiaqi Li, Dehai Min, Nan Hu, Yuxin Zhang, Guilin Qi, Yi Huang, Tongtong Wu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments New version is arXiv:2508.17905

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04183 2025-09-24 cs.CL cs.AI 62%

GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding

Ziyin Zhang, Hang Yu, Shijie Li, Peng Di, Jianguo Li, Rui Wang

机构 * Ant Group(蚂蚁集团) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00742 2025-09-23 cs.CL cs.LG 62%

Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents

Sarah Mercer, Daniel P. Martin, Phil Swatton

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15789 2025-09-22 cs.CL cs.LG 62%

UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations

Qiuyang Lu, Fangjian Shen, Zhengkai Tang, Qiang Liu, Hexuan Cheng, Hui Liu, Wushao Wen

机构 * United Nations(联合国)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments 5 pages, 1 figure, submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15975 2025-09-22 cs.CL cs.AI 62%

Sparsity May Be All You Need: Sparse Random Parameter Adaptation

Jesus Rios, Pierre Dognin, Ronny Luss, Karthikeyan N. Ramamurthy

机构 * IBM Research(IBM研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18848 2025-09-22 cs.LG cs.AI 62%

Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection

Alain Ryser, Thomas M. Sutter, Alexander Marx, Julia E. Vogt

机构 * Department of Computer Science ETH Zurich(计算机科学系,苏黎世联邦理工学院) Research Center Trustworthy Data Science and Security of the University Alliance Ruhr(可信数据科学与安全的鲁尔大学联盟研究中心) Department of Statistics TU Dortmund University(统计学系,多特蒙德技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research (TMLR) https://openreview.net/forum?id=Bt0zdsnWYc

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15474 2025-09-22 cs.CL cs.AI 62%

Subjective Behaviors and Preferences in LLM: Language of Browsing

Sai Sundaresan, Harshita Chopra, Atanu R. Sinha, Koustava Goswami, Nagasai Saketh Naidu, Raghav Karan, N Anushka

机构 * Adobe Research(Adobe研究院) University of Washington, Seattle(华盛顿大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13282 2025-09-17 cs.CL cs.CV cs.LG 62%

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan, Mir Rayat Imtiaz Hossain, Leonid Sigal, Giuseppe Carenini

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19331 2025-09-17 cs.CV cs.AI cs.CL 62%

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

Luca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara, Marcella Cornia, Lorenzo Baraldi, Fabrizio Falchi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ISTI-CNR(意大利国家研究委员会ISTI) University of Pisa(比萨大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10529 2025-09-16 cs.LG cs.AI cs.CV 62%

Mitigating Catastrophic Forgetting and Mode Collapse in Text-to-Image Diffusion via Latent Replay

Aoi Otani

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10523 2025-09-16 cs.LG cs.AI 62%

From Predictions to Explanations: Explainable AI for Autism Diagnosis and Identification of Critical Brain Regions

Kush Gupta, Amir Aly, Emmanuel Ifeachor, Rohit Shankar

机构 * University of Plymouth, Plymouth, UK(普利茅斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07373 2025-09-15 cs.LG cs.AI 62%

Atherosclerosis through Hierarchical Explainable Neural Network Analysis

Irsyad Adam, Steven Swee, Erika Yilin, Ethan Ji, William Speier, Dean Wang, Alex Bui, Wei Wang, Karol Watson, Peipei Ping

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12039 2025-09-15 cs.CR cs.AI cs.CL cs.SE 62%

Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection

Ira Ceka, Feitong Qiao, Anik Dey, Aastha Valecha, Gail Kaiser, Baishakhi Ray

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11792 2025-09-11 cs.RO cs.AI cs.LG 62%

Efficient and Generalized end-to-end Autonomous Driving System with Latent Deep Reinforcement Learning and Demonstrations

Zuojin Tang, Xiaoyu Chen, Yongqiang Li, Jianyu Chen

机构 * Shanghai Qizhi Institute(上海启智研究所) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息学院) Mogo Auto Intelligence and Telematics Information Technology Co., Ltd(摩戈智能与 telemetry 信息技术有限公司)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by ECML PKDD 2025 (Research Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07311 2025-09-10 cs.CL cs.AI 62%

Does This Look Familiar to You? Knowledge Analysis via Model Internal Representations

Sihyun Park

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02920 2025-09-10 cs.LG cs.CY cs.ET cs.SY eess.SY 62%

Event Detection and Classification for Long Range Sensing of Elephants Using Seismic Signal

Jaliya L. Wijayaraja, Janaka L. Wijekoon, Malitha Wijesundara

机构 * Sri Lanka Institute of Information Technology(斯里兰卡信息技术研究所) Victorian Institute of Technology(维多利亚技术学院) Department of System Design Engineering, Keio University(系统设计工程系,庆应大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CY、cs.LG

Comments This article has been accepted for publication in IEEE Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06918 2025-09-09 cs.LG cs.AI 62%

Tackling the Noisy Elephant in the Room: Label Noise-robust Out-of-Distribution Detection via Loss Correction and Low-rank Decomposition

Tarhib Al Azad, Shahana Ibrahim

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) University of Central Florida(中央佛罗里达大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04500 2025-09-08 cs.CL cs.AI 62%

Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts

Rushi Wang, Jiateng Liu, Cheng Qian, Yifan Shen, Yanzhou Pan, Zhaozhuo Xu, Ahmed Abbasi, Heng Ji, Denghui Zhang

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 36 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03805 2025-09-05 cs.CL cs.AI 62%

Measuring How (Not Just Whether) VLMs Build Common Ground

Saki Imai, Mert İnan, Anthony Sicilia, Malihe Alikhani

机构 * Northeastern University(东北大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03137 2025-09-04 cs.LG cs.AI nucl-ex physics.comp-ph physics.ins-det 62%

A Neural Network Approach to Multi-radionuclide TDCR Beta Spectroscopy

Li Yi, Qian Yang

机构 * Institute of Frontier and Interdisciplinary Science, Shandong University(前沿与交叉科学研究院,山东大学) Key Laboratory of Particle Physics and Particle Irradiation, Ministry of Education, Shandong University(粒子物理与粒子辐照重点实验室,教育部,山东大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03118 2025-09-04 cs.LG cs.AI cs.MA 62%

A Hierarchical Deep Reinforcement Learning Framework for Traffic Signal Control with Predictable Cycle Planning

Hankang Gu, Yuli Zhang, Chengming Wang, Ruiyuan Jiang, Ziheng Qiao, Pengfei Fan, Dongyao Jia

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19475 2025-09-04 cs.CV astro-ph.GA cs.AI cs.LG 62%

GalaxAlign: Mimicking Citizen Scientists' Multimodal Guidance for Galaxy Morphology Analysis

Ruoqi Wang, Haitao Wang, Qiong Luo

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) School of Computer Science and Engineering, Sun Yat-Sen University(中山大学计算机科学与工程学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01885 2025-09-03 cs.CL cs.AI 62%

Extracting OPQRST in Electronic Health Records using Large Language Models with Reasoning

Zhimeng Luo, Abhibha Gupta, Adam Frisch, Daqing He

机构 * School of Computing and Information University of Pittsburgh(计算与信息学院 西弗吉尼亚大学) Department of Emergency Medicine University of Pittsburgh(急诊医学系 西弗吉尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01592 2025-09-03 cs.CR cs.AI cs.LG cs.SY eess.SY 62%

Securing Radiation Detection Systems with an Efficient TinyML-Based IDS for Edge Devices

Einstein Rivas Pizarro, Wajiha Zaheer, Li Yang, Khalil El-Khatib, Glenn Harvel

机构 * Ontario Tech University(安大略技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Preprint author original pre review. Accepted and Presented at NPIC & HMIT 2025. The official proceedings version is available in the ANS Digital Library

Journal ref Proc. NPIC HMIT 2025 pp. 651-661 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00351 2025-09-03 cs.CV cs.AI cs.LG 62%

Target-Oriented Single Domain Generalization

Marzi Heidari, Yuhong Guo

机构 * School of Computer Science, Carleton University(计算机科学学院,卡尔顿大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16834 2025-09-03 cs.LG cs.AI physics.ao-ph 62%

Improving Significant Wave Height Prediction Using Chronos Models

Yilin Zhai, Hongyuan Shi, Chao Zhan, Qing Wang, Zaijin You, Nan Wang

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2403.07815 by other authors

Journal ref Ocean Engineering, Volume 341, Part 2, 1 December 2025, Article 122502

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20217 2025-08-29 cs.CL cs.AI 62%

Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models

Mohammad Amini, Babak Ahmadi, Xiaomeng Xiong, Yilin Zhang, Christopher Qiao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20018 2025-08-28 cs.AI cs.CL cs.CV cs.MA 62%

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control

Quanfeng Lu, Zhantao Ma, Shuai Zhong, Jin Wang, Dahai Yu, Michael K. Ng, Ping Luo

机构 * The University of Hong Kong(香港大学) Hong Kong Baptist University(香港 Baptist 大学) TCL Corporate Research (Hong Kong) Co., Ltd(TCL 香港研究院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 28 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19464 2025-08-28 cs.CL cs.AI 62%

Bridging Language Gaps: Enhancing Few-Shot Language Adaptation

Philipp Borchert, Jochen De Weerdt, Marie-Francine Moens

机构 * IESEG School of Management(IESEG管理学院) Research Centre for Information Systems Engineering(信息系统工程研究中心) Department of Computer Science(计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏