arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2510.23630 2025-10-29 cs.LG cs.AI cs.CL 67%

NUM2EVENT: Interpretable Event Reasoning from Numerical time-series

Ninghui Feng, Yiyan Qi

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA)) University of Nottingham Ningbo(诺丁汉大学宁波校区)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21184 2025-10-27 cs.LG cs.AI cs.CL stat.ML 67%

Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference

Stephen Zhao, Aidan Li, Rob Brekelmans, Roger Grosse

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19779 2025-10-23 cs.CL cs.AI cs.LG 67%

AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders

Yuezhou Hu, Jiaxin Guo, Xinyu Feng, Tuo Zhao

机构 * University of California, Berkeley(加州大学伯克利分校) Tsinghua University(清华大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18502 2025-10-22 cs.CV cs.AI cs.CL cs.LG 67%

Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation

Wei-Chia Chang, Yan-Ann Chen

机构 * Yuan Ze University(元智大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by The 38th Conference of Open Innovations Association FRUCT, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17692 2025-10-17 cs.CL cs.AI cs.LG 67%

MIO: A Foundation Model on Multimodal Tokens

Zekun Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jiashuo Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang

机构 * Beihang University(北航) M-A-P The Hong Kong Polytechnic University(香港理工大学) AIWaves University of Alberta(阿尔伯塔大学) University of Waterloo(滑铁卢大学) University of Manchester(曼彻斯特大学) Chinese Academy of Sciences(中国科学院) Peking University(北京大学) Shanghai AI Lab(上海AI实验室) Nanjing University(南京大学) Kuaishou Technology(快手科技)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2025 (Oral). Codes and models are available in https://github.com/MIO-Team/MIO

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13462 2025-10-16 cs.CR 67%

Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers

Xin Zhao, Xiaojun Chen, Bingshan Liu, Haoyu Gao, Zhendong Zhao, Yilong Chen

专题命中 其他安全 :safety(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08855 2025-10-13 cs.LG cs.AI cs.CL 67%

Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training

T. Ed Li, Junyu Ren

机构 * Yale University(耶鲁大学) University of Chicago(芝加哥大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments First submitted on February 10th, 2025 to ICLR 2025 Workshop (XAI4Science: From Understanding Model Behavior to Discovering New Scientific Knowledge). The paper was accepted but the workshop does not generate proceedings. Now uploading to arXiv to make the paper publicly available

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19102 2025-10-10 cs.IR cs.AI cs.CL cs.LG 67%

Distilling a Small Utility-Based Passage Selector to Enhance Retrieval-Augmented Generation

Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Baidu Inc(百度公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by SIGIR-AP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05763 2025-10-09 cs.CL cs.AI cs.LG 67%

GMLM: Bridging Graph Neural Networks and Language Models for Heterophilic Node Classification

Aarush Sinha

机构 * Department of Computer Science ,University of Copenhagen(计算机科学系,哥本哈根大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03727 2025-10-07 cs.AI cs.CL cs.CV cs.LG 67%

Bridging the Gap Between Multimodal Foundation Models and World Models

Xuehai He

机构 * Computer Science and Engineering University of California, Santa Cruz(计算机科学与工程大学加州大学圣克ruz分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21923 2025-10-06 cs.LG cs.AI cs.CL 67%

Multiplicative-Additive Constrained Models:Toward Joint Visualization of Interactive and Independent Effects

Fumin Wang

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15182 2025-09-30 cs.CL cs.AI cs.LG 67%

ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection

Jeonghye Kim, Sojeong Rhee, Minbeom Kim, Dohyung Kim, Sangmook Lee, Youngchul Sung, Kyomin Jung

机构 * KAIST(韩国科学技术院) Seoul National University(首尔国立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15243 2025-09-22 cs.CV 67%

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Muhammad Imran, Yugyung Lee

机构 * Computer Science, School of Science and Engineering, University of Missouri - Kansas City(计算机科学系,科学与工程学院,密苏里大学-堪萨斯城分校)

专题命中 其他安全 :alignment(abstract);safety(abstract)

Comments 8 pages, 6 figures, 3 tables

Journal ref Non-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09740 2025-09-15 q-bio.QM cs.AI cs.CL cs.LG 67%

HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets

Ying Yuan, Xing-Yue Monica Ge, Aaron Archer Waterman, Tommaso Biancalani, David Richmond, Yogesh Pandit, Avtar Singh, Russell Littman, Jin Liu, Jan-Christian Huetter, Vladimir Ermakov

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06620 2025-09-09 cs.LG cs.AI cs.CL 67%

MedualTime: A Dual-Adapter Language Model for Medical Time Series-Text Multimodal Learning

Jiexia Ye, Weiqi Zhang, Ziyue Li, Jia Li, Meng Zhao, Fugee Tsung

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Technical University of Munich(慕尼黑技术大学) Columbia University(哥伦比亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 6 figure, 3 tables

Journal ref IJCAI 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12226 2025-09-09 cs.CL cs.AI cs.CV cs.LG 67%

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu

机构 * Fudan University(复旦大学) Multimodal Art Projection Research Community(多模态艺术投影研究社区) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 28 pages, 16 figures, under review, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01186 2025-09-03 cs.CL cs.AI cs.CY 67%

Statutory Construction and Interpretation for Artificial Intelligence

Luxi He, Nimra Nadeem, Michel Liao, Howard Chen, Danqi Chen, Mariano-Florentino Cuéllar, Peter Henderson

机构 * Princeton University(普林斯顿大学) Carnegie Endowment for International Peace(国际和平研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21377 2025-09-01 cs.CL cs.AI cs.LG 67%

Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models

Shubham Sharma, Sneha Tuli, Narendra Badam

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16059 2025-08-25 cs.AI cs.CL cs.LG 67%

Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting

Zhuomin Chen, Dan Li, Jiahui Zhou, Shunyu Wu, Haozheng Ye, Jian Lou, See-Kiong Ng

机构 * School of Software Engineering Sun Yat-Sen University(软件工程学院中山大学) Institute of Data Science(数据科学研究院) School of Computing(计算机学院) National University of Singapore(新加坡国立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments To be published in CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12422 2025-08-19 cs.CV 67%

Illusions in Humans and AI: How Visual Perception Aligns and Diverges

Jianyi Yang, Junyi Ye, Ankan Dash, Guiling Wang

专题命中 其他安全 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10567 2025-08-15 cs.CV cs.RO 67%

SpaRC-AD: A Baseline for Radar-Camera Fusion in End-to-End Autonomous Driving

Philipp Wolters, Johannes Gilg, Torben Teepe, Gerhard Rigoll

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 其他安全 :alignment(abstract);safety(abstract)

Comments 8 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21298 2025-08-12 cs.SD cs.AI cs.CL cs.LG cs.MM eess.AS 67%

Exploring Adapter Design Tradeoffs for Low Resource Music Generation

Atharva Mehta, Shivam Chauhan, Monojit Choudhury

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14701 2025-08-12 cs.CL cs.AI cs.HC cs.LG q-bio.NC 67%

COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling

Baihan Lin, Djallel Bouneffouf, Yulia Landa, Rachel Jespersen, Cheryl Corcoran, Guillermo Cecchi

机构 * Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai(人工智能与人类健康系,伊坎医学院 Mount Sinai 分校) Department of Psychiatry, Icahn School of Medicine at Mount Sinai(精神病学系,伊坎医学院 Mount Sinai 分校) Department of Neuroscience, Icahn School of Medicine at Mount Sinai(神经科学系,伊坎医学院 Mount Sinai 分校) Berkman Klein Center for Internet & Society, Harvard University(互联网与社会研究中心,哈佛大学) IBM Research, T.J. Watson Research Center(IBM 研究,T.J. Watson 研究中心) Mental Illness Research, Education and Clinical Center, James J. Peters VA Medical Center(精神疾病研究、教育与临床中心,James J. Peters VA 医疗中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Translational Psychiatry, in press. This work extends our research series in computational psychiatry (e.g auto annotation in arXiv:2204.05522, topic extraction in arXiv:2204.10189, and diagnosis in arXiv:2210.15603) with the introduction of LLMs to complete the full cycle of interpreting and understanding psychotherapy strategies as a comprehensive analytical framework

Journal ref Transl Psychiatry 15, 166 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02215 2025-08-05 cs.LG cs.AI cs.CL 67%

LeanK: Learnable K Cache Channel Pruning for Efficient Decoding

Yike Zhang, Zhiyuan He, Huiqiang Jiang, Chengruidong Zhang, Yuqing Yang, Jianyong Wang, Lili Qiu

机构 * Tsinghua University(清华大学) Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20930 2025-07-31 cs.CL cs.AI cs.LG 67%

FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models

Likun Tan, Kuan-Wei Huang, Kevin Wu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21138 2025-07-30 cs.CL cs.AI cs.LG cs.SD eess.AS 67%

TTS-1 Technical Report

Oleg Atamanenko, Anna Chalova, Joseph Coombes, Nikki Cope, Phillip Dang, Zhifeng Deng, Jimmy Du, Michael Ermolenko, Feifan Fan, Yufei Feng, Cheryl Fichter, Pavel Filimonov, Louis Fischer, Kylan Gibbs, Valeria Gusarova, Pavel Karpik, Andreas Assad Kottner, Ian Lee, Oliver Louie, Jasmine Mai, Mikhail Mamontov, Suri Mao, Nurullah Morshed, Igor Poletaev, Florin Radu, Dmytro Semernia, Evgenii Shingarev, Vikram Sivaraja, Peter Skirko, Rinat Takhautdinov, Robert Villahermosa, Jean Wang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 20 pages, 10 figures. For associated modeling and training code, see https://github.com/inworld-ai/tts

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19568 2025-07-29 cs.CY cs.AI cs.CE cs.LG 67%

Programmable Virtual Humans Toward Human Physiologically-Based Drug Discovery

You Wu, Philip E. Bourne, Lei Xie

机构 * Ph.D. Program in Computer Science The Graduate Center The City University of New York(计算机科学博士项目 美国城市大学研究生中心 新 York 美国) School of Data Science & Department of Biomedical Engineering University of Virginia(数据科学学院与生物医学工程系 美国弗吉尼亚大学) School of Pharmacy and Pharmaceutical Sciences & Center for Drug Discovery Northeastern University(药学与药学科学学院及药物发现中心 美国东北大学) Department of Computer Science Hunter College The City University of New York(计算机科学系 美国城市大学霍顿学院) Helen & Robert Appel Alzheimer’s Disease Research Institute Feil Family Brain & Mind Research Institute Weill Cornell Medicine Cornell University(海伦与罗伯特·阿普尔阿尔茨海默病研究所 费尔家族脑与心灵研究研究所 威尔·康奈尔医学院 康奈尔大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15743 2025-07-22 cs.AI cs.CL cs.HC cs.LG 67%

Towards physician-centered oversight of conversational diagnostic AI

Elahe Vedadi, David Barrett, Natalie Harris, Ellery Wulczyn, Shashir Reddy, Roma Ruparel, Mike Schaekermann, Tim Strother, Ryutaro Tanno, Yash Sharma, Jihyeon Lee, Cían Hughes, Dylan Slack, Anil Palepu, Jan Freyberg, Khaled Saab, Valentin Liévin, Wei-Hung Weng, Tao Tu, Yun Liu, Nenad Tomasev, Kavita Kulkarni, S. Sara Mahdavi, Kelvin Guu, Joëlle Barral, Dale R. Webster, James Manyika, Avinatan Hassidim, Katherine Chou, Yossi Matias, Pushmeet Kohli, Adam Rodman, Vivek Natarajan, Alan Karthikesalingam, David Stutz

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究) Harvard Medical School, Beth Israel Deaconess Medical Center(哈佛医学院,贝塞斯达德acons医学中心)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14079 2025-07-21 cs.CL cs.AI cs.IR cs.LG 67%

DENSE: Longitudinal Progress Note Generation with Temporal Modeling of Heterogeneous Clinical Notes Across Hospital Visits

Garapati Keerthana, Manik Gupta

机构 * Department of Computer Science and Information Systems, Birla Institute of Technology and Science, Pilani, Hyderabad Campus(计算机科学与信息系统系,比拉理工学院,比拉理工学院,海得拉巴校区)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12774 2025-07-18 cs.LG cs.AI cs.CL 67%

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models

Weijieying Ren, Jingxi Zhu, Zehao Liu, Tianxiang Zhao, Vasant Honavar

机构 * Information Sciences and Technology, The Pennsylvania State University(信息科学与技术系,宾夕法尼亚州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏