arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2507.07689 2025-07-11 cs.SE 67%

From Domain Documents to Requirements: Retrieval-Augmented Generation in the Space Industry

Chetan Arora, Fanyu Wang, Chakkrit Tantithamthavorn, Aldeida Aleti, Shaun Kenyon

专题命中 其他安全 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00078 2025-07-02 cs.LG cs.AI cs.CL 67%

The language of time: a language model perspective on time-series foundation models

Yi Xie, Yun Xiong, Zejian Shi, Hao Niu, Zhengfu Liu

机构 * College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学) Shanghai Key Laboratory of Data Science(上海数据科学重点实验室) ZCTech School of Mathematics and Statistics, Beijing Institute of Technology(数学与统计学学院,北京理工大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00002 2025-07-02 cs.LG cs.AI cs.CL 67%

Hypertokens: Holographic Associative Memory in Tokenized LLMs

Christopher James Augeri

机构 * Sloop FX, Inc.(Sloop FX 公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments preprint as accepted to https://qnlp.ai/ - Quantum AI and NLP Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21620 2025-06-30 cs.CL cs.AI cs.CY cs.SI physics.soc-ph 67%

How Large Language Models play humans in online conversations: a simulated study of the 2016 US politics on Reddit

Daniele Cirulli, Giulio Cimini, Giovanni Palermo

机构 * Enrico Fermi Research Center(恩里科·费米研究中心) Physics Department and INFN, University of Rome Tor Vergata(罗马托尔维加塔大学物理系和INFN)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17930 2025-06-24 cs.AI cs.CL cs.LG cs.NE cs.RO 67%

Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective

Jianyu Wang, Zhiqiang Hu, Lidong Bing

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎扑实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML 2025, and Code will be released at: https://github.com/jianyu-cs/PromptQuine/

Journal ref Forty-second International Conference on Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14997 2025-06-19 cs.CY cs.CL cs.LG 67%

Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings

Harbin Hong, Sebastian Caldas, Liu Leqi

机构 * Princeton University(普林斯顿大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12927 2025-06-17 cs.AI cs.CL cs.LG 67%

Sectoral Coupling in Linguistic State Space

Sebastian Dumbrava

机构 * Sebastian Dumbrava(独立研究者)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 56 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13509 2025-06-17 cs.CL cs.AI cs.LG 67%

ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series

Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Wei Bi, Yida Xu, Guo Li, Xian Yang

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanxi University(山西大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海研究院) Tencent AI Lab(腾讯人工智能实验室) Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) The University of Manchester(曼彻斯特大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments This paper is accepted by ACL2025(Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09223 2025-06-17 cs.CL cs.AI cs.LG 67%

Foundations of Large Language Models

Tong Xiao, Jingbo Zhu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10155 2025-06-13 cs.CL cs.AI cs.LG 67%

Measuring Corporate Human Capital Disclosures: Lexicon, Data, Code, and Research Opportunities

Elizabeth Demers, Victor Xiaoqi Wang, Kean Wu

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 50 pages, 6 figures, 5 tables

Journal ref Journal of Information Systems 38 (2024) 163-186

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01067 2025-06-12 cs.AI cs.CL cs.CV cs.HC cs.LG 67%

Human-like object concept representations emerge naturally in multimodal large language models

Changde Du, Kaicheng Fu, Bincheng Wen, Yi Sun, Jie Peng, Wei Wei, Ying Gao, Shengpei Wang, Chuncheng Zhang, Jinpeng Li, Shuang Qiu, Le Chang, Huiguang He

机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启发智能技术重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Neuroscience, State Key Laboratory of Brain Cognition and Brain-Inspired Intelligence Technology(神经科学研究所) CAS Center for Excellence in Brain Science and Intelligence Technology(中国科学院脑科学与智能技术卓越创新中心) Chinese Academy of Sciences(中国科学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published on Nature Machine Intelligence

Journal ref Nature Machine Intelligence, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08147 2025-06-11 cs.CL cs.AI cs.LG 67%

Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models

Muhammad Usman, Muhammad Ahmad, M. Shahiki Tash, Irina Gelbukh, Rolando Quintero Tellez, Grigori Sidorov

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00653 2025-06-06 cs.LG cs.AI cs.CL 67%

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models

Femi Bello, Anubrata Das, Fanzhi Zeng, Fangcong Yin, Liu Leqi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13928 2025-06-04 cs.CV cs.AI cs.CL cs.LG 67%

Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images

Shengguang Wu, Fan-Yun Sun, Kaiyue Wen, Nick Haber

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to ACL 2025 Main. Project Website: https://s-vco.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01662 2025-06-03 cs.CY cs.AI cs.LG 67%

Explainable AI Systems Must Be Contestable: Here's How to Make It Happen

Catarina Moreira, Anna Palatkina, Dacia Braca, Dylan M. Walsh, Peter J. Leihn, Fang Chen, Nina C. Hubig

机构 * Data Science Institute UTS Australia(UTS澳大利亚数据科学研究院) IT:U Austria(奥地利IT:U学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04794 2025-06-03 cs.CL cs.AI cs.LG 67%

KnowCoder-X: Boosting Multilingual Information Extraction via Code

Yuxin Zuo, Wenxuan Jiang, Wenxuan Liu, Zixuan Li, Long Bai, Hanbin Wang, Yutao Zeng, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng

机构 * Key Laboratory of Network Data Science and Technology, Institute of Computing Technology, Chinese Academy of Sciences(网络数据科学与技术重点实验室,计算技术研究所,中国科学院) State Key Laboratory of AI Safety(人工智能安全国家重点实验室) School of Computer Science, University of Chinese Academy of Sciences(中国科学院大学计算机科学学院) School of Software, Northeastern University(东北大学软件学院) School of Software, Peking University(北京大学软件学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18949 2025-05-27 cs.CL cs.AI cs.LG 67%

The Price of Format: Diversity Collapse in LLMs

Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, Jingbo Shang

机构 * University of California, San Diego(加州大学圣迭戈分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13515 2025-05-21 cs.LG cs.AI cs.CL 67%

LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades

Yanan Li, Fanxu Meng, Muhan Zhang, Shiai Zhu, Shangguang Wang, Mengwei Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Peking University(北京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01636 2025-05-06 cs.AI cs.CL cs.LG 67%

Structured Prompting and Feedback-Guided Reasoning with LLMs for Data Interpretation

Amit Rath

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 21 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02908 2025-05-01 cs.LG cs.AI cs.CL 67%

Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Sampling

Kaiwen Zheng, Yongxin Chen, Hanzi Mao, Ming-Yu Liu, Jun Zhu, Qinsheng Zhang

机构 * Department of Computer Science & Technology, Institute for AI, Tsinghua University(1 计算机科学与技术系,人工智能研究院,清华大学) NVIDIA(2 英特尔)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15929 2025-04-23 cs.AI cs.CL cs.LG 67%

Certifying Knowledge Comprehension in LLMs

Isha Chaudhary, Vedaant V. Jain, Gagandeep Singh

机构 * Siebel School of Computing and Data Science University of Illinois, Urbana-Champaign(计算机与数据科学学院 伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03026 2025-04-16 cs.RO cs.AI cs.CL cs.CV cs.LG 67%

LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Hao Sha, Yao Mu, Yuxuan Jiang, Li Chen, Chenfeng Xu, Ping Luo, Shengbo Eben Li, Masayoshi Tomizuka, Wei Zhan, Mingyu Ding

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08192 2025-04-14 cs.LG cs.AI cs.CL cs.CR 67%

SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07097 2025-04-10 cs.LG cs.AI cs.CL math.PR stat.ML 67%

Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning

Nikhil Shivakumar Nayak, Krishnateja Killamsetty, Ligong Han, Abhishek Bhandwaldar, Prateek Chanda, Kai Xu, Hao Wang, Aldo Pareja, Oleg Silkin, Mustafa Eyceoz, Akash Srivastava

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 pages, 13 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15485 2025-04-09 cs.CV cs.AI cs.CL cs.LG 67%

TULIP: Towards Unified Language-Image Pretraining

Zineng Tang, Long Lian, Seun Eisape, XuDong Wang, Roei Herzig, Adam Yala, Alane Suhr, Trevor Darrell, David M. Chan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments (v2) Clarified fine-tuning process, updated appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20779 2025-04-01 cs.CL cs.AI cs.LG q-bio.NC 67%

Triple Phase Transitions: Understanding the Learning Dynamics of Large Language Models from a Neuroscience Perspective

Yuko Nakagi, Keigo Tada, Sota Yoshino, Shinji Nishimoto, Yu Takagi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 46 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15140 2025-04-01 cs.CV cs.AI cs.CL cs.LG 67%

Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

Hulingxiao He, Geng Li, Zijun Geng, Jinglin Xu, Yuxin Peng

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published as a conference paper at ICLR 2025. The model is available at https://huggingface.co/StevenHH2000/Finedefics

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13360 2025-03-31 cs.CV cs.AI cs.CL cs.LG cs.MM 67%

RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models

Haoran Hao, Jiaming Han, Changsheng Li, Yu-Feng Li, Xiangyu Yue

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by CVPR 2025. Code: https://github.com/Hoar012/RAP-MLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07807 2025-03-27 cs.CL cs.AI cs.LG 67%

Training Domain Draft Models for Speculative Decoding: Best Practices and Insights

Fenglu Hong, Ravi Raju, Jonathan Lingjie Li, Bo Li, Urmish Thakker, Avinash Ravichandran, Swayambhoo Jain, Changran Hu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published as a workshop paper at SCOPE - ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16970 2025-03-19 cs.CL cs.AI cs.LG 67%

Towards Aligning Language Models with Textual Feedback

Saüc Abadal Lloret, Shehzaad Dhuliawala, Keerthiram Murugesan, Mrinmaya Sachan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏