arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2503.09638 2025-03-14 cs.RO cs.AI cs.LG 62%

Edge AI-Powered Real-Time Decision-Making for Autonomous Vehicles in Adverse Weather Conditions

Milad Rahmati

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06729 2025-03-11 cs.HC cs.AI cs.CY cs.ET 62%

ACAI for SBOs: AI Co-creation for Advertising and Inspiration for Small Business Owners

Nimisha Karnatak, Adrien Baranes, Rob Marchant, Triona Butler, Kristen Olson

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04684 2025-03-11 cs.LG cs.AI 62%

G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary Diffusion

Mengdi Liu, Zhangyang Gao, Hong Chang, Stan Z. Li, Shiguang Shan, Xilin Chen

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05966 2025-03-11 cs.LG cs.AI 62%

FLOPS: Forward Learning with OPtimal Sampling

Tao Ren, Zishi Zhang, Jinyang Jiang, Guanghao Li, Zeliang Zhang, Mingqian Feng, Yijie Peng

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in the Thirteenth International Conference on Learning Representations(ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05870 2025-03-11 cs.CR cs.CL cs.LG 62%

Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents

Avital Shafran, Roei Schuster, Vitaly Shmatikov

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments To appear in USENIX Security Symposium 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04780 2025-03-10 cs.CL cs.AI physics.atom-ph 62%

MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model

Sumin Ha, Jun Hyeong Kim, Yinhua Piao, Sun Kim

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07335 2025-03-10 cs.LG cs.AI 62%

TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding

Haochuan Zhang, Chunhua Yang, Jie Han, Liyang Qin, Xiaoli Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12360 2025-03-07 cs.CV cs.AI cs.LG 62%

Detecting Systematic Weaknesses in Vision Models along Predefined Human-Understandable Dimensions

Sujan Sai Gannamaneni, Rohil Prakash Rao, Michael Mock, Maram Akila, Stefan Wrobel

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10877 2025-03-07 cs.CL cs.AI 62%

Improving Data Efficiency via Curating LLM-Driven Rating Systems

Jinlong Pang, Jiaheng Wei, Ankit Parag Shah, Zhaowei Zhu, Yaxuan Wang, Chen Qian, Yang Liu, Yujia Bao, Wei Wei

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12106 2025-03-07 cs.CL cs.AI 62%

Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models

Haoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang, Xin Zhang, Guojie Song

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17671 2025-03-06 cs.LG cs.AI 62%

Transfer of Reinforcement Learning-Based Controllers from Model- to Hardware-in-the-Loop

Mario Picerno, Lucas Koch, Kevin Badalian, Marius Wegener, Joschka Schaub, Charles Robert Koch, Jakob Andert

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Transactions on Vehicular Technology (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05316 2025-03-06 cs.LG cs.AI cs.CE q-bio.BM 62%

Aligning Large Language Models and Geometric Deep Models for Protein Representation

Dong Shu, Bingbing Duan, Kai Guo, Kaixiong Zhou, Jiliang Tang, Mengnan Du

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 37 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17037 2025-03-04 cs.CY cs.AI cs.HC 62%

Standardised schema and taxonomy for AI incident databases in critical digital infrastructure

Avinash Agarwal, Manisha J. Nene

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY

Comments 6 pages, 3 tables. Accepted at the 2024 IEEE Pune Section International Conference (PuneCon)

Journal ref IEEE Pune Section International Conference (PuneCon), Pune, India, 2024, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13213 2025-03-04 cs.AI cs.LG 62%

LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch

Caigao Jiang, Xiang Shu, Hong Qian, Xingyu Lu, Jun Zhou, Aimin Zhou, Yang Yu

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15998 2025-03-04 cs.CV cs.AI cs.LG cs.RO 62%

Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, Yilin Zhao, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, Guilin Liu

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Github: https://github.com/NVlabs/Eagle, HuggingFace: https://huggingface.co/NVEagle

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00177 2025-03-04 cs.LG cs.AI 62%

Steering Large Language Model Activations in Sparse Spaces

Reza Bayat, Ali Rahimi-Kalahroudi, Mohammad Pezeshki, Sarath Chandar, Pascal Vincent

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12373 2025-03-03 cs.LG cs.AI 62%

Cell-ontology guided transcriptome foundation model

Xinyu Yuan, Zhihao Zhan, Zuobai Zhang, Manqi Zhou, Jianan Zhao, Boyu Han, Yue Li, Jian Tang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2024 as Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20186 2025-02-28 cs.CL cs.LG 62%

Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge

Yan-Lun Chen, Yi-Ru Wei, Chia-Yi Hsu, Chia-Mu Yu, Chun-Ying Huang, Ying-Dar Lin, Yu-Sung Wu, Wei-Bin Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19347 2025-02-27 cs.CL cs.AI 62%

Controlled Diversity: Length-optimized Natural Language Generation

Diana Marie Schenke, Timo Baumann

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ISCA/ITG Workshop on Diversity in Large Speech and Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15766 2025-02-27 cs.LG cs.CL 62%

Learning Harmonized Representations for Speculative Sampling

Lefan Zhang, Xiaodan Wang, Yanhua Huang, Ruiwen Xu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Published as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17900 2025-02-26 cs.LG cs.AI 62%

Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputs

Che Liu, Cheng Ouyang, Zhongwei Wan, Haozhe Wang, Wenjia Bai, Rossella Arcucci

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03777 2025-02-26 cs.CL cs.AI 62%

Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling

Yuxuan Yao, Han Wu, Mingyang Liu, Sichun Luo, Xiongwei Han, Jie Liu, Zhijiang Guo, Linqi Song

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08038 2025-02-25 cs.LG cs.CL cs.SI 62%

Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized Approach

Hang Gao, Chenhao Zhang, Fengge Wu, Junsuo Zhao, Changwen Zheng, Huaping Liu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16797 2025-02-25 cs.CL cs.AI 62%

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

Alireza Amiri-Margavi, Iman Jebellat, Ehsan Jebellat, Seyed Pouyan Mousavi Davoudi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16198 2025-02-25 cs.NI cs.AI cs.ET cs.LG 62%

An Autonomous Network Orchestration Framework Integrating Large Language Models with Continual Reinforcement Learning

Masoud Shokrnezhad, Tarik Taleb

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments IEEE Communications Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15700 2025-02-25 cs.IR cs.AI cs.CL 62%

Sustainable Digitalization of Business with Multi-Agent RAG and LLM

Muhammad Arslan, Saba Munawar, Christophe Cruz

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18328 2025-02-24 cs.CL cs.CY 62%

Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring

Xuansheng Wu, Padmaja Pravin Saraf, Gyeonggeon Lee, Ehsan Latif, Ninghao Liu, Xiaoming Zhai

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted by Technology, Knowledge, and Learning (TKNL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08877 2025-02-24 cs.SE cs.CL cs.LG 62%

Aligning the Objective of LLM-based Program Repair

Junjielong Xu, Ying Fu, Shin Hwei Tan, Pinjia He

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by ICSE'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14893 2025-02-24 cs.CV cs.AI cs.LG cs.SD eess.AS 62%

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

Mingni Tang, Jiajia Li, Lu Yang, Zhiqiang Zhang, Jinghao Tian, Zuchao Li, Lefei Zhang, Ping Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13619 2025-02-20 cs.CL cs.AI 62%

Complex Ontology Matching with Large Language Model Embeddings

Guilherme Sousa, Rinaldo Lima, Cassia Trojahn

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏