arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7968 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7968 篇

2501.06164 2025-11-04 cs.LG cs.AI 76%

Model Alignment Search

Satchel Grant

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11018 2025-10-21 cs.CL cs.AI 76%

GRIFFIN: Effective Token Alignment for Faster Speculative Decoding

Shijing Hu, Jingyang Li, Xingyu Xie, Zhihui Lu, Kim-Chuan Toh, Pan Zhou

机构 * Fudan University(复旦大学) National University of Singapore(新加坡国立大学) Singapore Management University(新加坡管理学院)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21857 2025-10-15 cs.CV cs.AI cs.LG 76%

SPADE: Spatial Transcriptomics and Pathology Alignment Using a Mixture of Data Experts for an Expressive Latent Space

Ekaterina Redekop, Mara Pleasure, Zichen Wang, Kimberly Flores, Anthony Sisk, William Speier, Corey W. Arnold

机构 * Biomedical AI Research Lab, University of California, Los Angeles(生物医学人工智能研究实验室,加州大学洛杉矶分校) Department of Pathology, University of California, Los Angeles(病理学系,加州大学洛杉矶分校) Department of Radiology, University of California, Los Angeles(放射学系,加州大学洛杉矶分校) Department of Bioengineering, University of California, Los Angeles(生物工程系,加州大学洛杉矶分校) UCLA Medical Informatics Home Area, University of California, Los Angeles(UCLA医学信息学家庭区域,加州大学洛杉矶分校) Department of Computational Medicine, University of California, Los Angeles(计算医学系,加州大学洛杉矶分校)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16633 2025-09-23 cs.CV cs.AI cs.CL 76%

When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs

Abhirama Subramanyam Penamakuri, Navlika Singh, Piyush Arora, Anand Mishra

机构 * Indian Institute of Technology Jodhpur(印度理工学院朱诺尔)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted to EMNLP (Main) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10057 2025-08-15 q-bio.NC cs.AI cs.CL 76%

Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

Christopher Pinier, Sonia Acuña Vargas, Mariia Steeghs-Turchina, Dora Matzke, Claire E. Stevenson, Michael D. Nunez

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Presented at the 8th Annual Conference on Cognitive Computational Neuroscience (August 12-15, 2025; Amsterdam, The Netherlands); 20 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07525 2025-07-23 cs.CV cs.AI cs.LG 76%

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

Difei Gu, Yunhe Gao, Yang Zhou, Mu Zhou, Dimitris Metaxas

机构 * Rutgers University(罗格斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08665 2025-07-14 cs.CL cs.AI 76%

KELPS: A Framework for Verified Multi-Language Autoformalization via Semantic-Syntactic Alignment

Jiyao Zhang, Chengli Zhong, Hui Xu, Qige Li, Yi Zhou

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) USTC Knowledge Computing Lab(中国科学技术大学知识计算实验室)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted by the ICML 2025 AI4MATH Workshop. 22 pages, 16 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01069 2025-06-03 cs.CV cs.AI cs.LG 76%

Revolutionizing Blood Banks: AI-Driven Fingerprint-Blood Group Correlation for Enhanced Safety

Malik A. Altayar, Muhyeeddin Alqaraleh, Mowafaq Salem Alzboon, Wesam T. Almagharbeh

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Journal ref Data and Metadata [Internet]. 2025 Apr. 7 [cited 2025 Jun. 1];4:894

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24040 2025-06-02 cs.CL cs.AI 76%

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Yuexing Hao, Kumail Alhamoud, Hyewon Jeong, Haoran Zhang, Isha Puri, Philip Torr, Mike Schaekermann, Ariel D. Stern, Marzyeh Ghassemi

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14468 2025-04-22 cs.CL cs.LG eess.SP q-bio.NC 76%

sEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment

Yijun Liu

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

Comments Accepted for poster presentation at the CVPR 2025 Workshop on Multimodal Foundation Models (MMFM3)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13825 2025-04-21 cs.CL cs.LG 76%

Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models

Junjie Yang, Junhao Song, Xudong Han, Ziqian Bi, Tianyang Wang, Chia Xin Liang, Xinyuan Song, Yichao Zhang, Qian Niu, Benji Peng, Keyu Chen, Ming Liu

机构 * Pingtan Research Institute of Xiamen University(厦门大学滨海研究院) Imperial College London(伦敦帝国理工学院) University of Sussex(苏塞克斯大学) Purdue University(普渡大学) University of Liverpool(利物浦大学) JTB Technology Corp.(JTB技术公司) Emory University(埃默里大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) Kyoto University(京都大学) AppCubic Georgia Institute of Technology(佐治亚理工学院) AI Agent Lab(AI代理实验室)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05803 2025-04-18 cs.LG cs.AI cs.CV math.ST stat.TH 76%

Test-time Alignment of Diffusion Models without Reward Over-optimization

Sunwoo Kim, Minkyu Kim, Dongmin Park

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments ICLR 2025 (Spotlight). The Thirteenth International Conference on Learning Representations. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12900 2025-02-19 cs.CL cs.AI cs.SD 76%

Soundwave: Less is More for Speech-Text Alignment in LLMs

Yuhao Zhang, Zhiheng Liu, Fan Bu, Ruiyu Zhang, Benyou Wang, Haizhou Li

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11205 2025-02-18 cs.LG cs.CY 76%

Deep Contrastive Learning for Feature Alignment: Insights from Housing-Household Relationship Inference

Xiao Qian, Shangjia Dong, Rachel Davidson

专题命中 其他安全 :alignment(title);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18539 2025-01-31 cs.CL cs.AI cs.IR 76%

Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval Method

Peter Baile Chen, Yi Zhang, Michael Cafarella, Dan Roth

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12273 2025-01-22 cs.CL cs.AI 76%

Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement

Maosong Cao, Taolin Zhang, Mo Li, Chuyu Zhang, Yunxin Liu, Haodong Duan, Songyang Zhang, Kai Chen

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Tech Report. Github: https://github.com/InternLM/Condor

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06846 2024-12-11 cs.LG cs.AI 76%

Classifier-free guidance in LLMs Safety

Roman Smirnov

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01101 2024-11-19 cs.LG cs.AI 76%

Feature Alignment: Rethinking Efficient Active Learning via Proxy in the Context of Pre-trained Models

Ziting Wen, Oscar Pizarro, Stefan Williams

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accepted by Transactions on Machine Learning Research (TMLR, 2024) https://openreview.net/forum?id=PNcgJMJcdl

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10852 2024-10-16 cs.CL cs.AI 76%

SafeLLM: Domain-Specific Safety Monitoring for Large Language Models: A Case Study of Offshore Wind Maintenance

Connor Walker, Callum Rothon, Koorosh Aslansefat, Yiannis Papadopoulos, Nina Dethlefs

专题命中 其他安全 :safety(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16645 2024-09-26 cs.LG cs.AI 76%

Task Addition in Multi-Task Learning by Geometrical Alignment

Soorin Yim, Dae-Woong Jeong, Sung Moon Ko, Sumin Lee, Hyunseung Kim, Chanhui Lee, Sehui Han

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments 11 pages, 5 figures, Accepted at AI for Science Workshop at 41st International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12097 2024-09-20 cs.CL cs.IR cs.LG cs.SI 76%

Skill matching at scale: freelancer-project alignment for efficient multilingual candidate retrieval

Warren Jouanneau, Marc Palyart, Emma Jouffroy

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08078 2024-08-20 cs.CV cs.AI cs.CL 76%

Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation

Wenting Chen, Linlin Shen, Jingyang Lin, Jiebo Luo, Xiang Li, Yixuan Yuan

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024

Journal ref https://aclanthology.org/2024.acl-long.514/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06725 2024-08-14 cs.AI cs.CL cs.CV 76%

Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations

Wei Pang, Ruixue Duan, Jinfu Yang, Ning Li

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments This article has been accepted in CAAI Transactions on Intelligence Technology! Article ID: CIT2_12370, Article DOI: 10.1049/cit2.12370

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06540 2024-08-14 physics.acc-ph cs.AI cs.LG 76%

Dynamic Exclusion of Low-Fidelity Data in Bayesian Optimization for Autonomous Beamline Alignment

Megha R. Narayanan, Thomas W. Morris

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments 12 pages, 6 figure sets

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19594 2024-07-31 cs.CL cs.AI 76%

Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01103 2024-06-05 cs.AI cs.HC cs.LG 76%

Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human Alignment

Chen Zhang, Qiang He, Zhou Yuan, Elvis S. Liu, Hong Wang, Jian Zhao, Yang Wang

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accept at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15823 2024-01-23 cs.CL cs.AI 76%

Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment

Ahmed ElBakry, Mohamed Gabr, Muhammad ElNokrashy, Badr AlKhamissi

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Proceedings of ArabicNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13627 2023-10-25 cs.CL cs.AI 76%

InstructAlign: High-and-Low Resource Language Alignment via Continual Crosslingual Instruction Tuning

Samuel Cahyawijaya, Holy Lovenia, Tiezheng Yu, Willy Chung, Pascale Fung

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16424 2023-09-29 cs.CL cs.AI cs.SI 76%

Prompt-and-Align: Prompt-Based Social Alignment for Few-Shot Fake News Detection

Jiaying Wu, Shen Li, Ailin Deng, Miao Xiong, Bryan Hooi

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted to CIKM 2023 (Full Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05831 2023-09-13 cs.LG cs.AI 76%

Studying Accuracy of Machine Learning Models Trained on Lab Lifting Data in Solving Real-World Problems Using Wearable Sensors for Workplace Safety

Joseph Bertrand, Nick Griffey, Ming-Lun Lu, Rashmi Jha

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments 7 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏