arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2408.16482 2025-09-09 cs.CL 79%

Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning

Rochelle Choenni, Ekaterina Shutova

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01897 2025-09-03 cs.LG 79%

Predicting NCAP Safety Ratings: An Analysis of Vehicle Characteristics and ADAS Features Using Machine Learning

Raunak Kunwar, Aera Kim LeBoulluec

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments 11 pages, 4 figures, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00622 2025-09-03 cs.AI cs.IR 79%

BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting

Shiqiao Zhou, Holger Schöner, Huanbo Lyu, Edouard Fouché, Shuo Wang

机构 * University of Birmingham(伯明翰大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00210 2025-09-03 cs.CV cs.AI 79%

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu, Qinhan Lv, Xiying Li

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09625 2025-08-27 cs.CV cs.CL 79%

Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment

Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu, Zequn Jie, Lin Ma, Xu Wang

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Meituan Inc.(美团公司) College of Electronics and Information Engineering, Shenzhen University(深圳大学电子与信息工程学院) Guangdong Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13768 2025-08-26 cs.CL 79%

MGT-Prism: Enhancing Domain Generalization for Machine-Generated Text Detection via Spectral Alignment

Shengchao Liu, Xiaoming Liu, Chengzhengxu Li, Zhaohan Zhang, Guoxin Ma, Yu Lan, Shuai Xiao

机构 * Faculty of Electronic and Information Engineering, Xi’an Jiaotong University(电子与信息工程学院,西安交通大学) Queen Mary University of London(伦敦大学玛丽女王学院) Alibaba(阿里巴巴)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15963 2025-08-25 cs.LG cs.CE cs.SY eess.SP eess.SY physics.ins-det 79%

Advancing rail safety: An onboard measurement system of rolling stock wheel flange wear based on dynamic machine learning algorithms

Celestin Nkundineza, James Ndodana Njaji, Samrawit Abubeker, Omar Gatera, Damien Hanyurwimfura

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments Journal article published in Transportation Research Record: The Journal of Transportation Research Board

Journal ref Transportation Research Record, 2679(7), 791-810 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12803 2025-08-19 cs.CL 79%

When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

Ahmed Elshabrawy, Hour Kaing, Haiyue Song, Alham Fikri Aji, Hideki Tanaka, Masao Utiyama, Raj Dabre

机构 * MBZUAI(马克斯·普朗克人工智能研究所) NICT, Japan(日本信息通信技术研究所) IIT Madras(印度理工学院Madras分校)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11414 2025-08-18 cs.CL 79%

Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions

Shangrui Nie, Florian Mai, David Kaczér, Charles Welch, Zhixue Zhao, Lucie Flek

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 7 pages 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11364 2025-08-18 cs.CL 79%

Feedback Indicators: The Alignment between Llama and a Teacher in Language Learning

Sylvio Rüdian, Yassin Elsir, Marvin Kretschmer, Sabine Cayrou, Niels Pinkwart

机构 * Humboldt-Universität zu Berlin Department of Computer Science(柏林洪堡大学计算机科学系) Humboldt-Universität zu Berlin Language Centre(柏林洪堡大学语言中心) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 11 pages, one table

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10667 2025-08-15 cs.CV cs.AI 79%

AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models

Shixiong Xu, Chenghao Zhang, Lubin Fan, Yuan Zhou, Bin Fan, Shiming Xiang, Gaofeng Meng, Jieping Ye

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) CASIA(中国科学院自动化所) Alibaba Cloud(阿里云) School of Intelligence Science and Technology(智能科学与技术学院) CAIR(中国科学院香港创新研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09203 2025-08-14 cs.LG 79%

Building Safer Sites: A Large-Scale Multi-Level Dataset for Construction Safety Research

Zhenhui Ou, Dawei Li, Zhen Tan, Wenlin Li, Huan Liu, Siyuan Song

机构 * Arizona State University(亚利桑那州立大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments The paper was accepted on the CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06895 2025-08-12 cs.CV cs.AI 79%

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09815 2025-08-11 cs.CL 79%

Statistical Coherence Alignment for Large Language Model Representation Learning Through Tensor Field Convergence

Jonathan Gale, Godfrey Aldington, Harriet Thistlewood, Thomas Tattershall, Basil Wentworth, Vincent Enoasmo

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09485 2025-08-07 cs.LG cs.AR 79%

GenEDA: Towards Generative Netlist Functional Reasoning via Cross-Modal Circuit Encoder-Decoder Alignment

Wenji Fang, Jing Wang, Yao Lu, Shang Liu, Zhiyao Xie

机构 * Hong Kong University of Science and Technology(香港理工大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Accepted by ICCAD'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21082 2025-07-30 cs.CY 79%

Safety Features for a Centralised AGI Project

Sarah Hastings-Woodhouse

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08030 2025-07-14 cs.CL cs.CE cs.HC 79%

A Systematic Analysis of Declining Medical Safety Messaging in Generative AI Models

Sonali Sharma, Ahmed M. Alaa, Roxana Daneshjou

机构 * Department of Medicine, University of British Columbia(英属哥伦比亚大学医学系) Department of Biomedical Data Science, Stanford School of Medicine(斯坦福医学院生物医学数据科学系) University of California, Berkeley(加州大学伯克利分校) University of California, San Francisco(加州大学旧金山分校) Department of Dermatology, Stanford School of Medicine(斯坦福医学院皮肤科)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00216 2025-07-02 cs.CL 79%

Towards Style Alignment in Cross-Cultural Translation

Shreya Havaldar, Adam Stein, Eric Wong, Lyle Ungar

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04443 2025-06-23 cs.AI 79%

POV Learning: Individual Alignment of Multimodal Models using Human Perception

Simon Werner, Katharina Christ, Laura Bernardy, Marion G. Müller, Achim Rettinger

机构 * Trier University(特里尔大学) University of Innsbruck(因斯布鲁克大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13679 2025-06-17 cs.RO cs.AI cs.CV 79%

ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Yuqing Wen, Kefan Gu, Haoxuan Liu, Yucheng Zhao, Tiancai Wang, Haoqiang Fan, Xiaoyan Sun

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学) Dexmal

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09081 2025-06-11 cs.CV cs.AI 79%

Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment

Xiaowei Bi, Zheyuan Xu

机构 * Northwestern University(西北大学) IEEE Member(IEEE会员)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04038 2025-06-05 cs.SE cs.AI 79%

Generating Automotive Code: Large Language Models for Software Development and Verification in Safety-Critical Systems

Sven Kirchner, Alois C. Knoll

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments 8 pages; Accepted for publication at the 36th IEEE Intelligent Vehicles Symposium (IV), Cluj-Napoca, Romania, June 22-25, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17055 2025-06-03 cs.CL 79%

Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning

Chengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen, Yuchen Hu, Bosheng Ding, Ruirui Chen, Shafiq Joty

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Princeton University(普林斯顿大学) Nanyang Technological University(南洋理工大学) Salesforce Research(Salesforce 研究) Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR), Singapore(高性能计算研究所,新加坡科技研究局)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24088 2025-06-02 cs.LG cs.CV 79%

Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting

Chen Huang, Skyler Seto, Hadi Pouransari, Mehrdad Farajtabar, Raviteja Vemulapalli, Fartash Faghri, Oncel Tuzel, Barry-John Theobald, Josh Susskind

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14878 2025-06-02 cs.CL 79%

Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual Alignment

Yongxin Huang, Kexin Wang, Goran Glavaš, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab) Department of Computer Science Technical University of Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany(技术大学达姆施塔特计算机科学系通用知识处理实验室(UKP实验室)和应用网络安全国家研究中心ATHENE德国) Center for AI and Data Science, University of Würzburg(人工智能与数据科学中心,乌尔姆大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted for ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12176 2025-05-30 cs.RO cs.AI 79%

Safety Implications of Explainable Artificial Intelligence in End-to-End Autonomous Driving

Shahin Atakishiyev, Mohammad Salameh, Randy Goebel

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments Accepted for publication in IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21172 2025-05-28 cs.CL 79%

TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment

Zheng Li, Mao Zheng, Mingyang Song, Wenjie Yang

机构 * Tencent Hunyuan(腾讯文言)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21055 2025-05-28 cs.AI 79%

Agent-Environment Alignment via Automated Interface Generation

Kaiming Liu, Xuanyu Lei, Ziyue Wang, Peng Li, Yang Liu

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20303 2025-05-28 cs.SE cs.AI 79%

Future of Code with Generative AI: Transparency and Safety in the Era of AI Generated Software

David Hanson

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12326 2025-05-27 cs.LG 79%

Understanding Why Large Language Models Can Be Ineffective in Time Series Analysis: The Impact of Modality Alignment

Liangwei Nathan Zheng, Chang George Dong, Wei Emma Zhang, Lin Yue, Miao Xu, Olaf Maennel, Weitong Chen

机构 * The University of Adelaide(阿德莱德大学) The University of Queensland(昆士兰大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏