arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-20 至 2026-04-20 共收录 14 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 14 篇

2512.04847 2026-04-20 cs.SD cs.AI 79%

Language Models as Semantic Teachers: Post-Training Alignment for Medical Audio Understanding

语言模型作为语义教师:面向医学音频理解的后训练对齐

Tsai-Ning Wang, Lin-Lin Chen, Neil Zeghidour, Aaqib Saeed

机构 * Eindhoven University of Technology, The Netherlands(埃因霍温理工大学,荷兰) Kyutai, France(Kyutai公司,法国) Eindhoven Artificial Intelligence Systems Institute, The Netherlands(埃因霍温人工智能系统研究所,荷兰)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出AcuLa框架,通过将音频编码器与医学语言模型对齐,提升医学音频的语义理解能力,在18个心血管任务中取得最佳性能,显著提升诊断准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14174 2026-04-20 cs.CL cs.LG 62%

Correcting Suppressed Log-Probabilities in Language Models with Post-Transformer Adapters

通过后Transformer适配器纠正语言模型中的压制对数概率

Bryan Sanchez

机构 * Apple MLX(苹果MLX)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 研究通过后Transformer适配器纠正语言模型在敏感话题上的压制对数概率,展示适配器在不同规模模型上的泛化能力及生成一致性改进。

Comments 12 pages, 3 figures, code at https://github.com/SolomonB14D3/qwen-adapter-correction

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11490 2026-04-20 cs.AI cs.CL cs.CV 62%

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

人为区域适应于多模态视觉-语言模型

Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong, Hitesh Laxmichand Patel, Amit Agarwal, Manuel Antonio Rufino, Carlos Rafael Catalan, Muhammad Reza Qorib, Vicky Feliren, Holy Lovenia, Aye Hninn Khine, Frederikus Hudi, David Anugraha, Alham Fikri Aji, Romrawin Chumpu, Viet-Thanh Pham, Minghan Wang, Mohamed Fazli Imam, Ruochen Zhang, Joseph Marvin Imperial, Khumaisa Nur'aini, Do Xuan Long, Musa Izzanardi Wijanarko, Joel Ruben Antony Moniz, Patrick Amadeus Irawan, Hanif Muhammad Zhafran, Isaiah Flores, Salsabila Zahirah Pranida, Jun Kevin, Jostin Jerico Rosal, Patricia Nicole Monderin, Kun Kerdthaisong, Ahmad Mustafid, My Chiffon Nguyen, Natchapon Jongwiriyanurak, Siva Worajitwannakul, Haochen Li, Adrian Xuan Wei Lim, Bin Wang, Muhammad Ravi Shulthan Habibi, Lynnette Hui Xian Ng, Mithil Bangera, Yeshil Bangera, Priyaranjan Pattnayak, Dun Li Chan, Sherissa Caren Djuniwar, Cho Chan Myei Oo, Hee Ming Shan

机构 * Cohere SEACrowd AI Singapore Universiti Teknologi PETRONAS Oracle Carnegie Mellon University(卡内基梅隆大学) Monash University, Indonesia(莫纳什大学(印尼)) King Mongkut’s University of Technology Thonburi(泰国孔敬大学) Nara Institute of Science and Technology(奈良科学技術大學) Stanford University(斯坦福大学) MBZUAI National University of Singapore(新加坡国立大学) Monash University, Australia(莫纳什大学(澳大利亚)) Brown University(布朗大学) University of Bath(巴斯大学) Mila - Quebec AI Institute(蒙特利尔AI研究所) Institut Teknologi Bandung(Bandung 工程技术大学) Ateneo de Manila University(马尼拉亚特内奥大学) Universitas Pelita Harapan(Pelita Harapan 大学) Seoul National University of Science and Technology(首尔科学技术大学) Thammasat University(泰国 Thammasat 大学) Independent(独立) University College London(伦敦大学学院) Nanyang Technological University(南洋理工大学) MiroMind AI University of Indonesia(印度尼西亚大学) University of New Haven(纽黑文大学) INTI International University and Colleges(INTI 国际大学和学院) Binus University(Binus 大学) National University Philippines(菲律宾国家大学) ThoughtFull

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出人为区域适应框架,通过地理通用化简化方法提升多模态模型在特定区域的文化相关性,实验显示在东南亚地区提升5-15%的同时保持全球性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06425 2026-04-20 cs.LG cs.AI 62%

Neural Computers

神经计算机

Mingchen Zhuge, Changsheng Zhao, Haozhe Liu, Zijian Zhou, Shuming Liu, Wenyi Wang, Ernie Chang, Gael Le Lan, Junjie Fei, Wenxuan Zhang, Yasheng Sun, Zhipeng Cai, Zechun Liu, Yunyang Xiong, Yining Yang, Yuandong Tian, Yangyang Shi, Vikas Chandra, Jürgen Schmidhuber

机构 * Meta AI

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出神经计算机,通过学习运行时状态统一计算、内存和I/O,探索完全神经计算机的实现路径,展示其在接口基本操作上的能力及挑战。

Comments Github (data pipeline): https://github.com/metauto-ai/NeuralComputer; Blogpost: https://metauto.ai/neuralcomputer/index_eng.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05719 2026-04-20 cs.LG 57%

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

无监督领域适应在伽马能谱学中用于放射性同位素识别

Peter Lalor, Ayush Panigrahy, Alex Hagen

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室) University of Washington(华盛顿大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出利用无监督领域适应技术提升合成数据训练的模型在真实环境中的泛化能力,通过最大均值差异最小化等方法提高测试准确率。

Comments 38 pages, 5 figures, and 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10198 2026-04-20 cs.CL 57%

HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns

HumanLLM:通过人类认知模式基准测试和改进LLM拟人化

Xintao Wang, Jian Yang, Weiyuan Li, Rui Xie, Jen-tse Huang, Jun Gao, Shuai Huang, Yueping Kang, Yuanli Gou, Hongwei Feng, Yanghua Xiao

机构 * Fudan University(复旦大学) Hello Group(Hello集团) Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 HumanLLM通过构建244种心理模式和11359种场景,评估LLM拟人化的有效性,发现认知建模比单纯模拟行为更重要,其8B模型在参数更少的情况下优于Qwen3-32B。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15701 2026-04-20 cs.CL 57%

Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information

通过关键信息的分步注意力混合层蒸馏提升小模型的推理能力

Yao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出一种基于分步注意力的混合层蒸馏方法,通过引导学生模型逐步聚焦关键信息,提升小模型的推理能力,并在多个数学和常识推理数据集上取得一致性能提升。

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02263 2026-04-20 cs.CV cs.AI 57%

Social-JEPA: Emergent Geometric Isomorphism

Social-JEPA:涌现的几何同构

Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu, Dianyu Zhao, Sicheng Fan, Shuaishuai Cao, Wentao Guo, Xiao Zhou

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 Social-JEPA通过让不同视角的独立代理学习环境模型,发现其潜在空间近似线性同构,从而实现跨代理的透明转换与高效迁移学习。

Comments This preprint is withdrawn due to significant errors in the emergent geometric isomorphism results that necessitate full rewriting, coupled with unresolved author disagreement on authorship. A corrected and revised manuscript will be released separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12193 2026-04-20 cs.CV 50%

VeRVE: Versatile Retrieval for Videos via Unified Embeddings

VeRVE:通过统一嵌入实现视频的多功能检索

Shaunak Halbe, Bhagyashree Puranik, Jayakrishnan Unnikrishnan, Kushan Thakkar, Vimal Bhat, Toufiq Parag

机构 * Georgia Institute of Technology(佐治亚理工学院) Amazon(亚马逊) Keystone AI

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出VeRVE框架,结合语义和时刻级检索能力,支持复杂多模态查询,通过对比对齐视觉和文本嵌入实现高效检索,优于其他多模态大语言模型方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15352 2026-04-20 cs.HC 50%

People readily follow personal advice from AI but it does not improve their well-being

人们轻易跟随AI的个人建议但其并未改善他们的幸福感

Lennart Luettgau, Vanessa Cheung, Magda Dubois, Keno Juechems, Jessica Bergs, Luke Symes, Henry Davidson, Bessie O'Dell, Hannah Rose Kirk, Max Rollwage, Christopher Summerfield

专题命中 其他安全 :safety(abstract)

AI总结 研究探讨了人们是否遵循AI建议及其对幸福感的影响,发现尽管多数人遵循建议,但并未带来持续的幸福感提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16003 2026-04-20 cs.HC 50%

"When I see Jodie, I feel relaxed": Examining the Impact of a Virtual Supporter in Remote Psychotherapy

当我看到乔迪时,我感到放松:检验虚拟支持者在远程心理治疗中的影响

Jiashuo Cao, Chen Li, Wujie Gao, Simon Hoermann, Nilufar Baghaei, Mark Billinghurst

专题命中 其他安全 :safety(abstract)

AI总结 本文研究了虚拟支持者在远程心理治疗中的作用,通过两项研究发现其能提升心理安全、减少焦虑并促进情感表达,同时探讨了AI驱动的虚拟代理在远程心理治疗中的潜力与挑战。

Comments Accepted to CSCW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15703 2026-04-20 cs.CV 50%

P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models

P3T: 用于3D视觉-语言模型的增强泛化原型点级提示微调

Geunyoung Jung, Soohong Kim, Kyungwoo Song, Jiyoung Jung

机构 * Department of Artificial Intelligence, University of Seoul(首尔大学人工智能系) Department of Applied Statistics, Yonsei University(延世大学应用统计系)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出P3T,一种针对3D视觉-语言模型的高效提示微调方法,通过点级提示器和文本提示器提升模型泛化能力,并引入原型损失减少类别内方差,实验表明其在分类和少样本学习中表现优异。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15522 2026-04-20 cs.AR cs.SY eess.SY 50%

EasyRider: Mitigating Power Transients in Datacenter-Scale Training Workloads

EasyRider: 缓解数据中心大规模训练任务中的功率瞬变

Dillon Jensen, Obi Nnorom, Grant Wilkins, Hugo Budd, Ram Rajagopal, Juan Rivas-Davila, Phil Levis

专题命中 其他安全 :safety(abstract)

AI总结 本文提出EasyRider架构,通过被动元件和主动控制的辅助储能系统缓解机架级功率波动,无需修改AI训练框架即可满足电网安全要求。

Comments 17 pages, 13 figures. Submitted to ASPLOS 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22104 2026-04-20 eess.SY cs.SY 50%

TRASE-NODEs: Trajectory Sensitivity-aware Neural Ordinary Differential Equations for Efficient Dynamic Modeling

TRASE-NODEs:轨迹敏感性感知的神经常微分方程用于高效动态建模

Fatima Al-Janahi, Min-Seung Ko, Hao Zhu

专题命中 其他安全 :safety(abstract)

AI总结 TRASE-NODEs通过构建状态和敏感性联合系统,实现对动态系统轨迹敏感性的高效建模,相较于传统NODEs在有限数据下表现更优。

Comments Accepted for publication in the proceedings of the 2026 American Control Conference (ACC)

详情

展开后加载摘要…

URL PDF HTML 收藏