arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-04 至 2026-03-04 共收录 18 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 18 篇

2506.17871 2026-03-04 cs.CL cs.AI cs.LG 83%

LLM Probability Concentration: How Alignment Shrinks the Generative Horizon

LLM概率集中:对齐如何缩小生成范围

Chenghao Yang, Sida Li, Ari Holtzman

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现对齐微调通过减少生成多样性,使LLM生成更一致,从而影响复杂推理稳定性。

Comments Codebase: https://github.com/yangalan123/LLMBranchingFactor. V3: Significantly rewrite the whole paper for a clearer structure. Correct problems in the theory parts (Remove emphasis on AEP, discussions on variable LLM generation lengths) and strengthen asymptotic analysis. Add Qwen and OLMo2 experiments. Preliminary SFT v.s. RL comparison to better understand the alignment effects on BF

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22903 2026-03-04 cs.IR cs.LG 79%

PSQE: A Theoretical-Practical Approach to Pseudo Seed Quality Enhancement for Unsupervised Multimodal Entity Alignment

PSQE:一种伪种子质量增强的理论-实践方法用于无监督多模态实体对齐

Yunpeng Hong, Chenyang Bu, Jie Zhang, Yi He, Di Wu, Xindong Wu

机构 * Key Laboratory of Knowledge Engineering with Big Data (the Ministry of Education of China), Hefei University of Technology(大数据知识工程重点实验室(教育部)、合肥工业大学) Department of Data Science, College of William and Mary(威廉与玛丽学院数据科学系) College of Computer and Information Science, Southwest University(西南大学计算机与信息科学学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 PSQE通过多模态信息和聚类重采样提升伪种子质量,改善无监督多模态实体对齐的精度与图覆盖平衡性。

Comments 2026 SIGKDD Accept

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02878 2026-03-04 cs.RO astro-ph.EP 78%

Emerging trends in Cislunar Space for Lunar Science Exploration and Space Robotics aiding Human Spaceflight Safety

月球科学探索与空间机器人助力人类航天安全的新兴趋势

Arsalan Muhammad, Yue Wang, Hai Huang, Hao Wang

专题命中 其他安全 :safety(title,abstract)

AI总结 本文探讨了人工智能与空间机器人在月球科学探索及载人航天安全中的应用,旨在推动可持续月球探索和未来火星任务的发展。

Comments Conference Proceedings of 2nd IAA Conference on AI in and for Space (2nd IAA SPAICE), Suzhou, China, 1-3 November, 2025

Journal ref Conference Proceedings of 2nd IAA Conference on AI in and for Space (2nd IAA SPAICE), Suzhou, China, 1-3 November, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02250 2026-03-04 cs.SD eess.AS 78%

SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models

SGPA: 基于频谱的语音对齐用于多模态大语言模型中可行的谢普利值解释

Paweł Pozorski, Jakub Muszyński, Maria Ganzha

机构 * Warsaw University of Technology(华沙技术大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 SGPA通过结合连接主义时间分类和频谱边界细化,实现了多模态大语言模型中可行的音频解释,显著减少了模型评估次数并保持了全局轮廓。

Comments Submitted for admission in Interspeech 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03230 2026-03-04 cs.LG cs.AI cs.CL math.OC 75%

DiaBlo: Diagonal Blocks Are Sufficient For Finetuning

DiaBlo: 对角块足以用于微调

Selcuk Gurses, Aozhong Zhang, Yanxia Deng, Xun Dong, Xin Li, Naigang Wang, Penghang Yin, Zi Yang

机构 * University at Albany, SUNY(纽约州立大学阿尔巴尼分校) IBM T. J. Watson Research Center(IBM 汤普逊·杰·沃森研究中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 DiaBlo是一种仅更新模型权重矩阵对角块的参数高效微调方法,通过消除低秩矩阵乘积需求,实现稳定收敛和高效训练。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16894 2026-03-04 cs.CL 74%

Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment

让LoRA更出色:通过自适应奇异值和专家混合优化对齐提升LoRA

Chenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu, Xiaoye Qu, Wei Wei, Yu Cheng

机构 * School of Computer Science \& Technology, Huazhong University of Science The Chinese University of Hong Kong Zhejiang University

专题命中 其他安全 :alignment(title);分类 cs.CL

AI总结 GOAT通过自适应奇异值和专家混合优化对齐提升LoRA性能,实现在多个任务上的最优表现。

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02532 2026-03-04 cs.CV 67%

EIMC: Efficient Instance-aware Multi-modal Collaborative Perception

EIMC: 高效实例感知多模态协作感知

Kang Yang, Peng Wang, Lantao Li, Tianci Bu, Chen Sun, Deying Li, Yongcai Wang

机构 * School of Information, Renmin University of China(中国人民大学信息学院) Sony Research and Development Center China(索尼(中国)研发有限公司) National University of Defense Technology(国防科技大学)

专题命中 其他安全 :alignment(abstract);safety(abstract)

AI总结 EIMC通过实例感知的多模态协作感知方法,提升自动驾驶安全性,减少带宽使用,实现高效且准确的3D感知。

Comments 9 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02765 2026-03-04 cs.LG cs.AI 62%

Next Embedding Prediction Makes World Models Stronger

下一步嵌入预测使世界模型更强大

George Bredis, Nikita Balagansky, Daniil Gavrilov, Ruslan Rakhimov

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 NE-Dreamer通过时间转换器预测下一步嵌入,无需解码器即可在复杂环境中提升基于模型的强化学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06084 2026-03-04 cs.CL cs.AI 62%

Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability

频谱调节:面向分布覆盖和上下文可引导性的后训练

Taylor Sorensen, Benjamin Newman, Jared Moore, Chan Park, Jillian Fisher, Niloofar Mireshghallah, Liwei Jiang, Yejin Choi

机构 * University of Washington(华盛顿大学) Stanford University(斯坦福大学) Microsoft Research(微软研究院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Spectrum Tuning方法,通过Spectrum Suite提升模型在多样化分布下的引导能力和输出空间覆盖性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03074 2026-03-04 cs.HC cs.AI 61%

Design Generative AI for Practitioners: Exploring Interaction Approaches Aligned with Creative Practice

为实践者设计生成式AI:探索与创意实践对齐的交互方法

Xiaohan Peng, Wendy E. Mackay, Janin Koch

机构 * LISN Université Paris-Saclay, CNRS, Inria(LISN 巴黎-萨克雷大学,法国国家科学研究中心,法国国家信息与自动化技术研究所) UMR 9189 CRIStAL Univ. Lille, Inria, CNRS, Centrale Lille(UMR 9189 CRIStAL 利摩日大学,法国国家信息与自动化技术研究所,法国国家科学研究中心,利摩日中央理工大学)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI

AI总结 本文提出三种交互方法,帮助设计师在不同阶段引导生成式AI与创意实践对齐,强调动态协商与主动/被动角色的适应性。

Comments Accepted to ACM CHI 2026 Workshop on Bidirectional Human-AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02787 2026-03-04 cs.AI 57%

Rethinking Code Similarity for Automated Algorithm Design with LLMs

重新思考基于LLM的自动算法设计中的代码相似性

Rui Zhang, Zhichao Lu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 BehaveSim通过分析问题解决行为轨迹,提升LLM-AAD性能并促进算法分析。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02553 2026-03-04 cs.RO cs.CV cs.HC cs.LG 57%

Give me scissors: Collision-Free Dual-Arm Surgical Assistive Robot for Instrument Delivery

给我剪刀:用于器械传递的碰撞自由双臂手术辅助机器人

Xuejin Luo, Shiquan Sun, Runshi Zhang, Ruizhi Zhang, Junchen Wang

机构 * School of Mechanical Engineering and Automation, Beihang University(机械工程与自动化学院,北航)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出了一种碰撞自由的双臂手术辅助机器人,通过视觉-语言模型生成轨迹并实现动态环境中的安全器械传递。

Comments 8 pages, 10 figures. Accepted by IEEE International Conference on Robotics and Automation (ICRA), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23652 2026-03-04 cs.CV cs.AI 57%

3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection

面向MRI多器官异常检测的3D模态感知预训练

Haowen Zhu, Ning Yin, Xiaogen Zhou

机构 * School of Electronic, Electrical Engineering and Physics, Fujian University of Technology(福建工程学院电子电气工程学院) School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Department of Medical Imaging, Suzhou Traditional Chinese Medicine Hospital, China(苏州中医医院影像科)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出MedMAP框架,通过3D MRI的模态感知预训练提升多器官异常检测性能,实验表明其优于现有视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02901 2026-03-04 cs.HC 50%

Speech recognition assisted by large language models to command software orally -- Application to an augmented and virtual reality web app for immersive molecular graphics

通过大型语言模型辅助的语音识别来控制软件——应用于增强和虚拟现实网页应用中的沉浸式分子图形

Fabio Cortes Rodriguez, Luciano Abriata

专题命中 其他安全 :safety(abstract)

AI总结 通过结合Chrome的语音识别和大型语言模型的函数调用方法,开发了一个语音控制网页应用,用于沉浸式分子图形。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02885 2026-03-04 cs.DC 50%

MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing

MuxTune: 通过时空骨干多路复用实现多租户数据中心中的高效多任务LLM微调

Chunyu Xue, Yi Pan, Weihao Cui, Quan Chen, Shulai Zhang, Bingsheng He, Minyi Guo

专题命中 其他安全 :alignment(abstract)

AI总结 MuxTune通过时空骨干多路复用实现多租户数据中心中多个PEFT任务的高效并发执行,显著提升吞吐量和内存利用率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13251 2026-03-04 cs.CV 50%

Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs

映射流:揭示视频大语言模型中的隐藏信息路径

Minji Kim, Taekyung Kim, Bohyung Han

机构 * NAVER AI Lab(NAVER人工智能实验室)

专题命中 其他安全 :alignment(abstract)

AI总结 本研究通过机理可解释性技术揭示视频大语言模型中信息流的隐藏路径,展示了时间推理机制及提升模型可解释性和泛化能力的贡献。

Comments ICLR 2026, 32 pages, 39 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26645 2026-03-04 cs.CV 50%

TTT3R: 3D Reconstruction as Test-Time Training

TTT3R: 3D重建作为测试时训练

Xingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger, Anpei Chen

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Uni of Tübingen, Tübingen AI Center(图宾根大学,图宾根人工智能中心)

专题命中 其他安全 :alignment(abstract)

AI总结 TTT3R通过测试时训练方法提升3D重建的长度泛化能力,实现全局姿态估计性能提升2倍,运行效率高且内存消耗低。

Comments Page: https://rover-xingyu.github.io/TTT3R/ Code: https://github.com/Inception3D/TTT3R

Journal ref International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02351 2026-03-04 cs.CV 50%

MERG3R: A Divide-and-Conquer Approach to Large-Scale Neural Visual Geometry

MERG3R: 一种用于大规模神经视觉几何的分而治之方法

Leo Kaixuan Cheng, Abdus Shaikh, Ruofan Liang, Zhijie Wu, Yushi Guan, Nandita Vijaykumar

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 其他安全 :alignment(abstract)

AI总结 MERG3R通过分而治之的方法提升大规模神经视觉几何重建的准确性、内存效率和可扩展性。

Comments Project page: https://leochengkx.github.io/MERG3R/

详情

展开后加载摘要…

URL PDF HTML 收藏