arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-08-18 至 2026-08-18 共收录 155 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 35 篇

2608.14893 2026-08-18 cs.SI physics.soc-ph 新提交 50%

Weaker Coherence, Weaker Reciprocity: Comparing the Semantic and Social Organization of Moltbook and Reddit

连贯性更弱,互惠性更弱:对比Moltbook与Reddit的语义及社交组织

Favio Di Ciocco, Lucas Díaz Celauro, Sebastián Pinto, Marcelo Kuperman, Pablo Balenzuela

专题命中 其他安全 :alignment(abstract)

AI总结 该研究对比AI智能体社交网络Moltbook与早期Reddit,发现Reddit在语义连贯性、多样性及交互互惠性上更优,且未被Moltbook复现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13989 2026-08-18 eess.AS 50%

Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems

你听到我意味着什么?量化指令-感知差距在指令引导的表达性文本到语音系统中

Yi-Cheng Lin, Huang-Cheng Chou, Tzu-Chieh Wei, Kuan-Yu Chen, Hung-yi Lee

专题命中 其他安全 :alignment(abstract)

AI总结 本文研究了指令引导的表达性文本到语音系统中指令与感知之间的差距,通过感知分析和人类评价揭示了指令可控性及系统改进方向。

Comments Accepted to ICASSP 2026

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 16472-16476

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04854 2026-08-18 cs.SE 版本更新 50%

Assessing Large Language Models for Stabilizing Numerical Expressions in Scientific Software

评估大语言模型在科学软件中稳定数值表达式的能力

Tien Nguyen, Kirshanthan Sundararajah, Muhammad Ali Gulzar

专题命中 其他安全 :safety(abstract)

AI总结 本文评估大语言模型在数值稳定性任务中的表现,发现其在检测和稳定不稳定的计算方面与传统方法相当,并在某些情况下表现更优,但对控制流和高精度字面量处理较弱。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00600 2026-08-18 cs.RO 版本更新 50%

I-Perceive: A Foundation Model for Active Perception with Language Instructions

I-Perceive:一种基于语言指令的主动感知基础模型

Yongxi Huang, Zhuohang Wang, Wenjing Tang, Xinyu He, Cewu Lu, Panpan Cai

机构 * Shanghai Innovation Institute(上海创新研究院) Beihang University(北京航空航天大学)

专题命中 其他安全 :alignment(abstract)

AI总结 I-Perceive是一种基于自然语言指令的主动感知基础模型,通过融合视觉-语言模型与几何基础模型,实现了对开放意图场景的有效感知与推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18005 2026-08-18 cs.CV 版本更新 50%

UrbanWorld2.0: A Multimodal Agentic Framework for Reality-Aligned 3D World Generation at City-Scale

RAISECity: 一种用于城市级现实对齐3D世界生成的多模态代理框架

Shengyuan Wang, Zhiheng Zheng, Yu Shang, Lixuan He, Yangcheng Yu, Fan Hangyu, Jie Feng, Qingmin Liao, Yong Li

机构 * College of AI, Tsinghua University(人工智能学院,清华大学) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系,北京研究院,清华大学)

专题命中 其他安全 :alignment(abstract)

AI总结 RAISECity通过多模态代理框架实现城市级3D世界生成,提升现实对齐、精度和性能,适用于沉浸媒体和具身智能应用。

Comments Accepted by ACM MM 2026, the code is available at: https://github.com/tsinghua-fib-lab/UrbanWorld2.0

详情

展开后加载摘要…

URL PDF HTML 收藏