Reentrant phase behavior in binary topological flocks with nonreciprocal alignment
专题命中 其他安全 :alignment(title)
Comments Supplemental movies are available on request
Journal ref Phys. Rev. Research 7, 023008 (2025)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(title)
Comments Supplemental movies are available on request
Journal ref Phys. Rev. Research 7, 023008 (2025)
专题命中 其他安全 :safety(title)
Comments A version of this work has been published in Machine Learning with Applications (MLWA)
Journal ref Machine Learning with Applications Volume 20, June 2025, 100640
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
Comments 4 pages, 3 figures
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
Comments 5 pages, 1 figures
专题命中 其他安全 :alignment(title)
Comments Accepted at WACV 2025
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
Journal ref Nature Communications 15, 7620 (2024)
专题命中 其他安全 :alignment(title)
Comments Accepted by ICML 2024
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
Comments Will be published In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA 2024)
专题命中 其他安全 :safety(title)
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
专题命中 其他安全 :alignment(title)
Comments Accepted in INTERSPEECH 2021
专题命中 其他安全 :safety(title)
专题命中 其他安全 :safety(title)
Comments 8 tables, 2 figures, 7 pages, accepted after peer review as a workshop paper in ACM Conference on Health, Inference, and Learning (CHIL) 2020 https://www.chilconference.org/agenda/
专题命中 其他安全 :safety(title)
Comments The 38th AIAA/IEEE Digital Avionics Systems Conference (DASC)
专题命中 其他安全 :alignment(title)
Comments 30 pages, 15 figures
专题命中 其他安全 :alignment(title)
Comments 6 figures
专题命中 其他安全 :safety(title)
人工智能引发的哲学眩晕
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CY
AI总结 该研究提出人工智能引发的哲学眩晕概念,分析其产生、传播路径,关联临床妄想案例,指出AI将参与重构人类认知环境,并提出哲学可修正性作为应对对策。
Comments 29 pages, no figures. Source formatting revised to improve arXiv HTML accessibility; article text unchanged
链式思维验证器的在线可学习性:正确性与完备性的权衡
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Toyota Technological Institute at Chicago(芝加哥丰田技术研究所) ; Northwestern University(西北大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG
AI总结 本文提出一种在线学习框架,用于学习链式思维验证器,通过检查解决方案的正确性,解决生成器与验证器之间的反馈循环导致的分布偏移问题,并引入新的Littlestone维度扩展以优化验证器的学习。
Comments The abstract has been abridged due to arXiv length constraints
长期测量:迈向对人机交互的纵向理解
机构 * Google Research(谷歌研究院) ; Cornell University(康奈尔大学) ; Stanford University(斯坦福大学) ; Google DeepMind(谷歌DeepMind)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI
AI总结 本研究针对语言模型融入生活引发的长期人机交互风险,结合社会科学测量与NLP计算方法,提出通过长期测量建模人类行为变化,实现问题行为在线检测以缓解用户长期风险。
诊断推理模型中的病理链式思维
机构 * Department of Epidemiology, CAUSALab, Harvard University, Boston, USA(流行病学系、CAUSALab、哈佛大学) ; Imperial College London, London, UK(伦敦帝国学院) ; McGill University, Montreal, Canada(麦吉尔大学) ; Geodesic Research, Cambridge, UK(Geodesic研究)
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文提出了一种评估链式思维推理模型中病态的实用工具,通过定义具体度量标准和训练特定模型生物来识别和区分三种不同的病态现象。
利用知识图谱和大语言模型自动化因果规范生成
机构 * Autonomous Industrial Systems Lab, Imperial College London(帝国理工学院自主工业系统实验室)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI
AI总结 提出一种语义AI框架,结合知识图谱与约束大语言模型,自动生成因果逻辑和操作安全叙述,减少手动工作。
压缩感知在大语言模型能力定位中的应用
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL
AI总结 研究通过压缩感知方法识别大语言模型中特定能力依赖的稀疏注意力头,发现关闭少量头可显著降低特定能力表现,揭示了模型模块化组织原则。