Incomplete Contracting and AI Alignment
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
Comments NIPS workshop on Multi-Modal Machine Learning, 2015
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
Comments Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD, 2014
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
专题命中 其他安全 :safety(title,abstract);分类 cs.AI
Comments Extended Abstract, 3 pages, Accepted at IBM Collaborative Academia Research Exchange (I-CARE)-2011, uses ACM-Proceeding style file
专题命中 其他安全 :safety(title,abstract);分类 cs.AI
Journal ref Journal Of Artificial Intelligence Research, Volume 17, pages 363-378, 2002
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
Journal ref Proceedings of the Workshop and Tutorial on Learning Context-Free Grammars (in association with the 14th European Conference on Machine Learning and the 7th European Conference on Principles and Practice of Knowledge Discovery in Databases (ECML/PKDD 2003), September 2003, Cavtat-Dubrovnik, Croata), editors: C. de la Higuera and P. Adriaans and M. van Zaanen and J. Oncina, pp 113-124
作为人工智能-智能体对齐层次的基础市场设计
专题命中 其他安全 :alignment(title,abstract)
AI总结 研究探讨市场中人工智能-智能体对齐,提出将基础市场设计视为该对齐层次,通过对市场核心形式化建模,利用理论计算机科学严谨性构建透明盒模型,支持激励分析与机制设计,使期望行为受青睐,不良行为难维持。
Comments Accepted as at the EC'26 Workshop on Incentive-Based AI Alignment, co-located with the 27th ACM Conference on Economics and Computation, Rome, Italy, July 2026. This version is prepared for public dissemination following workshop acceptance
视觉语言模型的深度预对齐
机构 * Tsinghua University ; Shanghai Qi Zhi Institute ; Taobao \& Tmall Group of Alibaba
专题命中 其他安全 :alignment(title,abstract)
AI总结 本文提出深度预对齐(DPA),通过替换传统ViT编码器为小型VLM作为感知器,实现视觉特征与目标大语言模型文本空间的深度对齐,提升了多模态基准性能,并降低了语言能力遗忘。
Comments Accepted by ICML 2026. Project Website: https://github.com/THUMAI-Lab/Deep-Pre-Alignment
专题命中 其他安全 :alignment(title,abstract)
Comments GitHub: https://github.com/SHI-Labs/Diffusion-Driven-Test-Time-Adaptation-via-Synthetic-Domain-Alignment
专题命中 其他安全 :safety(title,abstract)
Journal ref Safety Science, 2024, 177, pp.106576
专题命中 其他安全 :alignment(title,abstract)
Comments Update Tri-modal Alignment task
评估大型语言模型在复杂隐藏角色游戏中的表现
机构 * University of Göttingen(哥廷根大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
AI总结 本研究通过社交推理游戏《秘密希特勒》评估大型语言模型的推理、说服和欺骗能力,引入新指标并发现当前模型在复杂多轮操纵中效果不佳。
Comments Master's thesis, University of Göttingen
多智能体AI中的隐藏联盟:来自内部表示的谱诊断
机构 * Reciprocal Research(递归研究) ; Center for the Future of AI, Mind, and Society(人工智能、心智与社会未来中心) ; Florida Atlantic University(佛罗里达 Atlantic 大学) ; Biological and Computational Intelligence Center(生物与计算智能中心) ; National Intelligence University(国家情报大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG
AI总结 本文提出通过分析多智能体系统内部神经表示的谱分区方法,检测隐藏联盟结构,验证了该方法在强化学习和大语言模型中的有效性,揭示了代表层次结构。
Comments 18 pages
信息论视角下欺骗与混淆的区别
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG
AI总结 本文从信息论角度区分两种AI安全失效模式:欺骗对齐与目标漂移,揭示二者在人类-AI系统不同接口的信息分歧,提出形式化模型和思想实验,为大型语言模型对齐挑战提供新视角。
Comments Proceedings of the 14th IJCNLP and the 4th AACL (2025)
语言模型中的地位层级
机构 * Brigham Young University–Hawaii ; COLUMBIA UNIVERSITY(哥伦比亚大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
AI总结 语言模型在多智能体环境中会因地位线索形成层级,高地位分配反而降低高能力模型的服从,揭示AI系统中的新兴社会行为。
机构 * Department of Philosophy University of North Carolina at Chapel Hill(哲学系北卡罗来纳大学教堂山分校) ; Department of Computer Science University of North Carolina at Chapel Hill(计算机科学系北卡罗来纳大学教堂山分校) ; Department of Computer Science University of Texas at Austin(计算机科学系德克萨斯大学奥斯汀分校)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
Comments substantial expansions of sections 4 and 5, updated references, numerous smaller additions and clarifications
机构 * Stanford University(斯坦福大学) ; Mathematical Medicine Group(数学医学组) ; Department of Neurosurgery(神经外科系) ; Physician-Scientist Training Program(医师科学家培训计划)
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY
专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG
CLARA:基于VLM衍生理由的片段级多模态对齐用于仇恨视频检测
机构 * Institute for Analytics and Data Science, University of Essex(埃塞克斯大学分析与数据科学研究所) ; University of Exeter(埃克塞特大学) ; Queen Mary University of London(伦敦大学玛丽皇后学院)
专题命中 其他安全 :alignment(title,abstract)
AI总结 本研究提出CLARA框架,以片段级多模态对齐结合VLM衍生理由,在三个仇恨视频数据集上实现了优于现有最优方法的仇恨视频检测性能。
ConceptFormer:学习自适应潜在概念以实现视觉文档检索中的查询-文档对齐
机构 * Northeastern University(东北大学) ; Tsinghua University(清华大学) ; Peking University(北京大学)
专题命中 其他安全 :alignment(title,abstract)
AI总结 针对视觉文档检索中现有监督信号的局限,本文提出ConceptFormer框架,以自适应潜在概念为中间表示衔接语义鸿沟,在基准测试中较最强基线实现了16.7%、22.1%的NDCG@10相对提升,性能优异。
AlignJEPA:面向遥感基础模型的预测性视觉-语言对齐
机构 * Space Applications Centre, ISRO(印度空间研究组织空间应用中心) ; Indian Institute of Science Education and Research (IISER) Bhopal(印度科学教育与研究学院(博帕尔)) ; Indian Institute of Technology Bombay(印度理工学院孟买分校)
专题命中 其他安全 :alignment(title,abstract)
AI总结 该研究针对遥感基础模型与自然语言对齐不足的问题,提出受JEPA启发的AlignJEPA框架,采用轻量级预测对齐网络,结合语义预测与双向对比检索,实现了参数高效的视觉-语言对齐。
Comments 18 pages
别让AI代理占用你的文件:将信息和控制移交给文件系统以实现代理安全和自主性
专题命中 其他安全 :safety(title,abstract)
AI总结 本文研究了AI代理对文件系统的滥用问题,提出YoloFS文件系统通过三种技术提升代理的安全性和自主性,减少用户交互并提高任务完成率。
RefineAny3D:作为语义对齐的深度细化用于单目3D检测
机构 * Michigan State University(密歇根州立大学) ; University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 其他安全 :alignment(title,abstract)
AI总结 RefineAny3D将单目3D检测的深度细化转化为视觉对齐问题,通过VLM实现无需数值预测的深度修正,在多类检测工具上均有性能提升且可泛化。
LASA:面向域泛化语义分割的语言与源锚定对齐
机构 * Xiamen University(厦门大学)
专题命中 其他安全 :alignment(title,abstract)
AI总结 针对域泛化语义分割中传统方法损害特征完整性的问题,提出LASA框架,含三个协同组件,实验显示其性能优于现有最优方法。
Comments 10 pages, 4 figures
面向汽车基于模型的系统工程的有状态多智能体大语言模型,用于跨视图接口对齐
专题命中 其他安全 :alignment(title,abstract)
AI总结 针对LLMs在汽车MBSE中引发的架构漂移问题,提出有状态多智能体验证流水线,在ADAS场景中实现97%实体可追溯性等指标,证明对抗性审核可让LLMs可靠生成零错误MBSE架构。
AirSplat:基于鲁棒前馈3D高斯散射的对齐与评分
机构 * KAIST(韩国科学技术院)
专题命中 其他安全 :alignment(title,abstract)
AI总结 本文提出AirSplat框架,通过自一致性姿态对齐和基于评分的透明度匹配技术,提升无姿态视角合成的重建质量。
Comments Project page: https://kaist-viclab.github.io/airsplat-site, accepted to ECCV 2026
PatchAlign3D: 本地特征对齐用于密集3D形状理解
机构 * École polytechnique(巴黎政治学院) ; University of Virginia(弗吉尼亚大学) ; KAUST(王国立阿联酋科技大学)
专题命中 其他安全 :alignment(title,abstract)
AI总结 PatchAlign3D通过点云直接生成语言对齐的片段特征,实现高效零样本3D部分分割,优于传统渲染方法。
Comments CVPR 2026. Project website: https://souhail-hadgi.github.io/patchalign3dsite/
ROAD:用于3D形状生成的判别式语义的互目标对齐
机构 * Huazhong University of Science and Technology(华中科技大学) ; Megvii(旷视科技) ; Zhejiang University(浙江大学)
专题命中 其他安全 :alignment(title,abstract)
AI总结 ROAD框架通过迁移判别式3D基础模型的先验,采用互目标对齐策略,仅用1.5%训练数据就实现了高保真3D生成,大幅降低了计算开销。
AgentHOI:通过隐式表示对齐进行人类-物体交互视频生成的多智能体推理
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; Tencent HunYuan(腾讯混元)
专题命中 其他安全 :alignment(title,abstract)
AI总结 研究针对HOI视频生成中现有方法依赖显式运动控制的问题,提出AgentHOI,通过多智能体推理和隐式文本-运动对齐策略,实现文本驱动的HOI视频生成,提升了交互自然性等,改进了复杂场景下的表现。
Comments Under review. The code is available at https://github.com/bone-11/agenthoi