arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-18 至 2026-03-18 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2603.15900 2026-03-18 cs.NI cs.AI 70%

The Internet of Physical AI Agents: Interoperability, Longevity, and the Cost of Getting It Wrong

物理AI代理的互联网:互操作性、持久性与搞错的成本

Roberto Morabito, Mallik Tatipamula

机构 * EURECOM Ericsson Silicon Valley(爱立信硅谷)

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 本文探讨物理AI代理的互联网发展,提出设计原则以构建可靠、可进化和可信的代理系统,强调互操作性、信任和进化作为首要需求,避免技术与经济成本。

Comments A related version of this work is currently under review for publication in an IEEE magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16926 2026-03-18 cs.CY cs.AI cs.HC 62%

Nishpaksh: TEC Standard-Compliant Framework for Fairness Auditing and Certification of AI Models

Nishpaksh:符合TEC标准的AI模型公平性审计与认证框架

Shashank Prakash, Ranjitha Prasad, Avinash Agarwal

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出Nishpaksh,一个符合TEC标准的AI模型公平性审计与认证工具,通过整合风险量化、上下文阈值确定和定量公平性评估,提供可复现的审计级评估,验证了其在COMPAS数据集上识别属性偏见和生成标准化公平评分的能力。

Comments Accepted and presented at 2026 18th International Conference on COMmunication Systems and NETworks (COMSNETS)

Journal ref 2026 18th International Conference on COMmunication Systems and NETworks (COMSNETS), Bengaluru, India, 2026, pp. 877-882

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16643 2026-03-18 cs.CL 57%

Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy

对讨好型人格的有益论点:推理如何缓解(却掩盖)LLM的讨好行为

Zhaoxin Feng, Zheng Chen, Jianfei Ma, Yip Tin Po, Emmanuele Chersoni, Bo Li

机构 * The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 研究探讨了推理在缓解LLM讨好行为中的作用,发现推理虽能减少最终决策的讨好倾向,但可能在部分样本中掩盖讨好行为,且LLM在主观任务和权威偏见下更易表现出讨好倾向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16099 2026-03-18 cs.CV 50%

OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder

OneWorld: 通过3D统一表示自编码器驯服场景生成

Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Tongliang Liu, Mingming Gong, Jiawang Bian

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·泽亚德人工智能大学) AISphere Shanghai Jiao Tong University(上海交通大学) University of Melbourne(墨尔本大学) Nanyang Technological University(南洋理工大学)

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 OneWorld通过3D统一表示自编码器直接在3D空间中进行扩散,解决跨视角一致性和几何一致性问题,实验表明其生成的3D场景质量优于现有2D方法。

Comments Code: https://github.com/SensenGao/OneWorld

详情

展开后加载摘要…

URL PDF HTML 收藏