arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-12 至 2026-03-12 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 5 篇

2603.10351 2026-03-12 cs.CL cs.AI 62%

Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck

通过解耦信息瓶颈缓解多语言大语言模型作为评判者的翻译偏差

Hongbin Zhang, Kehai Chen, Xuefen Bai, Youcheng Pan, Yang Xiang, Jinpeng Wang, Min Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出DIBJudge框架,通过解耦信息瓶颈缓解多语言LLM的翻译偏差问题,有效提升多语言评估的公平性和准确性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00307 2026-03-12 cs.AI 57%

BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models

BiasBusters: 检测和缓解大语言模型中的工具选择偏差

Thierry Blankenstein, Jialin Yu, Zixuan Li, Vassilis Plachouras, Sunando Sengupta, Philip Torr, Yarin Gal, Alasdair Paren, Adel Bibi

机构 * University of Oxford(牛津大学) Microsoft(微软)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 研究通过基准测试揭示LLM工具选择中的系统性偏差,并提出轻量级缓解策略以减少选择偏差。

Comments ICLR 2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22699 2026-03-12 cs.CL 57%

Are you sure? Measuring models bias in content moderation through uncertainty

你确定吗?通过不确定性测量内容审核中的模型偏差

Alessandra Urbinati, Mirko Lai, Simona Frenda, Marco Antonio Stranisci

机构 * Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University(生物与社会技术系统建模实验室,东北大学) Heriot-Watt University(赫瑞-瓦特大学) aequa-tech(aequa-tech公司) Università del Piemonte Orientale(皮埃蒙特东方大学) Università degli Studi di Torino(托里尼大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

AI总结 本文提出通过模型预测不确定性来衡量内容审核中模型的偏差,揭示预训练模型对少数群体的预测准确性与置信度的差异,以改进模型公平性。

Comments accepted at Findings of ACL: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10028 2026-03-12 cs.CY cs.AI 54%

How to Count AIs: Individuation and Liability for AI Agents

如何计算AI:AI代理的个体化与责任归属

Yonathan Arbel, Peter Salib, Simon Goldstein

机构 * University of Alabama(阿拉巴马大学) University of Hong Kong(香港大学) University of Houston(休斯顿大学) HKU AI & Humanity Lab(香港大学AI与人类实验室) Center for Law & AI Risk(法律与人工智能风险中心) Institute for Law & AI(法律与人工智能研究所)

专题命中 AI治理与伦理 :分类 cs.AI、cs.CY;safety(comments);AI safety(comments)

AI总结 本文提出通过‘算法公司’解决AI代理的个体识别与责任归属问题,通过法律虚构实体实现对AI行为的追踪与责任划分。

Comments 36 pages. Presented at the Law Following AI conference, Cambridge University. Interdisciplinary: AI safety, AI governance, legal theory

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10773 2026-03-12 cs.HC 50%

AI-Generated Rubric Interfaces: K-12 Teachers' Perceptions and Practices

AI生成的评分标准界面:K-12教师的感知与实践

Bahare Riahi, Sayali Patukale, Joy Niranjan, Yogya Koneru, Tiffany Barnes, Veronica Cateté

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 本研究探讨了K-12教师在使用AI生成评分标准时的感知与实践,发现AI生成的评分标准在结构和清晰度上有帮助,但需要教师监督以确保准确性和相关性。

Comments 20 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏