Socially Responsible AI Algorithms: Issues, Purposes, and Challenges
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
Comments 45 pages, 8 figures
Journal ref Journal of Artificial Intelligence Research 71 (2021) 1137-1181
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
Comments 45 pages, 8 figures
Journal ref Journal of Artificial Intelligence Research 71 (2021) 1137-1181
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
Comments Accepted to the 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG
Comments Originally presented at the ICML 2020 Participatory Approaches to Machine Learning workshop
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
Comments 12 pages, 2 figures, 3 tables
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
Comments (Nahian and Frazier contributed equally to this work)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG
Comments In Proceedings of the AAAI 2021 Conference. Supersedes arXiv:1902.09980, arXiv:2001.07118
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
Comments A Computing Community Consortium (CCC) white paper, 5 pages
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments AISTATS 2020
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG
Comments Where is the Human? Bridging the Gap Between AI and HCI, CHI Workshop 2019
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG
Comments 4 pages, 4th Multidisciplinary Conference on Reinforcement Learning and Decision Making
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
Comments 31 pages
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG
Comments To appear in the 2019 ACM CHI Conference on Human Factors in Computing Systems (CHI 2019)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG
Comments Published in proceedings of the 16th International Conference on Autonomous Agents and Multi-agent Systems
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
Comments 14 pages, typo fixed in v2
Journal ref Big Data & Society 4(2). 2017
对基于AI的紧急警务调度中的种族偏见进行审计:对十一种大型语言模型的跨语言评估
机构 * Department of Industrial Engineering, Tsinghua University(清华大学工业工程系) ; School of Social Sciences, Tsinghua University(清华大学社会科学部) ; Department of Industrial Engineering, Federal University of Rio de Janeiro(里约热内卢联邦大学工业工程系)
专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.CL
AI总结 本文通过跨语言框架评估11种模型,在19800个输出中发现当事件严重性模糊时种族偏见系统性出现,但当操作优先级由通话内容确定时偏见消失。偏见程度因种族轴而异,宗教外观影响最大,性别次之,种族最小。语言间偏见转移不一致,性别偏见在中文中放大,种族偏见在英文中更明显。
Comments 26 pages, 7 figures. Submitted to Humanities and Social Sciences Communications (Nature) collection on Artificial Intelligence and Emerging Technologies in Public Safety. Code and data: https://github.com/williamguey/llmdispatchbias
超越静态沙箱:为自主AI代理的学得能力治理
机构 * Institute for Applied AI Research(应用人工智能研究所) ; Faculty of Computer and Information Science(计算机与信息科学学院) ; Ben-Gurion University of the Negev(贝内-约尔大学)
专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.AI
AI总结 本文提出Aethelgard框架,通过学得策略实现AI代理的最小必要能力集,解决能力过度配置问题。
Comments 17 pages (9 content pages), 2 figures, 7 tables. Submitted to NeurIPS 2026 Agent Safety Workshop. Code and dataset available at https://github.com/sidikbro/aethelgard
人工智能为你的社区发声:通过AI代理收集数据中心项目公众意见
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY;trustworthy(comments)
AI总结 本文提出AI代理调查框架,利用大型语言模型评估社区对数据中心项目的意见,以指导负责任的AI发展。
Comments 35 Pages. Accepted to NeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models (ResponsibleFM)
专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CY
Comments 31 pages. Accepted to the position paper track at ICML 2025. A previous version was presented at the Pluralistic Alignment Workshop at NeurIPS 2024. For ongoing work, see: https://democracylevels.org
机构 * Machine Learning Theory and Applications (MALTA) Lab, PUCRS(机器学习理论与应用(MALTA)实验室,PUCRS)
专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.CY;trustworthy(comments)
Comments Submitted to ACM Computing Surveys - Special Issue on Trustworthy AI
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL;safety(comments)
Comments NeurIPS 2024 SoLaR Workshop and NeurIPS 2024 Safety Gen AI Workshop
Journal ref NeurIPS Safe Generative AI Workshop 2024; Workshop on Socially Responsible Language Modelling Research 2024
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG;trustworthy(comments)
Comments Best paper runner-up at ICLR 2023 Workshop on Trustworthy and Reliable Large-Scale Machine Learning Models (Non Archival)
Journal ref Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.CY;trustworthy(journal_ref)
Comments TrustNLP at ACL 2023
Journal ref Proceedings at The Third Workshop on Trustworthy Natural Language Processing collocated at the 61st Annual Meeting Of The Association For Computational Linguistics. 2023
专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.LG
Comments 8 pages, 2 figures. To appear in 2022 Trustworthy and Socially Responsible Machine Learning (TSRML 2022) co-located with NeurIPS 2022
在块内思考:使用大语言模型生成符合法规的场景的RegulaRAG——以联合国第152号法规为例
机构 * Technical University of Munich(慕尼黑工业大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
AI总结 针对LLMs难以结合冗长分层标准的问题,提出RegulaRAG流水线,经实验其在UN R152数据集上元分数最高且鲁棒性强,优于基线系统。
天生不连贯?大型语言模型的道德自我一致性研究
机构 * Cornell University(康奈尔大学) ; Cornell Tech(康奈尔科技学院) ; Nanyang Technological University(南洋理工大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
AI总结 本研究针对GPT、Mistral、Llama等LLM,在义务论等三大伦理框架下,发现其道德推理存在最高78%的矛盾率,内部不连贯是AI对齐的必要前提。
Comments 88 pages; pages 16 to 88 are the appendix