Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives
机构 * Stanford University(斯坦福大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG
Comments 34 pages, 13 figures
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Stanford University(斯坦福大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG
Comments 34 pages, 13 figures
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
Comments This work has been accepted for publication as a full paper at the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
Comments 40 pages, 14 figures, 16 tables. To be published in Nature Scientific Reports
机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院) ; Institute of Entrepreneurship and Management, HES-SO Valais-Wallis(创业与管理研究所) ; Institute of Informatics, HES-SO Valais-Wallis(信息研究所)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
机构 * Department of Physics University of Basel(物理系 巴塞尔大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
Comments 11+25 pages, 4+11 figures
机构 * KIIT Deemed University(KIIT大学) ; Indian Institute of Technology (IIT) Bhubaneswar(印度理工学院(Bhubaneswar分校))
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY
Comments Accepted at ASI @ ICCV 2025
机构 * Namibia University of Science \& Technology 13 Jackson Kaujeua Windhoek Namibia 9000 ; Rhodes University Makhanda South Africa ; International University of Management Namibia ; Charles Darwin University Australia ; Namibia University of Science \& Technology ; Rhodes University ; International University of Management ; Charles Darwin University
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
Comments 2025 AAAI Conference on AI, Ethics, and Society
机构 * Microsoft Research AI for Science(微软研究院人工智能与科学研究中心) ; Novartis Biomedical Research(诺华生物医学研究) ; University of Cambridge(剑桥大学) ; Jagiellonian University(雅盖隆大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
Comments Conference version: AIES 2025 (non-archival track), 12 pages
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Apple(苹果公司)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * Independent Researcher(独立研究者)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
Comments Published in MIT Science Policy Review 6, 139-146 (2025)
Journal ref MIT Science Policy Review, 6. (2025)
机构 * Johannes Kepler University (JKU)(约翰内斯·开普勒大学) ; Linz Institute of Technology (LIT)(林茨技术研究所) ; University of Innsbruck(因斯布鲁克大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
Comments Under Review
机构 * CFCS, School of Computer Science, Peking University(计算机科学系,北京大学) ; School of Business, Jiangnan University(商学院,江南大学) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
Comments A shorter conference version is published in IJCAI 2025, titled 'Game Theory Meets Large Language Models: A Systematic Survey'
机构 * Huazhong University of Science and Technology(华中科技大学) ; Lehigh University(莱斯大学) ; The University of Hong Kong(香港大学) ; Jilin University(吉林大学) ; Southern University of Science and Technology(南方科技大学) ; Worcester Polytechnic Institute(沃思堡理工学院) ; LinkedIn Corporation(领英公司) ; Squirrel Ai Learning ; University of Georgia(佐治亚大学) ; Duke University(杜克大学) ; Michigan State University(密歇根州立大学) ; Salesforce Research(Salesforce研究) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Illinois at Chicago(伊利诺伊大学芝加哥分校) ; Microsoft Research(微软研究院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
Comments 87 pages, 21 figures, 9 tables
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
Comments This project has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement number 101142306. The project is also supported by the Center for Digital Narrative, which is funded by the Research Council of Norway through its Centres of Excellence scheme, project number 332643
Journal ref Open Research Europe 2025, 5:202 [version 1; peer review: awaiting peer review]
机构 * Independent Researcher(独立研究者)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Spotify
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Yale University(耶鲁大学) ; Hong Kong University of Science(香港科学大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
Comments 11 Pages, SIGKDD 2025
机构 * Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京中国) ; Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 印度昌迪加尔大学 摩哈利)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
机构 * FAIR at Meta(Meta 的 FAIR)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG
机构 * Journal for Language Technology and Computational Linguistics(语言技术与计算语言学期刊)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
Comments 14 pages, 2 tables
Journal ref JLCL 2025, Band 38(2)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG
机构 * Department of Computer Science and A.I, University of Granada(计算机科学与人工智能系,格拉纳达大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
Comments 37 pages, under review in WIREs Data Mining and Knowledge Discovery
Journal ref Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery (2025), 15(3), e70029
机构 * University of Tartu, Institute of Computer Science(塔尔图大学计算机科学研究所)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 25 pages, 6 figures
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY
机构 * School of Computer Science and DAIM University of Hull(计算机科学学院和DAIM赫尔大学)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG