PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection
机构 * Queen’s University(皇后大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
Comments Accepted to ICCV 2025
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Queen’s University(皇后大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
Comments Accepted to ICCV 2025
机构 * Faculty of Computing - Federal University of Mato Grosso do Sul(计算机学院 - 莫扎尔河大省联邦大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
Comments 12 pages; 2 figures; Preprint with the original submission accepted for publication at 39th Brazilian Symposium on Software Engineering (SBES)
机构 * Charité – Universitätsmedizin Berlin, Humboldt-Universität zu Berlin, Berlin Institute of Health (BIH), Berlin, Germany(柏林查理医院、洪堡-柏林大学、柏林健康研究所(BIH)、柏林)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
机构 * Università della Svizzera italiana \& University of Amsterdam ; University of Amsterdam The Netherland ; Unviersity of Amsterdam The Netherland ; University of Amsterdam
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
Comments Accepted by ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR 2025)
机构 * Department of EECS University of Arkansas Fayetteville(电子工程与计算机科学系美国阿肯色大学弗莱维尔分校) ; Department of CS Baylor University Waco(计算机科学系贝勒大学沃斯堡)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
Comments Accepted to ACL 2025
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
Comments Accepted to ICML 2025 position paper track
机构 * Tsinghua University, Beijing, China(清华大学) ; Beijing Zhongguancun Academy, Beijing, China(北京中关村学院) ; Shanghai Qi Zhi Institute, Shanghai, China(上海启智研究所)
专题命中 AI治理与伦理 :DPO(abstract);分类 cs.AI
Comments Published in ICML 2025
机构 * Zhengzhou University(郑州大学) ; Wuhan University(武汉大学)
专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CL
机构 * Minerva CQ
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
机构 * KAIST AI(韩国科学技术院人工智能研究所)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
Comments 20 pages including references and appendix; To appear in ACL 2025 main conference
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
机构 * INRIA, LMO, Université Paris-Saclay, Orsay, France(INRIA、LMO、巴黎-萨克雷大学、欧萨斯分校、法国) ; TML Lab, EPFL, Switzerland(TML实验室、瑞士联邦理工学院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
Comments ICML camera ready version
机构 * ScaDS.AI and TU Dresden(ScaDS.AI 和 梵高大学) ; LMU Munich(慕尼黑大学) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
机构 * School of Information University of Texas at Austin(信息学院得克萨斯大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
Journal ref CHI EA ' 2025: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
Journal ref ICML 2025
机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) ; AI Lab, China Mobile Communication Group Tianjin Co., Ltd.(中国移动通信集团天津有限公司人工智能实验室)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
Comments Accepted by ACL 2025
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
Comments 10 pages, 1 figure, 1 table, in review for AIES 2025, presented at TAIS 2025
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
Comments Updated abstract to fix a typo; no changes to the content of the paper
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
机构 * HKUST(GZ)(香港科技大学(广州)) ; CSE, HKUST(香港科技大学计算机科学与工程系) ; Xi’an Jiaotong University(西安交通大学) ; University of Pisa, IT(比萨大学) ; University of Trento, IT(特伦特大学) ; Nagoya University(名古屋大学) ; China University of Mining & Technology, Beijing(中国矿业大学(北京)) ; Tongji University(同济大学) ; SPIC Energy Science and Technology Research Institute(SPIC能源科学与技术研究院) ; Shanghai Jiao Tong University(上海交通大学) ; Fudan University(复旦大学) ; College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY
机构 * University of Western Australia(西澳大学) ; University of Melbourne(墨尔本大学) ; University of British Columbia(不列颠哥伦比亚大学) ; Australian National University(澳大利亚国立大学)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI
Comments This research is supported by the NISDRG project #20100007, funded by the Australian Government
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
机构 * MPI-SWS(马克斯·普朗克所社会科学研究院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG