CommentsComments: 11 pages; Code is available at https://github.com/mujocolab/mjlab ; Expanded sensor and domain randomization sections, added references, minor edits
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
机器去学习并不如你所想:生成式AI政策与研究的启示
A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen, Kevin Klyman, Matthew Jagielski, Katja Filippova, Ken Liu, Alexandra Chouldechova, Jamie Hayes, Yangsibo Huang, Eleni Triantafillou, Peter Kairouz, Nicole Elyse Mitchell, Niloofar Mireshghallah, Abigail Z. Jacobs, James Grimmelmann, Vitaly Shmatikov, Christopher De Sa, Ilia Shumailov, Andreas Terzis, Solon Barocas, Jennifer Wortman Vaughan, danah boyd, Yejin Choi, Sanmi Koyejo, Fernando Delgado, Percy Liang, Daniel E. Ho, Pamela Samuelson, Miles Brundage, David Bau, Seth Neel, Hanna Wallach, Amy B. Cyphert, Mark A. Lemley, Nicolas Papernot, Katherine Lee
机构
*
The GenLaw Center(GenLaw中心)
;
Microsoft Research(微软研究院)
;
Stanford University(斯坦福大学)
;
Google DeepMind(谷歌DeepMind)
;
Center for Democracy & Technology(民主与科技中心)
;
Princeton(普林斯顿)
;
Google(谷歌)
;
University of Washington(华盛顿大学)
;
University of Michigan(密歇根大学)
;
Cornell Tech(康奈尔科技)
;
Cornell Law School(康奈尔法学院)
;
Cornell University(康奈尔大学)
;
Lighthouse
;
Stanford Law School(斯坦福法学院)
;
UC Berkeley(伯克利大学)
;
Independent(独立研究者)
;
Northeastern University(东北大学)
;
Harvard Business School(哈佛商学院)
;
W. Virginia University College of Law(维珍尼亚大学法学院)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
基于描述的可解释性:通过CLIP探测CNN中的偏见
Patrick Koller, Amil V. Dravid, Guido M. Schuster, Aggelos K. Katsaggelos
机构
*
Northwestern University, Evanston, IL, USA(西北大学)
;
University of California, Berkeley, CA, USA(加州大学伯克利分校)
;
Eastern Switzerland University of Applied Sciences, Rapperswil, SG, CH(东瑞士应用科学大学)
AI总结
本文提出一种基于描述的XAI方法,通过CLIP模型探测CNN中的偏见,以提高模型鲁棒性。
CommentsAccepted and presented at the IEEE ICIP 2025 Satellite Workshop "Generative AI for World Simulations and Communications & Celebrating 40 Years of Excellence in Education: Honoring Prof. Aggelos Katsaggelos", Anchorage, USA, Sept 14, 2025. Camera-ready preprint; IEEE Xplore version to follow. Author variant: Amil Dravid. Code: https://github.com/patch0816/caption-driven-xai
Journal ref2025 IEEE International Conference on Image Processing Workshops (ICIPW), IEEE, 2025
Spurious Rewards: Rethinking Training Signals in RLVR
虚假奖励:重新思考强化学习中的训练信号
Rulin Shao, Shuyue Stella Li, Rui Xin, Scott Geng, Yiping Wang, Sewoong Oh, Simon Shaolei Du, Nathan Lambert, Sewon Min, Ranjay Krishna, Yulia Tsvetkov, Hannaneh Hajishirzi, Pang Wei Koh, Luke Zettlemoyer
机构
*
University of Washington, Seattle, WA, USA(华盛顿大学)
;
Allen Institute for Artificial Intelligence, Seattle, WA, USA(人工智能研究院)
;
University of California, Berkeley, Berkeley, CA, USA(加州大学伯克利分校)
Incentive-Aware Synthetic Control: Accurate Counterfactual Estimation via Incentivized Exploration
具有激励的合成控制:通过激励探索实现准确的反事实估计
Daniel Ngo, Keegan Harris, Anish Agarwal, Vasilis Syrgkanis, Zhiwei Steven Wu
机构
*
J.P. Morgan Chase AI Research(J.P. Morgan Chase人工智能研究)
;
University of California, Berkeley(加州大学伯克利分校)
;
Columbia University(哥伦比亚大学)
;
Stanford University(斯坦福大学)
;
Carnegie Mellon University(卡内基梅隆大学)