A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation
Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou, José Lezama, Agrim Gupta, Lijun Yu, Lu Jiang, Aren Jansen, Jacob Walker, Krishna Somandepalli
机构
*
Seoul National University(首尔国立大学)
;
Google DeepMind(谷歌DeepMind)
;
Stanford University(斯坦福大学)
;
Carnegie Mellon University(卡内基梅隆大学)
Journal refIn Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024)
NeSyPack: A Neuro-Symbolic Framework for Bimanual Logistics Packing
Bowei Li, Peiqi Yu, Zhenran Tang, Han Zhou, Yifan Sun, Ruixuan Liu, Changliu Liu
机构
*
Carnegie Mellon University(卡内基梅隆大学)
Comments10 pages, 5 figures. Accepted to the RSS 2025 Workshop on Benchmarking Robot Manipulation: Improving Interoperability and Modularity. First Prize in the WBCD competition at ICRA 2025. Equal contribution by Bowei Li and Peiqi Yu
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
Minsu Kim, Jee-weon Jung, Hyeongseop Rha, Soumi Maiti, Siddhant Arora, Xuankai Chang, Shinji Watanabe, Yong Man Ro
机构
*
Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))
;
Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学)
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Nikhil Kandpal, Brian Lester, Colin Raffel, Sebastian Majstorovic, Stella Biderman, Baber Abbasi, Luca Soldaini, Enrico Shippole, A. Feder Cooper, Aviya Skowron, John Kirchenbauer, Shayne Longpre, Lintang Sutawika, Alon Albalak, Zhenlin Xu, Guilherme Penedo, Loubna Ben Allal, Elie Bakouch, John David Pressman, Honglu Fan, Dashiell Stander, Guangyu Song, Aaron Gokaslan, Tom Goldstein, Brian R. Bartoldson, Bhavya Kailkhura, Tyler Murray
机构
*
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
Hugging Face
;
EleutherAI
;
The Allen Institute for Artificial Intelligence(人工智能研究所)
;
Teraflop AI
;
Cornell University(康奈尔大学)
;
University of Maryland, College Park(马里兰大学学院市分校)
;
MIT(麻省理工学院)
;
CMU(卡内基梅隆大学)
;
Lila Sciences
;
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality
Ziyan Wang, Zhicheng Zhang, Fei Fang, Yali Du
机构
*
Department of Informatics, King’s College London, London, United Kingdom(伦敦大学国王学院信息学系)
;
Societal Systems Department, Carnegie Mellon University, Pittsburgh, PA, USA(卡内基梅隆大学社会系统系)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, Xiang Yue
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
M-A-P
;
Nanyang Technological University(南洋理工大学)
;
University of Waterloo(滑铁卢大学)
;
The University of Manchester(曼彻斯特大学)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
Jiewen Hu, Leena Mathur, Paul Pu Liang, Louis-Philippe Morency
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Massachusetts Institute of Technology(麻省理工学院)
CommentsIEEE FG 2025, \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work