Mindstorms in Natural Language-Based Societies of Mind
自然语言基础的思维社会中的风暴
Mingchen Zhuge, Haozhe Liu, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Abed Al Kader Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li, Guohao Li, Shuming Liu, Jinjie Mai, Piotr Piękos, Aditya Ramesh, Imanol Schlag, Weimin Shi, Aleksandar Stanić, Wenyi Wang, Yuhui Wang, Mengmeng Xu, Deng-Ping Fan, Bernard Ghanem, Jürgen Schmidhuber
Commentspublished in Computational Visual Media Journal (CVMJ); 9 pages in main text + 7 pages of references + 38 pages of appendices, 14 figures in main text + 13 in appendices, 7 tables in appendices
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
鲁棒的音频视觉目标说话人提取与情感感知多注册融合
Zhan Jin, Bang Zeng, Peijun Yang, Jiarong Du, Wei Ju, Yao Tian, Juan Liu, Ming Li
机构
*
School of Computer Science, Wuhan University, Wuhan, China(1 武汉大学计算机学院,武汉,中国)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(2 香港中文大学(深圳)人工智能学院,中国)
;
School of Artificial Intelligence, Wuhan University, Wuhan, China(3 武汉大学人工智能学院,武汉,中国)
;
School of Cyber Science and Engineering, Wuhan University, Wuhan, China(4 武汉大学网络科学与工程学院,武汉,中国)
;
Digital Innovation Research Center, Duke Kunshan University, Kunshan, China(5 香港中文大学(深圳)数字创新研究中心,中国)
;
AI Center, OPPO, Beijing, China(6 OPPO人工智能中心,北京,中国)
Yunsheng Wang, Yuntao Shou, Yilong Tan, Wei Ai, Tao Meng, Keqin Li
机构
*
College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南林业科技大学)
;
Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)
机构
*
Google Mountain View CA USA(谷歌山景城)
;
University of Michigan Ann Arbor MI USA(密歇根大学安 Arbor分校)
;
Columbia University New York NY USA(哥伦比亚大学)
;
Google San Francisco CA USA(谷歌旧金山)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
高效音频-视觉语音分离与离散唇语语义及多尺度全局-局部注意力
Kai Li, Kejun Gao, Xiaolin Hu
机构
*
Department of Computer Science and Technology, Institute for AI, BNRist, Tsinghua University, Beijing(计算机科学与技术系,人工智能研究院,BNRist,清华大学,北京)
;
IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing(IDG/麦克戈文脑研究 institute,清华大学,北京)
;
Chinese Institute for Brain Research (CIBR), Beijing(中国脑科学研究院(CIBR),北京)
Communication-Efficient Multimodal Federated Learning: Joint Modality and Client Selection
高效多模态联邦学习:联合模态与客户端选择
Liangqi Yuan, Dong-Jun Han, Su Wang, Devesh Upadhyay, Christopher G. Brinton
机构
*
School of Electrical and Computer Engineering, Purdue University(普渡大学电气与计算机工程学院)
;
Department of Computer Science and Engineering, Yonsei University(延世大学计算机科学与工程系)
;
School of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程学院)
;
Saab Inc.(Saab公司)
机构
*
School of Mechanical Engineering, University of Science and Technology Beijing(北京科技大学机械工程学院)
;
Laboratory for Computational Sensing and Robotics, Johns Hopkins University(约翰霍普金斯大学计算感知与机器人实验室)
机构
*
Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程系,中国沈阳)
;
Key Laboratory of Intelligent Computing in Medical Image of Ministry of Education, Northeastern University, Shenyang, China(教育部医学图像智能计算重点实验室,东北大学,中国沈阳)
;
National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Shenyang, China(工业智能与系统优化国家级前沿科学中心,中国沈阳)
;
AiShiWeiLai AI Research, China(艾世维来人工智能研究,中国)
;
Amii, University of Alberta, Edmonton, Alberta, Canada(阿尔伯塔大学艾米人工智能研究所,加拿大埃德蒙顿,阿尔伯塔)
专题命中
多模态生成
:cross-modal(abstract);分类 cs.CV
AI总结
本文提出视觉引导的文本解耦框架,通过细粒度语义解耦提升医学图像生成的可控性和生成质量。
Comments10 pages, 7 figures. Currently under review