CommentsAuthor Accepted Manuscript. Accepted for publication in the Proceedings of the 34th ACM International Conference on Multimedia (ACM MM '26). This author-created manuscript is not the ACM Version of Record
Linked Multi-Modal Data on Russian Domestic and Foreign Policy Speeches
连接多模型数据:俄罗斯国内与对外政策演讲
Daria Blinova, Gayathri Emuru, Rakesh Emuru, Kushagradheer Shridheer Srivastava, Mina Rulis, Sunita Chandrasekaran, Benjamin E. Bagozzi
机构
*
University of Delaware, Department of Political Science & International Relations(德克萨斯大学,政治科学与国际关系系)
;
University of Delaware, Masters of Science in Data Science Program(德克萨斯大学,数据科学硕士项目)
;
University of Delaware, Department of Computer & Information Sciences(德克萨斯大学,计算机与信息科学系)
;
University of Pennsylvania, Department of Political Science(宾夕法尼亚大学,政治科学系)
机构
*
Peking University(北京大学)
;
Kling Team, Kuaishou Technology(快手团队)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Sun Yat-sen University(中山大学)
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
Omnilingual SONAR:跨语言与跨模态句子嵌入,连接大规模多语言文本与语音
Omnilingual SONAR Team, João Maria Janeiro, Pere-Lluís Huguet Cabot, Ioannis Tsiamas, Yen Meng, Vivek Iyer, Guillem Ramírez, Loic Barrault, Belen Alastruey, Xiang "Tony" Cao, Yu-An Chung, Marta R. Costa-Jussa, David Dale, Kevin Heffernan, Jaehyeong Jo, Artyom Kozhevnikov, Alexandre Mourachko, Christophe Ropers, Holger Schwenk, Paul-Ambroise Duquenne
机构
*
College of Computer Science, Nankai University, Tianjin, China(南开大学计算机科学学院,天津,中国)
;
Academy for Advanced Interdisciplinary Studies, Nankai University, Tianjin, China(南开大学先进跨学科研究学院,天津,中国)
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
跨模态一致性引导用于自动回归TTS模型中的鲁棒情绪控制
Yizhou Peng, Yukun Ma, Chong Zhang, Yi-Wen Chao, Chongjia Ni, Bin Ma, Eng Siong Chng
机构
*
Alibaba-NTU Global e-Sustainability CorpLab(阿里-国立大学全球可持续发展公司实验室)
;
Nanyang Technological University(南洋理工大学)
;
College of Computing and Data Science(计算与数据科学学院)
;
Alibaba(阿里)
;
Alibaba Inc.(阿里公司)
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
MeMo: 视觉受损条件下的实时视听目标说话人提取的注意力动量
Junjie Li, Wenxuan Wu, Shuai Wang, Zexu Pan, Kong Aik Lee, Helen Meng, Haizhou Li
机构
*
Department of Electrical and Electronic Engineering, Faculty of Engineering, The Hong Kong Polytechnic University(电子工程系,工程学院,香港理工大学)
;
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong(系统工程与工程管理系,香港中文大学)
;
School of Artificial Intelligence (SAI), The Chinese University of Hong Kong, Shenzhen(人工智能学院(SAI),香港中文大学深圳校区)
;
School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学)
;
Tongyi Lab, Alibaba Group, Singapore(通义实验室,阿里巴巴集团,新加坡)
Legible and Intuitive Multi-modal Robot State and Intent Communication Validated in Online and Real-world Studies
可读且直观的多模态机器人状态与意图通信:在线和真实世界研究验证
Tim Schreiter, Jens V. Rüppel, Andrey Rudenko, Martin Magnusson, Achim J. Lilienthal
机构
*
Chair of Perception for Intelligent Systems, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(慕尼黑工业大学慕尼黑机器人与机器智能研究所智能系统感知教席)
;
Centre for Applied Autonomous Sensor Systems (AASS), Örebro University(厄勒布鲁大学应用自主传感器系统中心)
;
Robotics Institute Germany (RIG)(德国机器人研究所)
Scaling to Multimodal and Multichannel Heart Sound Classification with Synthetic and Augmented Biosignals
利用合成与增强生物信号实现多模态和多通道心音分类的规模化
Milan Marocchi, Matthew Fynn, Kayapanda Mandana, Yue Rong
机构
*
School of Electrical Engineering, Computing, and Mathematical Sciences (EECMS), Faculty of Science and Engineering, Curtin University, Bentley, WA 6102, Australia(电气工程、计算与数学科学学院(EECMS),科学与工程学院, Curtin 大学,Bentley,WA 6102,澳大利亚)