机构
*
Ministry of Education Key Laboratory of Intelligent Networks and Network Security(教育部智能网络与网络安全重点实验室)
;
Centre for Frontier AI Research, Institute of High Performance Computing, Agency for Science, Technology and Research(前沿人工智能研究中心,高性能计算研究所,科技研究局)
机构
*
Kyoto University(京都大学)
;
Center for Information and Neural Networks(信息与神经网络中心)
;
National Institute of Information and Communications Technology(国立信息通信技术研究所)
;
The University of Osaka(大阪大学)
Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
通过教师引导的双路径实现语义噪声削减
Linge Wang, Yingying Chen, Bingke Zhu, Lu Zhou, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Objecteye Inc.(北京眼神科技有限公司)
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
让 MLLMs 失明:MLLM 内容审核中的对抗走私攻击
Zhiheng Li, Zongyang Ma, Yuntong Pan, Ziqi Zhang, Xiaolei Lv, Bo Li, Jun Gao, Jianing Zhang, Chunfeng Yuan, Bing Li, Weiming Hu
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(北京市多模态信息超智能安全重点实验室)
;
Hellogroup
;
University of Washington(华盛顿大学)
;
Jilin University(吉林大学)
;
ShanghaiTech University(上海科技大学)
An Attention-Assisted Multi-Modal Data Fusion Model for Real-Time Estimation of Underwater Sound Velocity
一种结合注意力机制的多模态数据融合模型用于实时估计水下声速
Pengfei Wu, Wei Huang, Yujie Shi, Hao Zhang
机构
*
Faculty of Information Science and Engineering, Ocean University of China(中国海洋大学信息科学与工程学部)
;
School of Environmental Science and Engineering, Ocean University of China(中国海洋大学环境科学与工程学院)
机构
*
Keio University(庆应义塾大学)
;
National Institute of Informatics(国立信息学研究所)
;
National Institute of Informatics Research and Development Center for Large Language Models(国立信息学研究所大型语言模型研发中心)
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
通过一次剪辑检索增强多模态大语言模型的长视频理解
Tao Chen, Shaobo Ju, Qiong Wu, Chenxin Fang, Kun Zhang, Jun Peng, Hui Li, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China(多媒体可信感知与高效计算教育部重点实验室)
;
Xiamen University(厦门大学)
Exploring Temporal Representation in Neural Processes for Multimodal Action Prediction
探索神经过程在多模态动作预测中的时间表示
Marco Gabriele Fedozzi, Yukie Nagai, Francesco Rea, Alessandra Sciutti
机构
*
DIBRIS Department, University of Genoa(热那亚大学DIBRIS系)
;
CONTACT Unit, Italian Institute of Technology(意大利理工学院CONTACT单元)
;
International Research Center for Neurointelligence, The University of Tokyo(东京大学国际神经智能研究中心)