PAD: Phase-Amplitude Decoupling Fusion for Multi-Modal Land Cover Classification
Huiling Zheng, Xian Zhong, Bin Liu, Yi Xiao, Bihan Wen, Xiaofeng Li
机构
*
Sanya Science and Education Innovation Park, Wuhan University of Technology, Sanya 572025, China(武汉理工大学三亚科学教育创新园)
;
School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China(武汉理工大学计算机科学与人工智能学院)
;
Key Laboratory of Ocean Circulation and Waves, Institute of Oceanology, Chinese Academy of Sciences, Qingdao 266071, China(中国科学院海洋循环与波浪重点实验室)
;
Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China(湖北省交通物联网重点实验室)
;
State Key Laboratory of Maritime Technology and Safety, Wuhan University of Technology, Wuhan 430063, China(武汉理工大学航海技术与安全国家重点实验室)
;
College of Oceanography and Ecological Science, Shanghai Ocean University, Shanghai 201306, China(上海海洋大学海洋科学与生态学院)
;
School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China(郑州大学计算机与人工智能学院)
;
Rapid-Rich Object Search Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798(南洋理工大学电子与电气工程学院)
机构
*
Zhejiang Gongshang University(浙江工商大学)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
;
Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波)
;
Meituan Inc.(美团公司)
;
National University of Singapore(新加坡国立大学)
Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji
机构
*
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
Institute of AI for Industries(工业人工智能研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis
Chengsheng Zhang, Linhao Qu, Xiaoyu Liu, Zhijian Song
机构
*
Digital Medical Research Center, School of Basic Medical Science, Fudan University, Shanghai 200032, China(复旦大学基础医学学院数字医学研究中心)
;
Shanghai Key Lab of Medical Image Computing(上海医学图像计算重点实验室)
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Imperial College London(帝国理工学院伦敦分校)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
Xu Li, Fan Lyu
机构
*
Khoury College of Computer Sciences, Northeastern University(东北大学克劳尔计算机科学学院)
;
New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Technical University of Munich(慕尼黑技术大学)
;
Columbia University(哥伦比亚大学)
GalaxAlign: Mimicking Citizen Scientists' Multimodal Guidance for Galaxy Morphology Analysis
Ruoqi Wang, Haitao Wang, Qiong Luo
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
School of Computer Science and Engineering, Sun Yat-Sen University(中山大学计算机科学与工程学院)
M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
Yanshu Li, Yi Cao, Hongyang He, Qisen Cheng, Xiang Fu, Xi Xiao, Tianyang Wang, Ruixiang Tang
机构
*
Brown University(布朗大学)
;
University of Warwick(沃里克大学)
;
Samsung US(三星美国分公司)
;
Boston University(波士顿大学)
;
University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
;
Rutgers University(罗格斯大学)
AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
Wei Li, Xun Gong, Jiao Li, Xiaobin Sun
机构
*
School of Computing and Artificial Intelligence(计算机与人工智能学院)
;
Southwest Jiaotong University(西南交通大学)
;
Department of Gastroenterology(消化内科部)
;
The Third People’s Hospital of Chengdu(成都第三人民医院)
Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection
Md. Mithun Hossain, Md. Shakil Hossain, Sudipto Chaki, M. F. Mridha
机构
*
Department of Computer Science and Engineering, Bangladesh University of Business and Technology(计算机科学与工程系,孟加拉国商业技术大学)
;
Department of Computer Science, American International University-Bangladesh(计算机科学系,美国国际大学-孟加拉国)
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception
Chuang Chen, Xiao Sun, Zhi Liu
机构
*
School of Artificial Intelligence, Anhui University(安徽大学人工智能学院)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥国家综合科学中心人工智能研究所)
;
Anhui Province Key Laboratory of Affective Computing and Advanced Intelligent Machines, School of Computer Science and Information Engineering, Hefei University of Technology(安徽省情感计算与先进智能机器重点实验室,合肥工业大学计算机科学与信息工程学院)
;
Department of Computer and Network Engineering, The University of Electro-Communications(电子通信大学计算机与网络工程系)
SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection
Peican Zhu, Yubo Jing, Le Cheng, Bin Chen, Xiaodong Cui, Lianwei Wu, Keke Tang
机构
*
School of Artificial Intelligence, Optics, and Electronics (iOPEN), Northwestern Polytechnical University(人工智能、光学与电子学院(iOPEN),西北工业大学)
;
School of Computer Science, Northwestern Polytechnical University(计算机学院,西北工业大学)
;
Unit 93212 of People’s Liberation Army of China(中国人民解放军第九三二一二单位)
;
School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学)
;
Cyberspace Institute of Advanced Technology, Guangzhou University(高级技术网络空间研究院,广州大学)
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Tsinghua University(清华大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)