Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models
重新审视遥感中的变化视觉问答问题:结构化与原生多模态Qwen模型
Yakoub Bazi, Mohamad M. Al Rahhal, Mansour Zuair, Faroun Mohamed
机构
*
Computer Engineering Department, College of Computer and Information Sciences, King Saud University(计算机工程系,计算机与信息科学学院,沙特国王大学)
;
Applied Computer Science Department, College of Applied Computer Science, King Saud University(应用计算机科学系,应用计算机科学学院,沙特国王大学)
Joycelyn Teo, Rui Cao, Zhenyun Deng, Zifeng Ding, Michael Sejr Schlichtkrull, Andreas Vlachos
机构
*
Defence Science and Technology Agency, Singapore(新加坡国防科学与技术局)
;
University of Cambridge, UK(英国剑桥大学)
;
Queen Mary University of London, UK(英国伦敦大学学院)
PestVL-Net: Enabling Multimodal Pest Learning via Fine-grained Vision-Language Interaction
PestVL-Net: 通过细粒度视觉-语言交互实现多模态害虫学习
Xueheng Li, Tao Hu, Ke Cao, Runsheng Qi, Huixin Zhang, Rui Li, Jie Zhang, Chengjun Xie
机构
*
Institute of Intelligent Machines, Hefei Institutes of Physical Science, Chinese Academy of Sciences(智能机器研究所,合肥物理科学研究院,中国科学院)
;
University of Science and Technology of China(中国科学技术大学)
;
Zhongke Hefei Institute of Technology Innovation Engineering(中科合肥技术创新工程研究院)
Bias-constrained multimodal intelligence for equitable and reliable clinical AI
具有偏见约束的多模态智能用于公平且可靠的临床AI
Cheng Li, Weijian Huang, Jiarun Liu, Hao Yang, Qi Yang, Song Wu, Ye Li, Hairong Zheng, Shanshan Wang
机构
*
Paul C. Lauterbur Research Center for Biomedical Imaging(Paul C. Lauterbur生物医学成像研究中心)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
;
Pengcheng Laboratory(鹏城实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Department of Radiology, Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院放射科)
;
Department of Urology, South China Hospital, Medical School, Shenzhen University(深圳大学南方医院泌尿科)
;
Huawei Technologies Co., Ltd.(华为技术有限公司)
Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models
基于延迟反馈的测试时扰动学习用于视觉-语言-动作模型
Zehua Zang, Xi Wang, Fuchun Sun, Xiao Xu, Lixiang Lium, Jiahuan Zhou, Jiangmeng Li
机构
*
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)
;
National Defense University(国防大学)
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA School of Artificial Intelligence, University of Chinese Academy of Sciences AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University Voyager Research, Didi Chuxing School of Information Science(多模态人工智能系统国家重点实验室,中国科学院自动化所人工智能学院,中国科学院大学,AutoLab,上海交通大学人工智能学院,Voyager Research,滴滴出行,信息科学与技术学院)
Source-Free Domain Adaptation with Vision-Language Prior
无源领域适应与视觉-语言先验
Song Tang, Yunxiang Bai, Wenxin Su, Mao Ye, Jianwei Zhang, Xiatian Zhu
机构
*
Institute of Machine Intelligence(机器智能研究所)
;
University of Shanghai for Science and Technology(上海科技大学)
;
Department of Informatics(信息学院)
;
Universität Hamburg(汉堡大学)
;
European Molecular Biology Laboratory(欧洲分子生物学实验室)
;
School of Computer Science and Engineering(计算机科学与工程学院)
;
University of Electronic Science and Technology of China(电子科技大学)
;
TAMS Group(多模态系统技术组)
;
Surrey Institute for People-Centred Artificial Intelligence(萨里大学以人为本的人工智能研究所)
;
Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音与信号处理中心)
;
University of Surrey(萨里大学)
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络与信息安全学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
AVRT:通过单模态教师模型实现音频-视觉推理迁移
Edson Araujo, Saurabhchand Bhati, M. Jehanzeb Mirza, Brian Kingsbury, Samuel Thomas, Rogerio Feris, James R. Glass, Hilde Kuehne
机构
*
University of Tübingen, Germany(图宾根大学)
;
MIT, Cambridge MA, USA(麻省理工学院)
;
IBM Research, USA(IBM研究院)
;
MIT-IBM Watson AI Lab, USA(麻省理工-IBM沃森人工智能实验室)
;
Tuebingen AI Center, Germany(图宾根人工智能中心)
Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach
多模态情感分析中的缺失模态:一种知识迁移方法
Weide Liu, Huijing Zhan
机构
*
College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算机与数据科学学院)
;
Singapore University of Social Sciences, Singapore(新加坡社会科学研究大学)