机构
*
Department of Computer Science and Engineering, Dhaka International University(达卡国际大学计算机科学与工程系)
;
Department of Computer Science and Engineering, Bangladesh University of Engineering and Technology(孟加拉国工程技术大学计算机科学与工程系)
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning
观看合成视频:针对零样本视频字幕生成的视觉合成跨模态表征对齐
Liangyu Fu, Junbo Wang, Yuke Li, Ya Jing, Xuecheng Wu, Zhiyong Wang
机构
*
School of Software, Northwestern Polytechnical University(西北工业大学软件学院)
;
School of Information Science and Technology, Beijing University of Technology(北京工业大学信息科学与技术学院)
;
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
;
School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
用于遥感图像理解的多模态大语言模型:领域特定还是通用?
Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li
机构
*
School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院)
;
Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室)
;
Yuelushan Center for Industrial Innovation(岳麓山工业创新中心)
机构
*
Singapore University of Technology and Design(新加坡科技设计大学)
;
Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局)
;
Nanyang Technological University(南洋理工大学)
;
Chongqing University(重庆大学)
CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction
用于多模态肺癌生存预测的CT-CLIP表示
Sofie Allgöwer, Mikael Johansson, Andreas Hallqvist, Jonas Andersson, Åse Johnsson, Ida Häggström, Jennifer Alvén
机构
*
Chalmers University of Technology(查尔姆斯理工大学)
;
Umeå University(于默奥大学)
;
Sahlgrenska University Hospital(萨尔格伦斯卡大学医院)
;
Sahlgrenska Academy at Gothenburg University(哥德堡大学萨尔格伦斯卡学院)
Qixiang Yin, Huanjin Yao, Yuchen Cai, Jianghao Chen, Ziyi Wang, Min Yang, Fei Su, Zhicheng Zhao
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
ByteDance(字节跳动)
;
USTC(中国科学技术大学)
;
Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室)
;
Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部互动技术与体验系统重点实验室)
;
Zhongguancun Academy(中关村科学城)
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
在自进化大型多模态模型中更加关注视觉标记
Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Aalto University(阿尔托大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林雪平大学)
Million-scale multimodal pollen microscopy with expert-guided foundation models
百万级多模态花粉显微镜图像与专家引导的基础模型
András Biricz, Björn Gedda, Donát Magyar, Antonio Spanu, János Fillinger, Péter Pollner, István Csabai
机构
*
Department of Physics of Complex Systems, ELTE Eötvös Loránd University(ELTE罗兰大学复杂物理系)
;
The Palynological Laboratory at the Swedish Museum of Natural History(瑞典自然历史博物馆孢粉学实验室)
;
National Centre for Public Health and Pharmacy(国家公共卫生与药品中心)
;
INRAE, UR 546 BioSP, Site Agroparc(法国国家农业、食品与环境研究院,UR 546 BioSP,阿格罗帕克园区)
;
National Korányi Institute for Pulmonology(国家科拉尼肺病研究所)
;
Health Data Science and AI Knowledge Centre, Health Services Management Training Centre, Faculty of Health and Public Administration, Semmelweis University(塞梅维什大学健康与公共管理学院卫生服务管理培训中心健康数据科学与人工智能知识中心)
;
Department of Biological Physics, ELTE Eötvös Loránd University(ELTE罗兰大学生物物理系)
专题命中
图文多模态
:multimodal(title,abstract);分类 cs.CV
AI总结
提出百万级多模态花粉显微镜数据集Pollen AI Atlas,结合专家引导的视觉-语言模型生成形态描述,实现跨区域、跨设置的高精度花粉识别与检索。
Comments31 pages, 5 main figures, supplementary information included. Submitted to Scientific Reports