CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction
用于多模态肺癌生存预测的CT-CLIP表示
Sofie Allgöwer, Mikael Johansson, Andreas Hallqvist, Jonas Andersson, Åse Johnsson, Ida Häggström, Jennifer Alvén
机构
*
Chalmers University of Technology(查尔姆斯理工大学)
;
Umeå University(于默奥大学)
;
Sahlgrenska University Hospital(萨尔格伦斯卡大学医院)
;
Sahlgrenska Academy at Gothenburg University(哥德堡大学萨尔格伦斯卡学院)
Qixiang Yin, Huanjin Yao, Yuchen Cai, Jianghao Chen, Ziyi Wang, Min Yang, Fei Su, Zhicheng Zhao
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
ByteDance(字节跳动)
;
USTC(中国科学技术大学)
;
Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室)
;
Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部互动技术与体验系统重点实验室)
;
Zhongguancun Academy(中关村科学城)
Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning
按需查看:多模态推理中视觉证据获取的认知调度框架
Yang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun, Yidong Chen, Jiayi Ji, Qian Chen, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception(多媒体可信感知实验室)
;
Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China(教育部高效计算实验室,厦门大学,361005,中国)
;
Sino-Russian ResearchCenter for Digital Economy(中俄数字经济研究中心)
;
Institute of Artificial Intelligence, Xiamen University, China(人工智能研究院,厦门大学,中国)
;
School of Informatics, Xiamen University, China(信息学院,厦门大学,中国)
;
School of Information Engineering, Xiamen Ocean Vocational College, Xiamen 361102, China(信息工程学院,厦门海洋职业技术学院,厦门361102,中国)
Learning from Acquisition: Metadata-driven Multimodal Pre-training for Cardiac MRI
从采集信息中学习:基于元数据驱动的心脏MRI多模态预训练
Xueyi Fu, Liwei Hu, Zi Wang, Guang Yang
机构
*
Department of Surgery & Cancer(外科与癌症系)
;
Bioengineering Department and Imperial-X(生物工程系和Imperial-X)
;
National Heart and Lung Institute(国家心脏和肺研究所)
;
Cardiovascular Research Centre(心血管研究中心)
;
School of Biomedical Engineering & Imaging Sciences(生物医学工程与成像科学学院)
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
在自进化大型多模态模型中更加关注视觉标记
Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Aalto University(阿尔托大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林雪平大学)
Million-scale multimodal pollen microscopy with expert-guided foundation models
百万级多模态花粉显微镜图像与专家引导的基础模型
András Biricz, Björn Gedda, Donát Magyar, Antonio Spanu, János Fillinger, Péter Pollner, István Csabai
机构
*
Department of Physics of Complex Systems, ELTE Eötvös Loránd University(ELTE罗兰大学复杂物理系)
;
The Palynological Laboratory at the Swedish Museum of Natural History(瑞典自然历史博物馆孢粉学实验室)
;
National Centre for Public Health and Pharmacy(国家公共卫生与药品中心)
;
INRAE, UR 546 BioSP, Site Agroparc(法国国家农业、食品与环境研究院,UR 546 BioSP,阿格罗帕克园区)
;
National Korányi Institute for Pulmonology(国家科拉尼肺病研究所)
;
Health Data Science and AI Knowledge Centre, Health Services Management Training Centre, Faculty of Health and Public Administration, Semmelweis University(塞梅维什大学健康与公共管理学院卫生服务管理培训中心健康数据科学与人工智能知识中心)
;
Department of Biological Physics, ELTE Eötvös Loránd University(ELTE罗兰大学生物物理系)
专题命中
图文多模态
:multimodal(title,abstract);分类 cs.CV
AI总结
提出百万级多模态花粉显微镜数据集Pollen AI Atlas,结合专家引导的视觉-语言模型生成形态描述,实现跨区域、跨设置的高精度花粉识别与检索。
Comments31 pages, 5 main figures, supplementary information included. Submitted to Scientific Reports
Text-guided Feature Disentanglement for Cross-modal Gait Recognition
文本引导的跨模态步态识别特征解耦
Zhiyang Lu, Ming Cheng
机构
*
Fujian Key Laboratory of Urban Intelligent Sensing and Computing, Xiamen University(福建城市智能感知与计算重点实验室,厦门大学)
;
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,中华人民共和国教育部,厦门大学)
SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning
SVSR:一种用于多模态推理的自我验证与自我修正范式
Zhe Qian, Nianbing Su, Zhonghua Wang, Hebei Li, Zhongxing Xu, Yueying Li, Fei Luo, Zhuohan Ouyang, Yanbiao Ma
机构
*
South China Agricultural University(华南农业大学)
;
University of Glasgow(格拉斯哥大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Monash University(莫纳什大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Defense Technology(国防科技大学)
;
Renmin University of China(中国人民大学)
;
South China Normal University(华南师范大学)
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
;
Guangzhou University(广州大学)
;
Wuhan University(武汉大学)
;
Nanjing University(南京大学)
M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models
M3DocDep: 多模态、多页、多文档依赖分块方法基于大视觉-语言模型
Joongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
机构
*
Human-inspired AI Research, Korea University(韩国大学人智AI研究所)
;
Computer Science and Engineering, Konkuk University(konkuk大学计算机科学与工程系)
;
Department of Computer Science and Engineering, Korea University(韩国大学计算机科学与工程系)
机构
*
Peking University Shanghai AI Laboratory(北京大学上海人工智能实验室)
;
Nanjing University(南京大学)
;
Peking University(北京大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Zhongguancun Academy Beijing Key Laboratory of Data Intelligence and Security(中关村北京数据智能与安全重点实验室)