Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection
融合后鸟瞰图特征稳定化用于鲁棒多模态3D检测
Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin
机构
*
Department of Electrical Engineering, University of South Florida(佛罗里达州立大学电气工程系)
;
Department of Mechanical Engineering, University of South Florida(佛罗里达州立大学机械工程系)
机构
*
School of Computer, Wuhan University(武汉大学计算机学院)
;
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
school of computer, National University of Defense Technology(国防科技大学计算机学院)
;
Shandong Provincial Key Laboratory of Computer Networks, Shandong Computer Science Center (National Supercomputing Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(山东省计算机网络重点实验室、山东省计算机科学中心(国家超级计算中心济南中心)、齐鲁工业大学(山东省科学院))
;
School of Computer and Artificial Intelligence, Zhengzhou University(郑州大学计算机与人工智能学院)
;
School of Information and Communication Engineering, North University of China(北方大学信息与通信工程学院)
;
National Key Laboratory of Electromagnetic Energy, Naval University of Engineering(电磁能国家重点实验室、海军工程大学)
;
Hexagon AB
Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations
Bo Ma, LuYao Liu, Simon Lau, Chandler Yuan, and XueY Cui, Rosie Zhang
机构
*
Department of Software \& Microelectronics, Peking University, Beijing, China
;
Economic Law School, China University of Political Science
;
Financial Media, Peking University, ChangSha, China
LatXGen: Towards Radiation-Free and Accurate Quantitative Analysis of Sagittal Spinal Alignment Via Cross-Modal Radiographic View Synthesis
Moxin Zhao, Nan Meng, Jason Pui Yin Cheung, Chris Yuk Kwan Tang, Chenxi Yu, Wenting Zhong, Pengyu Lu, Chang Shi, Yipeng Zhuang, Teng Zhang
机构
*
Department of Orthopaedics and Traumatology, The University of Hong Kong(香港大学骨科与创伤学系)
;
Department of Joint Surgery, Shandong Provincial Hospital Affiliated to Shandong First Medical University(山东省第一医科大学附属山东省人民医院骨科)
S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan
机构
*
Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,布拉克大学)
;
Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology(电气与电子工程系,孟加拉国工程与技术大学)
专题命中
多模态训练与对齐
:multi-modal(title);分类 cs.CV、cs.AI
CommentsSubmitted to IEEE Journal of Biomedical and Health Informatics (JBHI). This preprint includes few additional details not present in the journal submission
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
Moran Yanuka, Assaf Ben Kish, Yonatan Bitton, Idan Szpektor, Raja Giryes
机构
*
Tel Aviv University(特拉维夫大学)
;
Google Research(谷歌研究)
专题命中
多模态训练与对齐
:multimodal(title);分类 cs.CV、cs.CL
CommentsAccepted to NAACL 2025
Journal refProceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics, Human Language Technologies, Long Papers, pp. 10497-10518
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
Soham Walimbe, Britty Baby, Vinkle Srivastav, Nicolas Padoy
机构
*
University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心(CNRS),法国国家卫生研究院(INSERM),ICube,UMR7357,斯特拉斯堡)
;
Institute of Image-Guided Surgery, IHU Strasbourg, Strasbourg, France(影像引导手术研究所,斯特拉斯堡IHU,斯特拉斯堡)
机构
*
City University of Hong Kong, Dongguan Campus(香港城市大学东莞校区)
;
The University of Melbourne(墨尔本大学)
;
Tsinghua University(清华大学)
;
Baidu Inc.(百度公司)
;
University of Chinese Academy of Science(中国科学院大学)
;
National University of Singapore(新加坡国立大学)
专题命中
多模态训练与对齐
:multi-modal(title);分类 cs.CV、cs.AI
Comments11 pages, 4 figures, Submitted to ACM MM 2025