arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2409.02145 2024-09-05 cs.LG cs.AI 83%

A Multimodal Object-level Contrast Learning Method for Cancer Survival Risk Prediction

Zekang Yang, Hong Liu, Xiangdong Wang

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05264 2024-08-13 cs.CR cs.CV 83%

Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security

Yihe Fan, Yuxin Cao, Ziyu Zhao, Ziyao Liu, Shaofeng Li

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 8 pages, 1 figure. Accepted to 2024 IEEE International Conference on Systems, Man, and Cybernetics

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00749 2024-07-25 cs.CV 83%

Multimodal Query-guided Object Localization

Aditay Tripathi, Rajath R Dani, Anand Mishra, Anirban Chakraborty

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to MMTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00369 2024-07-02 cs.CL 83%

How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models

Jaeyoung Lee, Ximing Lu, Jack Hessel, Faeze Brahman, Youngjae Yu, Yonatan Bisk, Yejin Choi, Saadia Gabriel

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02774 2024-05-22 eess.IV cs.CV physics.med-ph 83%

Spatial and Modal Optimal Transport for Fast Cross-Modal MRI Reconstruction

Qi Wang, Zhijie Wen, Jun Shi, Qian Wang, Dinggang Shen, Shihui Ying

专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14250 2024-01-26 cs.CV 83%

JUMP: A joint multimodal registration pipeline for neuroimaging with minimal preprocessing

Adria Casamitjana, Juan Eugenio Iglesias, Raul Tudela, Aida Ninerola-Baizan, Roser Sala-Llonch

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10616 2023-12-19 cs.CV 83%

DistilVPR: Cross-Modal Knowledge Distillation for Visual Place Recognition

Sijie Wang, Rui She, Qiyu Kang, Xingchao Jian, Kai Zhao, Yang Song, Wee Peng Tay

专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08343 2023-12-19 eess.IV cs.CV q-bio.QM 83%

Enhancing CT Image synthesis from multi-modal MRI data based on a multi-task neural network framework

Zhuoyao Xin, Christopher Wu, Dong Liu, Chunming Gu, Jia Guo, Jun Hua

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 4 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01605 2023-12-05 cs.CV cs.LG 83%

TextAug: Test time Text Augmentation for Multimodal Person Re-identification

Mulham Fawakherji, Eduard Vazquez, Pasquale Giampa, Binod Bhattarai

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12320 2023-11-22 cs.AI 83%

A Survey on Multimodal Large Language Models for Autonomous Driving

Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, Tianren Gao, Erlong Li, Kun Tang, Zhipeng Cao, Tong Zhou, Ao Liu, Xinrui Yan, Shuqi Mei, Jianguo Cao, Ziran Wang, Chao Zheng

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09067 2023-09-06 cs.CV 83%

2nd Place Winning Solution for the CVPR2023 Visual Anomaly and Novelty Detection Challenge: Multimodal Prompting for Data-centric Anomaly Detection

Yunkang Cao, Xiaohao Xu, Chen Sun, Yuqi Cheng, Liang Gao, Weiming Shen

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments The first two author contribute equally. CVPR workshop challenge report. arXiv admin note: substantial text overlap with arXiv:2305.10724

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09931 2023-07-20 cs.CV cs.LG eess.IV 83%

DISA: DIfferentiable Similarity Approximation for Universal Multimodal Registration

Matteo Ronchetti, Wolfgang Wein, Nassir Navab, Oliver Zettinig, Raphael Prevost

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments This preprint was submitted to MICCAI 2023. The Version of Record of this contribution will be published in Springer LNCS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04062 2023-04-11 cs.LG cs.AI 83%

Predicting multiple sclerosis disease severity with multimodal deep neural networks

Kai Zhang, John A. Lincoln, Xiaoqian Jiang, Elmer V. Bernstam, Shayan Shams

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10457 2023-03-21 cs.CV cs.RO 83%

Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation

Haozhi Cao, Yuecong Xu, Jianfei Yang, Pengyu Yin, Shenghai Yuan, Lihua Xie

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 15 pages, 6 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03299 2023-01-25 cs.LG cs.AI 83%

Multimodal learning with graphs

Yasha Ektefaie, George Dasoulas, Ayush Noori, Maha Farhat, Marinka Zitnik

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 27 pages, 5 figures, 2 boxes

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.07253 2022-06-23 cs.CV 83%

Cross-modal Learning for Domain Adaptation in 3D Semantic Segmentation

Maximilian Jaritz, Tuan-Hung Vu, Raoul de Charette, Émilie Wirbel, Patrick Pérez

专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments TPAMI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07595 2022-06-16 eess.IV cs.CV cs.LG 83%

BIO-CXRNET: A Robust Multimodal Stacking Machine Learning Technique for Mortality Risk Prediction of COVID-19 Patients using Chest X-Ray Images and Clinical Data

Tawsifur Rahman, Muhammad E. H. Chowdhury, Amith Khandakar, Zaid Bin Mahbub, Md Sakib Abrar Hossain, Abraham Alhatou, Eynas Abdalla, Sreekumar Muthiyal, Khandaker Farzana Islam, Saad Bin Abul Kashem, Muhammad Salman Khan, Susu M. Zughaier, Maqsud Hossain

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 25 pages, 8 Tables, 10 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04072 2022-05-10 cs.CV 83%

Beyond Bounding Box: Multimodal Knowledge Learning for Object Detection

Weixin Feng, Xingyuan Bu, Chenchen Zhang, Xubin Li

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Submitted to CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.09666 2021-12-06 cs.CL 83%

Multimodal End-to-End Sparse Model for Emotion Recognition

Wenliang Dai, Samuel Cahyawijaya, Zihan Liu, Pascale Fung

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.04086 2021-11-09 cs.LG cs.AI cs.DB cs.MA 83%

Meta Cross-Modal Hashing on Long-Tailed Data

Runmin Wang, Guoxian Yu, Carlotta Domeniconi, Xiangliang Zhang

专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06423 2020-11-13 cs.DB cs.AI 83%

Turning Transport Data to Comply with EU Standards while Enabling a Multimodal Transport Knowledge Graph

Mario Scrocca, Marco Comerio, Alessio Carenini, Irene Celino

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments International Semantic Web Conference (ISWC 2020) - In Use Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.13806 2020-07-14 cs.CV 83%

X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data

Danfeng Hong, Naoto Yokoya, Gui-Song Xia, Jocelyn Chanussot, Xiao Xiang Zhu

专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing,2020,167:12-23

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.02024 2019-12-05 cs.CV 83%

Template co-updating in multi-modal human activity recognition systems

Annalisa Franco, Antonio Magnani, Dario Maio

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.13733 2019-10-01 cs.MM cs.LG 83%

Cross-Modal Subspace Learning with Scheduled Adaptive Margin Constraints

David Semedo, João Magalhães

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.MM

Comments To appear in ACM MM 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.10802 2018-09-05 cs.CL 83%

The MeMAD Submission to the WMT18 Multimodal Translation Task

Stig-Arne Grönroos, Benoit Huet, Mikko Kurimo, Jorma Laaksonen, Bernard Merialdo, Phu Pham, Mats Sjöberg, Umut Sulubacak, Jörg Tiedemann, Raphael Troncy, Raúl Vázquez

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments To appear in WMT18

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.03847 2018-06-12 cs.RO cs.CL 83%

A Multimodal Classifier Generative Adversarial Network for Carry and Place Tasks from Ambiguous Language Instructions

Aly Magassouba, Komei Sugiura, Hisashi Kawai

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments 9 pages, 7 figures, accepted for IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.11730 2018-05-31 stat.ML cs.AI cs.LG 83%

Learn to Combine Modalities in Multimodal Deep Learning

Kuan Liu, Yanen Li, Ning Xu, Prem Natarajan

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1602.02720 2016-11-03 cs.CV 83%

Multimodal Remote Sensing Image Registration with Accuracy Estimation at Local and Global Scales

M. L. Uss, B. Vozel, V. V. Lukin, K. Chehdi

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 48 pages, 8 figures, 5 tables, 51 references Revised arguments in sections 2 and 3. Additional test cases added in Section 4; comparison with the state-of-the-art improved. References added. Conclusions unchanged. Proofread

详情

展开后加载摘要…

URL PDF HTML 收藏
1001.0443 2010-01-14 cs.MM 83%

Discovering Knowledge from Multi-modal Lecture Recordings

Rajkumar Kannan, Christian Guetl

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.MM

Comments First International Conference on Data Engineering and Management 2008, India

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27938 2026-07-31 cs.HC 新提交 82%

VizPilot: Automated Onboarding for SVG-based Composite Visualizations using Multimodal LLMs

VizPilot:基于多模态大语言模型的SVG复合可视化自动引导系统

Nishaanthini Gnanavel, Yong Wang

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 VizPilot是基于多模态大语言模型的SVG复合可视化自动引导工具,通过双模块实现自动生成交互式引导,经评估可降低创作工作量并减轻用户认知负荷,提升复合可视化可用性。

Comments Accepted by IEEE VIS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏