arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2406.14098 2024-07-08 cs.CV 79%

HeartBeat: Towards Controllable Echocardiography Video Synthesis with Multimodal Conditions-Guided Diffusion Models

Xinrui Zhou, Yuhao Huang, Wufeng Xue, Haoran Dou, Jun Cheng, Han Zhou, Dong Ni

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03836 2024-07-08 cs.CV cs.LG 79%

ADAPT: Multimodal Learning for Detecting Physiological Changes under Missing Modalities

Julie Mordacq, Leo Milecki, Maria Vakalopoulou, Steve Oudot, Vicky Kalogeiton

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at MIDL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12235 2024-07-02 cs.CV 79%

Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM

Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo, Chuchu Han, Xiaonan Huang, Changxin Gao, Yuehuan Wang, Nong Sang

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15556 2024-06-25 cs.CV 79%

Open-Vocabulary Temporal Action Localization using Multimodal Guidance

Akshita Gupta, Aditya Arora, Sanath Narayan, Salman Khan, Fahad Shahbaz Khan, Graham W. Taylor

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09781 2024-06-17 cs.CV 79%

GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

Yiqi Wu, Xiaodan Hu, Ziming Fu, Siling Zhou, Jiangong Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09711 2024-06-17 cs.CV 79%

AnimalFormer: Multimodal Vision Framework for Behavior-based Precision Livestock Farming

Ahmed Qazi, Taha Razzaq, Asim Iqbal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09076 2024-06-14 cs.CL 79%

3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection

Thye Shan Ng, Feiqi Cao, Soyeon Caren Han

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05496 2024-06-11 cs.CL 79%

Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities

Sai Munikoti, Ian Stewart, Sameera Horawalavithana, Henry Kvinge, Tegan Emerson, Sandra E Thompson, Karl Pazdernik

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments 25 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04575 2024-06-10 cs.LG cs.AI stat.AP stat.ML 79%

Optimization of geological carbon storage operations with multimodal latent dynamic model and deep reinforcement learning

Zhongzheng Wang, Yuntian Chen, Guodong Chen, Dongxiao Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04236 2024-06-07 cs.CV 79%

Understanding Information Storage and Transfer in Multi-modal Large Language Models

Samyadeep Basu, Martin Grayson, Cecily Morrison, Besmira Nushi, Soheil Feizi, Daniela Massiceti

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15813 2024-05-28 cs.CV 79%

From CNNs to Transformers in Multimodal Human Action Recognition: A Survey

Muhammad Bilal Shaikh, Syed Mohammed Shamsul Islam, Douglas Chai, Naveed Akhtar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 5 figures and 3 Tables. To appear in ACM Trans. Multimedia Comput. Commun. Appl.(TOMM) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12708 2024-05-22 cs.CV 79%

Multimodal video analysis for crowd anomaly detection using open access tourism cameras

Alejandro Dionis-Ros, Joan Vila-Francés, Rafael Magdalena-Benedicto, Fernando Mateo, Antonio J. Serrano-López

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11776 2024-05-20 cs.LG cs.CV 79%

3D object quality prediction for Metal Jet Printer with Multimodal thermal encoder

Rachel, Chen, Wenjia Zheng, Sandeep Jalui, Pavan Suri, Jun Zeng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03272 2024-05-07 cs.CV 79%

WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Yuanhan Zhang, Kaichen Zhang, Bo Li, Fanyi Pu, Christopher Arif Setiadharma, Jingkang Yang, Ziwei Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03091 2024-05-07 cs.CV cs.LG 79%

Research on Image Recognition Technology Based on Multimodal Deep Learning

Jinyin Wang, Xingchen Li, Yixuan Jin, Yihao Zhong, Keke Zhang, Chang Zhou

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14471 2024-04-29 cs.CV 79%

Narrative Action Evaluation with Prompt-Guided Multimodal Interaction

Shiyi Zhang, Sule Bai, Guangyi Chen, Lei Chen, Jiwen Lu, Junle Wang, Yansong Tang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05726 2024-04-25 cs.CV 79%

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, Ser-Nam Lim

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2024. Project Page https://boheumd.github.io/MA-LMM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11470 2024-04-18 cs.CV 79%

Exploring Missing Modality in Multimodal Egocentric Datasets

Merey Ramazanova, Alejandro Pardo, Humam Alwassel, Bernard Ghanem

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09276 2024-04-18 cs.CV 79%

Transformer-based Multimodal Change Detection with Multitask Consistency Constraints

Biyuan Liu, Huaixin Chen, Kun Li, Michael Ying Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10429 2024-04-17 cs.AI 79%

MEEL: Multi-Modal Event Evolution Learning

Zhengwei Tao, Zhi Jin, Junqiang Huang, Xiancai Chen, Xiaoying Bai, Haiyan Zhao, Yifan Zhang, Chongyang Tao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07610 2024-04-12 cs.CV 79%

Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval

Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi, Seong Tae Kim

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07540 2024-04-09 cs.LG cs.CV q-bio.QM 79%

Tensor-based Multimodal Learning for Prediction of Pulmonary Arterial Wedge Pressure from Cardiac MRI

Prasun C. Tripathi, Mohammod N. I. Suvon, Lawrence Schobs, Shuo Zhou, Samer Alabed, Andrew J. Swift, Haiping Lu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05698 2024-04-05 cs.CV 79%

Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities

AJ Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo, Victor Gomes, Anelia Angelova

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19811 2024-04-01 cs.CV 79%

X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin, Eric Sauser, Bernt Schiele, Shugao Ma

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12328 2024-03-26 cs.LG cs.AI 79%

A survey on knowledge-enhanced multimodal learning

Maria Lymperaiou, Giorgos Stamou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14598 2024-03-22 cs.CV 79%

PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

Zheng Zhang, Yeyao Ma, Enming Zhang, Xiang Bai

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14082 2024-03-22 cs.CV 79%

EventDance: Unsupervised Source-free Cross-modal Adaptation for Event-based Object Recognition

Xu Zheng, Lin Wang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13507 2024-03-22 cs.CV 79%

FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs

Jinmin Li, Kuofeng Gao, Yang Bai, Jingyun Zhang, Shu-tao Xia, Yisen Wang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16198 2024-03-08 cs.CV cs.LG 79%

Multi-modal learning for geospatial vegetation forecasting

Vitus Benson, Claire Robin, Christian Requena-Mesa, Lazaro Alonso, Nuno Carvalhais, José Cortés, Zhihan Gao, Nora Linscheid, Mélanie Weynants, Markus Reichstein

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2024. We provide open source code and pre-trained weights to reproduce our experimental results under https://github.com/vitusbenson/greenearthnet

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02483 2024-03-07 cs.CV 79%

EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model

Guozhang Li, Xinpeng Ding, De Cheng, Jie Li, Nannan Wang, Xinbo Gao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏