arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2401.16076 2024-01-31 cs.CV cs.MM 81%

Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas

Carlo Bretti, Pascal Mettes, Hendrik Vincent Koops, Daan Odijk, Nanne van Noord

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments MMM24

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04965 2024-01-22 cs.CL cs.AI cs.LG 81%

MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks

Jingyuan Qi, Minqian Liu, Ying Shen, Zhiyang Xu, Lifu Huang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by AAAI 2024. 11 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09861 2024-01-19 cs.CV cs.AI 81%

Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models

Li Sun, Liuan Wang, Jun Sun, Takayuki Okatani

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08581 2024-01-18 cs.CV cs.AI cs.LG 81%

Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision

Yi Cao, Swetava Ganguli, Vipul Pandey

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Extended abstract accepted for presentation at BayLearn 2023. 3 pages, 7 figures. Abstract based on IEEE IGARSS 2023 research track paper: arXiv:2304.13143

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06194 2024-01-15 cs.LG cs.AI cs.CL 81%

CrisisKAN: Knowledge-infused and Explainable Multimodal Attention Network for Crisis Event Classification

Shubham Gupta, Nandini Saini, Suman Kundu, Debasis Das

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10763 2024-01-11 cs.CV cs.AI cs.LG eess.IV 81%

Actor-agnostic Multi-label Action Recognition with Multi-modal Query

Anindya Mondal, Sauradip Nag, Joaquin M Prada, Xiatian Zhu, Anjan Dutta

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Published at the 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Paris, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03177 2024-01-09 cs.CV cs.CL 81%

Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks

Qian Li, Lixin Su, Jiashu Zhao, Long Xia, Hengyi Cai, Suqi Cheng, Hengzhu Tang, Junfeng Wang, Dawei Yin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03040 2024-01-09 cs.LG cs.AI cs.CV cs.DC 81%

AccidentGPT: Large Multi-Modal Foundation Model for Traffic Accident Analysis

Kebin Wu, Wenbin Li, Xiaofei Xiao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17117 2023-12-29 cs.CV cs.AI 81%

Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long Videos

Houlun Chen, Xin Wang, Hong Chen, Zihan Song, Jia Jia, Wenwu Zhu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01575 2023-12-05 cs.CL cs.CV 81%

A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video

Keito Kudo, Haruki Nagasawa, Jun Suzuki, Nobuyuki Shimizu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10899 2023-11-22 cs.CV cs.CL cs.LG 81%

Extraction and Summarization of Explicit Video Content using Multi-Modal Deep Learning

Shaunak Joshi, Raghav Gaggar

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11845 2023-09-27 cs.SD cs.LG cs.MM eess.AS 81%

TMac: Temporal Multi-Modal Graph Learning for Acoustic Event Classification

Meng Liu, Ke Liang, Dayu Hu, Hao Yu, Yue Liu, Lingyuan Meng, Wenxuan Tu, Sihang Zhou, Xinwang Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.MM、eess.AS

Comments This work has been accepted by ACM MM 2023 for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05880 2023-09-11 cs.CL cs.AI 81%

TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real World

Hongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu, Jingyuan Wen, Yixin Xu, Di Hu, Ruihua Song, Wayne Xin Zhao, Qin Jin, Zhiwu Lu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14395 2023-08-29 cs.MM cs.CV 81%

UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery Localization

Rui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu, Yang Zhou, Qiang Zeng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 11 pages, 8 figures, 66 references. This paper has been accepted for ACM MM 2023

Journal ref Proceedings of the 31st ACM International Conference on Multimedia (MM '23), October 29-November 3, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14160 2023-08-29 cs.CV cs.AI 81%

A Unified Transformer-based Network for multimodal Emotion Recognition

Kamran Ali, Charles E. Hughes

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13309 2023-08-21 cs.CV cs.AI 81%

History Aware Multimodal Transformer for Vision-and-Language Navigation

Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan Laptev

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in NeurIPS 2021; project page at https://cshizhe.github.io/projects/vln_hamt.html; corrected a typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00732 2023-08-14 cs.IR cs.AI cs.CL 81%

Kuaipedia: a Large-scale Multi-modal Short-video Encyclopedia

Haojie Pan, Zepeng Zhai, Yuzhou Zhang, Ruiji Fu, Ming Liu, Yangqiu Song, Zhongyuan Wang, Bing Qin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02469 2023-08-01 cs.CV cs.CL 81%

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06385 2023-07-20 cs.CV cs.LG cs.SD eess.AS 81%

Temporal Label-Refinement for Weakly-Supervised Audio-Visual Event Localization

Kalyan Ramakrishnan

专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00701 2023-06-26 astro-ph.IM astro-ph.CO cs.AI cs.CV cs.LG gr-qc 81%

DeepGraviLens: a Multi-Modal Architecture for Classifying Gravitational Lensing Data

Nicolò Oreste Pinciroli Vago, Piero Fraternali

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this article is published in Neural Computing and Applications, and is available online at https://doi.org/10.1007/s00521-023-08766-9

Journal ref Neural Comput & Applic (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12647 2023-06-08 cs.CV cs.AI 81%

Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering

Yang Liu, Guanbin Li, Liang Lin

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 17 pages, 9 figures. This work has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). The datasets, code and models are available at https://github.com/HCPLab-SYSU/CMCIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00355 2023-05-05 cs.CV cs.AI 81%

MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer

Yifang Xu, Yunzhuo Sun, Yang Li, Yilei Shi, Xiaoxiang Zhu, Sidan Du

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01476 2023-05-03 cs.SD cs.MM eess.AS 81%

Deep Learning Based Multimodal with Two-phase Training Strategy for Daily Life Video Classification

Lam Pham, Trang Le, Cam Le, Dat Ngo, Weissenfeld Axel, Alexander Schindler

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07775 2023-04-28 cs.CV cs.MM 81%

Robust Cross-Modal Knowledge Distillation for Unconstrained Videos

Wenke Xia, Xingjian Li, Andong Deng, Haoyi Xiong, Dejing Dou, Di Hu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14369 2023-03-28 cs.CV cs.MM 81%

Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu, Xiangyang Ji, Li Yuan, Jie Chen

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments CVPR 2023 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11381 2023-03-22 cs.CV cs.CL cs.LG 81%

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, Lijuan Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10777 2023-01-12 cs.CV cs.AI cs.LG cs.RO cs.SY eess.SY 81%

ParkPredict+: Multimodal Intent and Motion Prediction for Vehicles in Parking Lots with CNN and Transformer

Xu Shen, Matthew Lacayo, Nidhir Guggilla, Francesco Borrelli

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published at IEEE ITSC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14714 2022-10-27 cs.CV cs.MM 81%

TAMFormer: Multi-Modal Transformer with Learned Attention Mask for Early Intent Prediction

Nada Osman, Guglielmo Camporese, Lamberto Ballan

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02494 2022-10-20 cs.CV cs.CL 81%

Progressive Video Summarization via Multimodal Self-supervised Learning

Li Haopeng, Ke Qiuhong, Gong Mingming, Tom Drummond

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12694 2022-09-27 cs.CV cs.AI 81%

Multi-modal Video Chapter Generation

Xiao Cao, Zitan Chen, Canyu Le, Lei Meng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏