arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4726 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4726 篇

2410.16592 2024-10-23 cs.LG cs.CL cs.CY 74%

ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding

Andrew Kan, Christopher Kan, Zaid Nabulsi

专题命中 视频多模态 :multi-modal(title);分类 cs.CL

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12420 2024-10-03 cs.CL cs.LG 74%

MMUTF: Multimodal Multimedia Event Argument Extraction with Unified Template Filling

Philipp Seeberger, Dominik Wagner, Korbinian Riedhammer

专题命中 视频多模态 :multimodal(title);分类 cs.CL

Comments Accepted to Findings of EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00017 2024-10-02 cs.CV eess.SP 74%

Multimodal Power Outage Prediction for Rapid Disaster Response and Resource Allocation

Alejandro Aparcedo, Christian Lopez, Abhinav Kotta, Mengjie Li

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 7 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02690 2024-09-05 cs.SI cs.CL 74%

Detecting Calls to Action in Multimodal Content: Analysis of the 2021 German Federal Election Campaign on Instagram

Michael Achmann-Denkler, Jakob Fehle, Mario Haim, Christian Wolff

专题命中 视频多模态 :multimodal(title);分类 cs.CL

Comments Accepted Archival Paper for the CPSS Workshop at KONVENS 2024. Camera Ready Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14930 2024-08-29 cs.CV 74%

CMTA: Cross-Modal Temporal Alignment for Event-guided Video Deblurring

Taewoo Kim, Hoonhee Cho, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments Accepted in ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15377 2024-08-15 cs.CV 74%

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Chenting Wang, Guo Chen, Baoqi Pei, Ziang Yan, Rongkun Zheng, Jilan Xu, Zun Wang, Yansong Shi, Tianxiang Jiang, Songze Li, Hongjie Zhang, Yifei Huang, Yu Qiao, Yali Wang, Limin Wang

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments a technical report about video understanding (accepted to ECCV2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15921 2024-07-02 cs.CV 74%

PUDD: Towards Robust Multi-modal Prototype-based Deepfake Detection

Alvaro Lopez Pellcier, Yi Li, Plamen Angelov

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments CVPR2024

Journal ref CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09585 2024-06-12 cs.CV 74%

StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential Equation

Yining Shi, Kun Jiang, Ke Wang, Jiusi Li, Yunlong Wang, Mengmeng Yang, Diange Yang

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments cvpr2024 poster (highlight), code at https://github.com/synsin0/StreamingFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00822 2024-06-12 cs.AI cs.HC cs.RO 74%

Open-Ended Multi-Modal Relational Reasoning for Video Question Answering

Haozheng Luo, Ruiyang Qin, Chenwei Xu, Guo Ye, Zening Luo

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10825 2024-03-19 cs.CV 74%

Affective Behaviour Analysis via Integrating Multi-Modal Knowledge

Wei Zhang, Feng Qiu, Chen Liu, Lincheng Li, Heming Du, Tiancheng Guo, Xin Yu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08369 2024-02-14 cs.AI 74%

One-shot Imitation in a Non-Stationary Environment via Multi-Modal Skill

Sangwoo Shin, Daehee Lee, Minjong Yoo, Woo Kyung Kim, Honguk Woo

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Comments ICML-2023 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12419 2024-01-24 cs.CV 74%

Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)

Shih-Han Chou, Matthew Kowal, Yasmin Niknam, Diana Moyano, Shayaan Mehdi, Richard Pito, Cheng Zhang, Ian Knopke, Sedef Akinli Kocak, Leonid Sigal, Yalda Mohsenzadeh

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14634 2023-12-25 cs.RO cs.AI 74%

Mining multi-modal communication patterns in interaction with explainable and non-explainable robots

Suna Bensch, Amanda Eriksson

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

Journal ref IEEE RO-MAN 2023, 32nd IEEE International conference on Robot and Human Interactive Communication; Workshop Human-Robot Interaction for Explainability in Robotics, Busan, Korea, August 28-31, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01910 2023-11-27 cs.CV 74%

Multimodal Generation of Novel Action Appearances for Synthetic-to-Real Recognition of Activities of Daily Living

Zdravko Marinov, David Schneider, Alina Roitberg, Rainer Stiefelhagen

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 8 pages, 7 figures, to be published in IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05494 2023-11-10 cs.CV cs.RO 74%

Object-centric Cross-modal Feature Distillation for Event-based Object Detection

Lei Li, Alexander Liniger, Mario Millhaeusler, Vagia Tsiminaki, Yuanyou Li, Dengxin Dai

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.00796 2023-09-26 cs.CV 74%

Multimodal Visual Concept Learning with Weakly Supervised Techniques

Giorgos Bouritsas, Petros Koutras, Athanasia Zlatintsi, Petros Maragos

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments CVPR 2018

Journal ref Proc. IEEE/CVF Conf. Comp. Vis. Patt. Rec. (CVPR) pp. 4914 - 4923 (2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09306 2023-07-19 cs.CV cs.LG cs.RO 74%

EigenTrajectory: Low-Rank Descriptors for Multi-Modal Trajectory Forecasting

Inhwan Bae, Jean Oh, Hae-Gon Jeon

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08063 2023-03-21 cs.CV 74%

MINOTAUR: Multi-task Video Grounding From Multimodal Queries

Raghav Goyal, Effrosyni Mavroudi, Xitong Yang, Sainbayar Sukhbaatar, Leonid Sigal, Matt Feiszli, Lorenzo Torresani, Du Tran

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 22 pages, 8 figures and 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02530 2023-03-07 cs.CV cs.LG eess.SP stat.AP stat.ML 74%

MSED: a multi-modal sleep event detection model for clinical sleep analysis

Alexander Neergaard Olesen, Poul Jennum, Emmanuel Mignot, Helge B. D. Sorensen

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments 10 pages, 4 figures. Accepted for publication in IEEE Transactions on Biomedical Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13606 2023-02-01 cs.CV 74%

Multi-video Moment Ranking with Multimodal Clue

Danyang Hou, Liang Pang, Yanyan Lan, Huawei Shen, Xueqi Cheng

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 9 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12460 2022-10-25 cs.CL 74%

Collaborative Reasoning on Multi-Modal Semantic Graphs for Video-Grounded Dialogue Generation

Xueliang Zhao, Yuxuan Wang, Chongyang Tao, Chenshuo Wang, Dongyan Zhao

专题命中 视频多模态 :multi-modal(title);分类 cs.CL

Comments To appear at EMNLP 2022 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08113 2022-10-18 cs.CV 74%

Instance Segmentation with Cross-Modal Consistency

Alex Zihao Zhu, Vincent Casser, Reza Mahjourian, Henrik Kretzschmar, Sören Pirk

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments 8 pages, 9 figures, 5 tables. Presented at IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04246 2022-08-09 cs.CV cs.LG 74%

Snowpack Estimation in Key Mountainous Water Basins from Openly-Available, Multimodal Data Sources

Malachy Moran, Kayla Woputz, Derrick Hee, Manuela Girotto, Paolo D'Odorico, Ritwik Gupta, Daniel Feldman, Puya Vahabi, Alberto Todeschini, Colorado J Reed

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments Accepted Oral Presentation at CVPR 2022 MultiEarth

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.13274 2022-07-15 cs.LG cs.AI 74%

Evaluating Multimodal Interactive Agents

Josh Abramson, Arun Ahuja, Federico Carnevale, Petko Georgiev, Alex Goldin, Alden Hung, Jessica Landon, Timothy Lillicrap, Alistair Muldal, Blake Richards, Adam Santoro, Tamara von Glehn, Greg Wayne, Nathaniel Wong, Chen Yan

专题命中 视频多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.07096 2022-05-17 cs.CV cs.RO 74%

Multi-modal curb detection and filtering

Sandipan Das, Navid Mahabadi, Saikat Chatterjee, Maurice Fallon

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15829 2022-03-31 cs.CV 74%

An EEG-Based Multi-Modal Emotion Database with Both Posed and Authentic Facial Actions for Emotion Analysis

Xiaotian Li, Xiang Zhang, Huiyuan Yang, Wenna Duan, Weiying Dai, Lijun Yin

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Journal ref FG2021(long Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07086 2022-03-15 cs.CV 74%

MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization

Alexander Kunitsyn, Maksim Kalashnikov, Maksim Dzabraev, Andrei Ivaniuta

专题命中 视频多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.02252 2022-01-25 cs.CV 74%

Discourse Parsing in Videos: A Multi-modal Appraoch

Arjun R. Akula, Song-Chun Zhu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Accepted in CVPR 2019 Workshop on Language and Vision (Oral Presentation)

Journal ref CVPR 2019 Workshop on Language and Vision (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12341 2021-11-25 cs.CV 74%

EvDistill: Asynchronous Events to End-task Learning via Bidirectional Reconstruction-guided Cross-modal Knowledge Distillation

Lin Wang, Yujeong Chae, Sung-Hoon Yoon, Tae-Kyun Kim, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

Comments CVPR 2021 (updated references in this version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12083 2021-11-24 cs.RO cs.CV cs.LG 74%

VISTA 2.0: An Open, Data-driven Simulator for Multimodal Sensing and Policy Learning for Autonomous Vehicles

Alexander Amini, Tsun-Hsuan Wang, Igor Gilitschenski, Wilko Schwarting, Zhijian Liu, Song Han, Sertac Karaman, Daniela Rus

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments First two authors contributed equally. Code and project website is available here: https://vista.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏