arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2410.16116 2024-10-22 astro-ph.SR astro-ph.IM cs.AI cs.CV 81%

Multimodal Flare Forecasting with Deep Learning

Grégoire Francisco, Sabrina Guastavino, Teresa Barata, João Fernandes, Dario Del Moro

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14150 2024-10-21 cs.AI cs.CL 81%

Utilizing Large Language Models for Event Deconstruction to Enhance Multimodal Aspect-Based Sentiment Analysis

Xiaoyong Huang, Heli Sun, Qunshu Gao, Wenjie Huang, Ruichen Cao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09379 2024-10-15 cs.CV cs.AI 81%

Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering

Ting Yu, Kunhao Fu, Jian Zhang, Qingming Huang, Jun Yu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Transactions on Image Processing

Journal ref Transactions on Image Processing, vol. 33, pp. 3115-3129, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07405 2024-10-11 cs.CV cs.AI 81%

Exploring Efficient Foundational Multi-modal Models for Video Summarization

Karan Samel, Apoorva Beedu, Nitish Sontakke, Irfan Essa

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05267 2024-10-08 cs.CL cs.CV 81%

Grounding Partially-Defined Events in Multimodal Data

Kate Sanders, Reno Kriz, David Etter, Hannah Recknor, Alexander Martin, Cameron Carpenter, Jingyang Lin, Benjamin Van Durme

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Preprint; 9 pages; 2024 EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16050 2024-10-04 cs.CV cs.CL 81%

Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

Yuxuan Wang, Yueqian Wang, Pengfei Wu, Jianxin Liang, Dongyan Zhao, Yang Liu, Zilong Zheng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments To appear at EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16099 2024-09-25 cs.CV cs.AI 81%

Neuromorphic Drone Detection: an Event-RGB Multimodal Approach

Gabriele Magrini, Federico Becattini, Pietro Pala, Alberto Del Bimbo, Antonio Porta

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at NeVi Workshop at ECCV24

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09446 2024-09-17 cs.CV cs.AI 81%

MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction

Yan Feng, Alexander Carballo, Keisuke Fujii, Robin Karlsson, Ming Ding, Kazuya Takeda

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21757 2024-09-13 cs.CV cs.MM 81%

Learning Video Context as Interleaved Multimodal Sequences

Kevin Qinghong Lin, Pengchuan Zhang, Difei Gao, Xide Xia, Joya Chen, Ziteng Gao, Jinheng Xie, Xuhong Xiao, Mike Zheng Shou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01560 2024-09-04 cs.CV cs.AI 81%

Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models

Bin Fu, Qiyang Wan, Jialin Li, Ruiping Wang, Xilin Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 39 pages, 28 figures, 4 tables. Accepted at The 35th British Machine Vision Conference (BMVC 2024). Project page at https://fubin29.github.io/Blocks-as-Probes/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10579 2024-09-04 cs.CL cs.AI 81%

DER-GCN: Dialogue and Event Relation-Aware Graph Convolutional Neural Network for Multimodal Dialogue Emotion Recognition

Wei Ai, Yuntao Shou, Tao Meng, Nan Yin, Keqin Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17428 2024-08-27 cs.CV cs.AI 81%

SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation

Qi Liu, Xinchen Liu, Kun Liu, Xiaoyan Gu, Wu Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07981 2024-08-16 cs.CV cs.AI 81%

LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning

Jiajie Li, Garrett Skinner, Gene Yang, Brian R Quaranto, Steven D Schwaitzberg, Peter C W Kim, Jinjun Xiong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04388 2024-08-09 cs.MM cs.AI cs.IR 81%

MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language Models

Haoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin, Yang Yang, Tat-Seng Chua

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07895 2024-07-30 cs.CV cs.CL cs.LG 81%

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, Chunyuan Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Project Page: https://llava-vl.github.io/blog/2024-06-16-llava-next-interleave/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06157 2024-07-09 cs.CV cs.AI 81%

Temporal Grounding of Activities using Multimodal Large Language Models

Young Chol Song

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04923 2024-07-09 cs.CV cs.CL 81%

OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding

Tiancheng Zhao, Qianqian Zhang, Kyusong Lee, Peng Liu, Lu Zhang, Chunxin Fang, Jiajia Liao, Kelei Jiang, Yibo Ma, Ruochen Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04697 2024-07-08 cs.CV cs.MM 81%

VCoME: Verbal Video Composition with Multimodal Editing Effects

Weibo Gong, Xiaojie Jin, Xin Li, Dongliang He, Xinglong Wu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18139 2024-06-27 cs.CL cs.CV 81%

LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Zhongwei Wan, Ziang Wu, Che Liu, Jinfa Huang, Zhihong Zhu, Peng Jin, Longyue Wang, Li Yuan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04873 2024-05-22 cs.CV cs.AI 81%

Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition

Xinzhe Ni, Yong Liu, Hao Wen, Yatai Ji, Jing Xiao, Yujiu Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICMR 2024 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05130 2024-05-09 cs.CV cs.MM 81%

Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection

Shengyang Sun, Xiaojin Gong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16557 2024-04-26 cs.CV cs.AI 81%

Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples

Kuofeng Gao, Jindong Gu, Yang Bai, Shu-Tao Xia, Philip Torr, Wei Liu, Zhifeng Li

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2401.11170

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07484 2024-04-12 cs.MM cs.AI 81%

Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios

Yuan Zhang, Xiaomei Tao, Hanxu Ai, Tao Chen, Yanling Gan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04763 2024-04-09 cs.CV cs.AI 81%

GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling

Hritik Bansal, Po-Nien Kung, P. Jeffrey Brantingham, Kai-Wei Chang, Nanyun Peng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 20 pages, 15 Figures, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01258 2024-04-03 cs.CV cs.AI 81%

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander Hauptmann, Yonatan Bisk, Yiming Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19221 2024-03-29 cs.CV cs.AI 81%

Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality

Sishuo Chen, Lei Li, Shuhuai Ren, Rundong Gao, Yuanxin Liu, Xiaohan Bi, Xu Sun, Lu Hou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Code available at https://github.com/lancopku/MR-VPC

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15096 2024-02-26 cs.LG cs.CV cs.MM 81%

Multimodal Transformer With a Low-Computational-Cost Guarantee

Sungjin Park, Edward Choi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted to ICASSP 2024 (5 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08486 2024-02-22 cs.CL cs.CV 81%

Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals

Te-Lin Wu, Alex Spangher, Pegah Alipoormolabashi, Marjorie Freedman, Ralph Weischedel, Nanyun Peng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments In Proceedings of the Conference of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05653 2024-02-07 cs.CV cs.MM 81%

Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos

Junbin Zhang, Pei-Hsuan Tsai, Meng-Hsun Tsai

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments 13 pages, 3 figures, 9 tables. Published on Applied Intelligence

Journal ref Applied Intelligence(2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16150 2024-02-05 cs.CV cs.AI cs.LG eess.IV 81%

Multimodal video and IMU kinematic dataset on daily life activities using affordable devices (VIDIMU)

Mario Martínez-Zarzuela, Javier González-Alonso, Míriam Antón-Rodríguez, Francisco J. Díaz-Pernas, Henning Müller, Cristina Simón-Martínez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Sci Data 10, 648 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏