arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2010.05406 2020-10-13 cs.CL cs.CV 81%

VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles

Mingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan, Dongyan Zhao, Rui Yan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by The 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09748 2020-08-25 cs.CV cs.AI cs.LG eess.IV eess.SP 81%

Multidomain Multimodal Fusion For Human Action Recognition Using Inertial Sensors

Zeeshan Ahmad, Naimul Khan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02678 2020-04-29 cs.CV cs.MM 81%

A Local-to-Global Approach to Multi-modal Movie Scene Segmentation

Anyi Rao, Linning Xu, Yu Xiong, Guodong Xu, Qingqiu Huang, Bolei Zhou, Dahua Lin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments CVPR2020. Project page: https://anyirao.com/projects/SceneSeg.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02205 2020-04-07 cs.CV cs.LG cs.MM 81%

Deep Multimodal Feature Encoding for Video Ordering

Vivek Sharma, Makarand Tapaswi, Rainer Stiefelhagen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments IEEE International Conference on Computer Vision (ICCV) Workshop on Large Scale Holistic Video Understanding. The datasets and code are available at https://github.com/vivoutlaw/tcbp

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.08854 2019-11-21 cs.CV cs.MM 81%

The dynamics of the stomatognathic system from 4D multimodal data

Agnieszka A. Tomaka, Leszek Luchowski, Dariusz Pojda, Michał Tarnawski, Krzysztof Domino

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Chapter 3 in A.Gadomski (ed.): Multiscale Locomotion: Its Active-Matter Addressing Physical Principles; UTP University of Science & Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02932 2019-10-08 cs.CV cs.IR cs.MM 81%

Multi-Modal Machine Learning for Flood Detection in News, Social Media and Satellite Sequences

Kashif Ahmad, Konstantin Pogorelov, Mohib Ullah, Michael Riegler, Nicola Conci, Johannes Langguth, Ala Al-Fuqaha

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Journal ref MediaEval 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.02511 2019-03-07 cs.CV cs.AI cs.LG 81%

Learning multimodal representations for sample-efficient recognition of human actions

Miguel Vasco, Francisco S. Melo, David Martins de Matos, Ana Paiva, Tetsunari Inamura

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 6 figures, submitted to 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.12563 2018-12-03 cs.CV cs.MM 81%

Deep Multimodal Learning: An Effective Method for Video Classification

Tianqi Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.03931 2017-12-12 cs.LG cs.AI cs.CV cs.GR cs.RO 81%

MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments

Manolis Savva, Angel X. Chang, Alexey Dosovitskiy, Thomas Funkhouser, Vladlen Koltun

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments MINOS is a simulator designed to support research on end-to-end navigation

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.05861 2017-10-25 cs.CV cs.LG cs.MM 81%

Continuous Multimodal Emotion Recognition Approach for AVEC 2017

Narotam Singh, Nittin Singh, Abhinav Dhall

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 4 pages, 3 figures, arXiv:1605.06778, arXiv:1512.03385

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.07200 2017-09-22 cs.CV cs.LG cs.MM 81%

Temporal Multimodal Fusion for Video Emotion Classification in the Wild

Valentin Vielzeuf, Stéphane Pateux, Frédéric Jurie

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Journal ref ACM - ICMI 2017, Nov 2017, Glasgow, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.04508 2017-06-15 cs.MM cs.CV 81%

Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification

Yu-Gang Jiang, Zuxuan Wu, Jinhui Tang, Zechao Li, Xiangyang Xue, Shih-Fu Chang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.07475 2017-02-27 cs.RO cs.AI cs.CV 81%

Sequence-based Multimodal Apprenticeship Learning For Robot Perception and Decision Making

Fei Han, Xue Yang, Yu Zhang, Hao Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures, accepted by ICRA'17

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.06603 2016-01-26 cs.MM cs.CV 81%

Egocentric Activity Recognition with Multimodal Fisher Vector

Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar, Bappaditya Mandal, Jie Lin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 5 pages, 4 figures, ICASSP 2016 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.00599 2016-01-05 cs.CV cs.IR cs.MM 81%

Multimodal Classification of Events in Social Media

Matthias Zeppelzauer, Daniel Schopfhauser

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Preprint of accepted manuscript for the Elsevier Image and Vision Computing Journal (IMAVIS). The paper will be published by IMAVIS under DOI 10.1016/j.imavis.2015.12.004

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.04024 2015-12-01 cs.CL cs.CV 81%

Multimodal Skip-gram Using Convolutional Pseudowords

Zachary Seymour, Yingming Li, Zhongfei Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.04831 2015-07-20 cs.CV cs.LG cs.MM cs.SD 81%

Deep Multimodal Speaker Naming

Yongtao Hu, Jimmy Ren, Jingwen Dai, Chang Yuan, Li Xu, Wenping Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15349 2024-04-25 eess.SP cs.LG cs.MM 80%

A Survey on Multimodal Wearable Sensor-based Human Action Recognition

Jianyuan Ni, Hao Tang, Syed Tousiful Haque, Yan Yan, Anne H. H. Ngu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments Multimodal Survey for Wearable Sensor-based Human Action Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05430 2023-09-26 cs.CV 80%

Ensemble Modeling for Multimodal Visual Action Recognition

Jyoti Kini, Sarah Fleischer, Ishan Dave, Mubarak Shah

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 22nd International Conference on Image Analysis and Processing Workshops - Multimodal Action Recognition on the MECCANO Dataset, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07483 2023-07-19 cs.CV 80%

Multimodal Distillation for Egocentric Action Recognition

Gorjan Radevski, Dusan Grujicic, Marie-Francine Moens, Matthew Blaschko, Tinne Tuytelaars

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2023; Codebase released at https://github.com/gorjanradevski/multimodal-distillation

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04916 2023-07-12 cs.CV eess.IV 80%

Rapid Deforestation and Burned Area Detection using Deep Multimodal Learning on Satellite Imagery

Gabor Fodor, Marcos V. Conde

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2023 Workshop on Multimodal Learning for Earth and Environment (MultiEarth)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07951 2021-09-17 cs.CV 80%

Overview of Tencent Multi-modal Ads Video Understanding Challenge

Zhenzhi Wang, Liyu Wu, Zhimin Li, Jiangfeng Xiong, Qinglin Lu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8-page extended version of our challenge paper in ACM MM 2021. It presents the overview of grand challenge "Multi-modal Ads Video Understanding" in ACM MM 2021. Our grand challenge is also the Tencent Advertising Algorithm Competition (TAAC) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13570 2019-11-25 cs.LG cs.AI cs.NE stat.ML 80%

Factorized Inference in Deep Markov Models for Incomplete Multimodal Time Series

Tan Zhi-Xuan, Harold Soh, Desmond C. Ong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 4 figures, accepted to AAAI 2020, code available at: https://github.com/ztangent/multimodal-dmm

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02908 2019-03-14 cs.RO cs.CV 80%

Multi-Modal Obstacle Detection in Unstructured Environments with Conditional Random Fields

Mikkel Kragh, James Underwood

专题命中 视频多模态 :multi-modal(title);multimodal(abstract,comments);分类 cs.CV

Comments This is the accepted version of the following article: Kragh M, Underwood J. Multimodal obstacle detection in unstructured environments with conditional random fields. J Field Robotics. 2019, 1-20., which has been published in final form at https://doi.org/10.1002/rob.21866

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00721 2018-05-03 cs.CV 80%

Joint Surgical Gesture and Task Classification with Multi-Task and Multimodal Learning

Duygu Sarikaya, Khurshid A. Guru, Jason J. Corso

专题命中 视频多模态 :multimodal(title,comments);multi-modal(abstract);分类 cs.CV

Comments Keywords Robot-Assisted Surgery, Surgical Gesture Classification, Multi-task Learning, Multimodal Learning, Long Short-term Recurrent Neural Networks, Convolutional Neural Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.05212 2016-05-18 cs.LG cs.CV 80%

Multimodal Sparse Coding for Event Detection

Youngjune Gwon, William Campbell, Kevin Brady, Douglas Sturim, Miriam Cha, H. T. Kung

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Multimodal Machine Learning Workshop at NIPS 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13210 2026-08-14 cs.CV cs.AI cs.MM 新提交 80%

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

NARU:用于理解日语超长视频中叙事演变与文化细微差别的基准

Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma

机构 * The University of Tokyo(东京大学) Kyushu University(九州大学) Macau University of Science and Technology(澳门科技大学) Infinimind Japan Inc.(Infinimind日本公司) University of Alberta(阿尔伯塔大学)

专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI、cs.MM

AI总结 该研究推出NARU基准,涵盖155个146.8小时日语视频的1481个问题,评估模型在长程叙事整合与文化推理上的局限,为MLLM开发提供测试平台。

Comments Yuheng Huang and Jianlang Chen contributed equally to this work. More details available on the project's website https://ma-labo.github.io/naru/ and https://infinimind.io/en/company/news/2026/narubench-release

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12290 2026-08-13 cs.CV cs.AI cs.MM 新提交 80%

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

超越试错:面向图像到视频一致性的智能体优化

Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson

机构 * Google Cloud(谷歌云) Google DeepMind(谷歌DeepMind)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 针对图像到视频模型试错效率低的问题,提出Agentic Self-Improvement框架,通过两阶段优化提升视频与文本一致性,生成视频胜率达69%,为视频生成模型提供实用可控的优化方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11534 2026-08-11 cs.CV cs.CL cs.MM 版本更新 80%

HFS: Holistic Query-Aware Frame Selection for Efficient Video Understanding

HFS: 为高效视频推理的全局查询感知帧选择

Yiqing Yang, Yun Li, Daiqing Qi, Lehan Yang, Tianlong Wang, Wenhao Zhang, Sheng Li, Kin-man Lam

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 HFS提出一种端到端可训练的帧选择框架,通过任务自适应方法提升视频推理效率。

Comments Accepted to the Main Track of ACM Multimedia 2026 (ACM MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04515 2026-08-06 cs.CV cs.AI cs.CL 新提交 80%

CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding

CARVE:用于高效3D医学体积理解的视觉证据跨切片各向异性重分配

Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang

专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.CL、cs.AI

AI总结 针对3D医学体积理解中切片式MLLM的视觉令牌冗余问题,提出无需训练的CARVE框架,通过跨切片各向异性重分配压缩80%令牌,在AMOS-MM等基准上性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏