arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2306.17778 2024-01-23 cs.CV cs.LG 57%

Look, Remember and Reason: Grounded reasoning in videos with language models

Apratim Bhattacharyya, Sunny Panchal, Mingu Lee, Reza Pourreza, Pulkit Madan, Roland Memisevic

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments To appear at ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13336 2024-01-23 cs.LG cs.AI 57%

Dual-Branched Spatio-temporal Fusion Network for Multi-horizon Tropical Cyclone Track Forecast

Zili Liu, Kun Hao, Xiaoyi Geng, Zhenwei Shi

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08010 2024-01-22 cs.CV cs.LG 57%

EZ-CLIP: Efficient Zeroshot Video Action Recognition

Shahzad Ahmad, Sukalpa Chanda, Yogesh S Rawat

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09732 2024-01-19 cs.CV 57%

Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation

Zesen Cheng, Kehan Li, Hao Li, Peng Jin, Chang Liu, Xiawu Zheng, Rongrong Ji, Jie Chen

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08449 2024-01-17 cs.MM 57%

CLIPRerank: An Extremely Simple Method for Improving Ad-hoc Video Search

Aozhu Chen, Fangming Zhou, Ziyuan Wang, Xirong Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.MM

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04718 2024-01-12 cs.CV 57%

Jump Cut Smoothing for Talking Heads

Xiaojuan Wang, Taesung Park, Yang Zhou, Eli Shechtman, Richard Zhang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Correct typos in the caption of Figure 1; Change the project website address. Project page: https://jeanne-wang.github.io/jumpcutsmoothing/

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03397 2024-01-11 cs.LG cs.AI stat.AP 57%

Predicting the Skies: A Novel Model for Flight-Level Passenger Traffic Forecasting

Sina Ehsani, Elina Sergeeva, Wendy Murdy, Benjamin Fox

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 6 figures, to be published

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04150 2024-01-10 cs.CV 57%

Two-stream joint matching method based on contrastive learning for few-shot action recognition

Long Deng, Ziqiang Li, Bingxin Zhou, Zhongming Chen, Ao Li, Yongxin Ge

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14277 2024-01-10 cs.CV 57%

G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory

Hongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li, Zhihong Zhu, Yuexian Zou

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments ICCV2023 oral, release the code

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.05401 2024-01-09 cs.CV 57%

Benchmarking Joint Face Spoofing and Forgery Detection with Visual and Physiological Cues

Zitong Yu, Rizhao Cai, Zhi Li, Wenhan Yang, Jingang Shi, Alex C. Kot

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Dependable and Secure Computing (TDSC). Corresponding authors: Zitong Yu and Wenhan Yang

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09310 2024-01-08 cs.CV 57%

Language-Assisted Deep Learning for Autistic Behaviors Recognition

Andong Deng, Taojiannan Yang, Chen Chen, Qian Chen, Leslie Neely, Sakiko Oyama

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Smart Health Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02138 2024-01-05 cs.CV 57%

Explore Human Parsing Modality for Action Recognition

Jinfu Liu, Runwei Ding, Yuhang Wen, Nan Dai, Fanyang Meng, Shen Zhao, Mengyuan Liu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2307.07977

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01993 2024-01-05 cs.RO cs.AI 57%

On Time-Indexing as Inductive Bias in Deep RL for Sequential Manipulation Tasks

M. Nomaan Qureshi, Ben Eisner, David Held

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01529 2024-01-04 cs.CV 57%

Glance and Focus: Memory Prompting for Multi-Event Video Question Answering

Ziyi Bai, Ruiping Wang, Xilin Chen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted in NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00271 2024-01-02 cs.CV 57%

HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid Explorations

Yilan Dong, Chunlin Yu, Ruiyang Ha, Ye Shi, Yuexin Ma, Lan Xu, Yanwei Fu, Jingya Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16071 2023-12-27 cs.NE cs.AI cs.GR cs.LG 57%

Event-based Shape from Polarization with Spiking Neural Networks

Peng Kang, Srutarshi Banerjee, Henry Chopp, Aggelos Katsaggelos, Oliver Cossairt

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15670 2023-12-27 cs.CV 57%

Open-Vocabulary Video Relation Extraction

Wentao Tian, Zheng Wang, Yuqian Fu, Jingjing Chen, Lechao Cheng

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments accpeted by AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12155 2023-12-20 cs.CV 57%

Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval

Zhihang Liu, Jun Li, Hongtao Xie, Pandeng Li, Jiannan Ge, Sun-Ao Liu, Guoqing Jin

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05797 2023-12-19 cs.CV 57%

Multimodality in Online Education: A Comparative Study

Praneeta Immadisetty, Pooja Rajesh, Akshita Gupta, Anala M R, Soumya A, K. N. Subramanya

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09525 2023-12-18 cs.CV 57%

Hierarchical Graph Pattern Understanding for Zero-Shot VOS

Gensheng Pei, Fumin Shen, Yazhou Yao, Tao Chen, Xian-Sheng Hua, Heng-Tao Shen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments accepted by IEEE Transactions on Image Processing

Journal ref IEEE Transactions on Image Processing 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08644 2023-12-15 cs.CV 57%

Generative Model-based Feature Knowledge Distillation for Action Recognition

Guiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao, Shusen Yang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted on AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01971 2023-12-13 cs.CV 57%

District-scale surface temperatures generated from high-resolution longitudinal thermal infrared images

Subin Lin, Vasantha Ramani, Miguel Martin, Pandarasamy Arjunan, Adrian Chong, Filip Biljecki, Marcel Ignatius, Kameshwar Poolla, Clayton Miller

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Journal ref Sci Data 10, 859 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06699 2023-12-13 cs.CV cs.LG 57%

Leveraging Generative Language Models for Weakly Supervised Sentence Component Analysis in Video-Language Joint Learning

Zaber Ibn Abdul Hakim, Najibul Haque Sarker, Rahul Pratap Singh, Bishmoy Paul, Ali Dabouei, Min Xu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06598 2023-12-12 cs.CV 57%

Early Action Recognition with Action Prototypes

Guglielmo Camporese, Alessandro Bergamo, Xunyu Lin, Joseph Tighe, Davide Modolo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01244 2023-12-05 cs.CL 57%

Challenges and Applications of Automated Extraction of Socio-political Events from Text (CASE 2023): Workshop and Shared Task Report

Ali Hürriyetoğlu, Hristo Tanev, Osman Mutlu, Surendrabikram Thapa, Fiona Anting Tan, Erdem Yörük

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments https://aclanthology.org/2023.case-1.22

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01083 2023-12-05 cs.CV 57%

Consistency Prototype Module and Motion Compensation for Few-Shot Action Recognition (CLIP-CP$\mathbf{M^2}$C)

Fei Guo, Li Zhu, YiKang Wang, Han Qi

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00937 2023-12-05 cs.CV 57%

Zero-Shot Video Question Answering with Procedural Programs

Rohan Choudhury, Koichiro Niinuma, Kris M. Kitani, László A. Jeni

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00055 2023-12-04 cs.CV cs.LG cs.RO 57%

LEAP: LLM-Generation of Egocentric Action Programs

Eadom Dessalene, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Dataset: https://drive.google.com/drive/folders/1Cpkw_TI1IIxXdzor0pOXG3rWJWuKU5Ex?usp=drive_link

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17942 2023-12-01 cs.CV 57%

Object-based (yet Class-agnostic) Video Domain Adaptation

Dantong Niu, Amir Bar, Roei Herzig, Trevor Darrell, Anna Rohrbach

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00450 2023-11-30 cs.CV 57%

Sketch-based Video Object Localization

Sangmin Woo, So-Yeong Jeon, Jinyoung Park, Minji Son, Sumin Lee, Changick Kim

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments WACV 2024; Code: https://github.com/sangminwoo/SVOL

详情

展开后加载摘要…

URL PDF HTML 收藏