arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2312.08367 2024-10-02 cs.CV 57%

ViLA: Efficient Video-Language Alignment for Video Question Answering

Xijun Wang, Junbang Liang, Chun-Kai Wang, Kenan Deng, Yu Lou, Ming Lin, Shan Yang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17987 2024-09-27 cs.CV cs.HC 57%

LLM4Brain: Training a Large Language Model for Brain Video Understanding

Ruizhe Zheng, Lichao Sun

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ECCV2024 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14109 2024-09-27 cs.CV 57%

Vision-Language Models Assisted Unsupervised Video Anomaly Detection

Yalong Jiang, Liquan Mao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16145 2024-09-25 cs.CV 57%

Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment

Yuxiao Chen, Kai Li, Wentao Bao, Deep Patel, Yu Kong, Martin Renqiang Min, Dimitris N. Metaxas

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13606 2024-09-23 cs.CV cs.LG 57%

Towards Child-Inclusive Clinical Video Understanding for Autism Spectrum Disorder

Aditya Kommineni, Digbalay Bose, Tiantian Feng, So Hyun Kim, Helen Tager-Flusberg, Somer Bishop, Catherine Lord, Sudarsana Kadiri, Shrikanth Narayanan

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10741 2024-09-18 cs.SE cs.CL 57%

NaviQAte: Functionality-Guided Web Application Navigation

Mobina Shahbandeh, Parsa Alian, Noor Nashid, Ali Mesbah

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03612 2024-09-18 cs.CV cs.LG 57%

JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling

Seok Hwan Lee, Taein Son, Soo Won Seo, Jisong Kim, Jun Won Choi

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 31 pages, 10 figures, update references

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15051 2024-09-17 cs.CV 57%

Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval

Yiyang Jiang, Wengyu Zhang, Xulu Zhang, Xiaoyong Wei, Chang Wen Chen, Qing Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08585 2024-09-16 cs.CV eess.IV 57%

Optimizing 4D Lookup Table for Low-light Video Enhancement via Wavelet Priori

Jinhong He, Minglong Xue, Wenhai Wang, Mingliang Zhou

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08277 2024-09-13 cs.CV 57%

Depth on Demand: Streaming Dense Depth from a Low Frame Rate Active Sensor

Andrea Conti, Matteo Poggi, Valerio Cambareri, Stefano Mattoccia

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted for publication at the European Conference on Computer Vision (ECCV) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07239 2024-09-12 cs.CV 57%

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

Yang Liu, Pengxiang Ding, Siteng Huang, Min Zhang, Han Zhao, Donglin Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05921 2024-09-11 cs.LG cs.AI 57%

STLLM-DF: A Spatial-Temporal Large Language Model with Diffusion for Enhanced Multi-Mode Traffic System Forecasting

Zhiqi Shao, Haoning Xi, Haohui Lu, Ze Wang, Michael G. H. Bell, Junbin Gao

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 26 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04900 2024-09-10 cs.CV 57%

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

Nina Shvetsova, Anna Kukleva, Xudong Hong, Christian Rupprecht, Bernt Schiele, Hilde Kuehne

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments https://github.com/ninatu/howtocaption

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02385 2024-09-05 cs.CV 57%

Unified Framework with Consistency across Modalities for Human Activity Recognition

Tuyen Tran, Thao Minh Le, Hung Tran, Truyen Tran

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to BMVC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00180 2024-09-05 cs.CV cs.LG 57%

MMA-MRNNet: Harnessing Multiple Models of Affect and Dynamic Masked RNN for Precise Facial Expression Intensity Estimation

Dimitrios Kollias, Andreas Psaroudakis, Anastasios Arsenos, Paraskevi Theofilou, Chunchang Shao, Guanyu Hu, Ioannis Patras

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00851 2024-09-04 cs.IR cs.LG cs.SD eess.AS 57%

Dissecting Temporal Understanding in Text-to-Audio Retrieval

Andreea-Maria Oncescu, João F. Henriques, A. Sophia Koepke

专题命中 视频多模态 :multimodal(abstract);分类 eess.AS

Comments 9 pages, 5 figures, ACM Multimedia 2024, https://www.robots.ox.ac.uk/~vgg/research/audio-retrieval/dtu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16380 2024-08-30 cs.CV 57%

Exploiting temporal information to detect conversational groups in videos and predict the next speaker

Lucrezia Tosato, Victor Fortier, Isabelle Bloch, Catherine Pelachaud

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to Pattern Recognition Letter, 8 pages, 10 figures

Journal ref Pattern Recognition Letters Volume 177, January 2024, Pages 164 168

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16219 2024-08-30 cs.CV 57%

Training-free Video Temporal Grounding using Large-scale Pre-trained Models

Minghang Zheng, Xinhao Cai, Qingchao Chen, Yuxin Peng, Yang Liu

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12471 2024-08-29 cs.CV 57%

Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection

Ee Yeo Keat, Zhang Hao, Alexander Matyasko, Basura Fernando

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15209 2024-08-28 cs.MM 57%

Sec2Sec Co-attention for Video-Based Apparent Affective Prediction

Mingwei Sun, Kunpeng Zhang

专题命中 视频多模态 :audio-visual(abstract);分类 cs.MM

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14916 2024-08-28 cs.CV 57%

Towards Real-world Event-guided Low-light Video Enhancement and Deblurring

Taewoo Kim, Jaeseok Jeong, Hoonhee Cho, Yuhwan Jeong, Kuk-Jin Yoon

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted in ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14743 2024-08-28 cs.CV cs.IR 57%

Personalized Video Summarization using Text-Based Queries and Conditional Modeling

Jia-Hong Huang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Ph.D. thesis, 137 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14469 2024-08-27 cs.CV 57%

Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos

Qirui Chen, Shangzhe Di, Weidi Xie

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11344 2024-08-27 cs.CV 57%

Centering the Value of Every Modality: Towards Efficient and Resilient Modality-agnostic Semantic Segmentation

Xu Zheng, Yuanhuiyi Lyu, Jiazhou Zhou, Lin Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11384 2024-08-22 cs.LG cs.AI 57%

Data-Centric Machine Learning for Earth Observation: Necessary and Sufficient Features

Hiba Najjar, Marlon Nuske, Andreas Dengel

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at MACLEAN workshop, ECML/PKDD 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10541 2024-08-21 cs.CV 57%

The Instance-centric Transformer for the RVOS Track of LSVOS Challenge: 3rd Place Solution

Bin Cao, Yisi Zhang, Hanyi Wang, Xingjian He, Jing Liu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2406.13939

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08500 2024-08-19 cs.CV 57%

CoSEC: A Coaxial Stereo Event Camera Dataset for Autonomous Driving

Shihan Peng, Hanyu Zhou, Hao Dong, Zhiwei Shi, Haoyue Liu, Yuxing Duan, Yi Chang, Luxin Yan

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19988 2024-08-19 cs.MM 57%

HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR Headsets

Yili Jin, Xize Duan, Fangxin Wang, Xue Liu

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

Comments Accepted by ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15706 2024-08-16 cs.CV 57%

Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition

Jinfu Liu, Chen Chen, Mengyuan Liu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07600 2024-08-15 cs.CV 57%

Disentangle and denoise: Tackling context misalignment for video moment retrieval

Kaijing Ma, Han Fang, Xianghao Zang, Chao Ban, Lanxiang Zhou, Zhongjiang He, Yongxiang Li, Hao Sun, Zerun Feng, Xingsong Hou

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏