arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2412.00927 2024-12-03 cs.CV 57%

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation

Weiming Ren, Huan Yang, Jie Min, Cong Wei, Wenhu Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://tiger-ai-lab.github.io/VISTA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13800 2024-12-03 cs.CV 57%

Aligning Step-by-Step Instructional Diagrams to Video Demonstrations

Jiahao Zhang, Anoop Cherian, Yanbin Liu, Yizhak Ben-Shabat, Cristian Rodriguez, Stephen Gould

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to CVPR'23

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16872 2024-11-28 cs.IR cs.AI cs.ET 57%

Enabling Adoption of Regenerative Agriculture through Soil Carbon Copilots

Margaret Capetz, Swati Sharma, Rafael Padilha, Peder Olsen, Jessica Wolk, Emre Kiciman, Ranveer Chandra

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17481 2024-11-27 cs.CV 57%

Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding

Mengzhao Wang, Huafeng Li, Yafei Zhang, Jinxing Li, Minghong Xie, Dapeng Tao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments This work has been accepted with mandatory minor revisions by TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16989 2024-11-27 cs.CV cs.LG 57%

CMAViT: Integrating Climate, Managment, and Remote Sensing Data for Crop Yield Estimation with Multimodel Vision Transformers

Hamid Kamangir, Brent. S. Sams, Nick Dokoozlian, Luis Sanchez, J. Mason. Earles

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11288 2024-11-27 cs.CV 57%

Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action Recognition

Yang Chen, Jingcai Guo, Song Guo, Dacheng Tao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15436 2024-11-26 cs.CV 57%

ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance

Haijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian, Jian Yang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04292 2024-11-25 cs.LG cs.AI 57%

AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies

Xixi Hu, Bo Liu, Xingchao Liu, Qiang Liu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments NeuRIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09817 2024-11-25 cs.CV stat.AP 57%

Interpretable machine learning for time-to-event prediction in medicine and healthcare

Hubert Baniecki, Bartlomiej Sobieski, Patryk Szatkowski, Przemyslaw Bombinski, Przemyslaw Biecek

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments An extended version of an AIME 2023 paper submitted to Artificial Intelligence in Medicine

Journal ref Artificial Intelligence in Medicine, vol. 159, 103026, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14505 2024-11-25 cs.CV 57%

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Weiheng Lu, Jian Li, An Yu, Ming-Ching Chang, Shengpeng Ji, Min Xia

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14153 2024-11-22 eess.AS 57%

MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation

Hengyi Hong, Qing Wang, Jun Du, Ruoyu Wei, Mingqi Cai, Xin Fang

专题命中 视频多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13917 2024-11-22 cs.MM 57%

SpikEmo: Enhancing Emotion Recognition With Spiking Temporal Dynamics in Conversations

Xiaomin Yu, Feiyang Wang, Ziyue Qiao

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12778 2024-11-21 cs.HC cs.AI 57%

Lucia: A Temporal Computing Platform for Contextual Intelligence

Weizhe Lin, Junxiao Shen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18861 2024-11-12 cs.CV 57%

Visual Mamba: A Survey and New Outlooks

Rui Xu, Shu Yang, Yihui Wang, Yu Cai, Bo Du, Hao Chen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06560 2024-11-11 stat.AP cs.AI 57%

TCKAN:A Novel Integrated Network Model for Predicting Mortality Risk in Sepsis Patients

Fanglin Dong

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04549 2024-11-08 cs.RO cs.AI cs.LG 57%

Vision Language Models are In-Context Value Learners

Yecheng Jason Ma, Joey Hejna, Ayzaan Wahid, Chuyuan Fu, Dhruv Shah, Jacky Liang, Zhuo Xu, Sean Kirmani, Peng Xu, Danny Driess, Ted Xiao, Jonathan Tompson, Osbert Bastani, Dinesh Jayaraman, Wenhao Yu, Tingnan Zhang, Dorsa Sadigh, Fei Xia

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Project website and demo: https://generative-value-learning.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.20024 2024-11-05 cs.CV 57%

eMoE-Tracker: Environmental MoE-based Transformer for Robust Event-guided Object Tracking

Yucheng Chen, Lin Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments RGB-event single object tracking

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04640 2024-11-01 cs.RO cs.AI cs.LG 57%

Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress

Christopher Agia, Rohan Sinha, Jingyun Yang, Zi-ang Cao, Rika Antonova, Marco Pavone, Jeannette Bohg

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Project page: https://sites.google.com/stanford.edu/sentinel. 35 pages, 9 figures. Accepted to the Conference on Robot Learning (CoRL) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15223 2024-11-01 cs.CV cs.LG cs.RO 57%

iVideoGPT: Interactive VideoGPTs are Scalable World Models

Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, Mingsheng Long

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2024. Code is available at project website: https://thuml.github.io/iVideoGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06651 2024-10-28 eess.SP cs.AI cs.HC cs.LG 57%

Accoustate: Auto-annotation of IMU-generated Activity Signatures under Smart Infrastructure

Soumyajit Chatterjee, Arun Singh, Bivas Mitra, Sandip Chakraborty

专题命中 视频多模态 :cross-modal(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

Journal ref IEEE DCOSS-IoT 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06040 2024-10-28 cs.CV 57%

Vript: A Video Is Worth Thousands of Words

Dongjie Yang, Suyuan Huang, Chengqiang Lu, Xiaodong Han, Haoxin Zhang, Yan Gao, Yao Hu, Hai Zhao

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS Dataset & Benchmark track

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02982 2024-10-28 cs.CV 57%

CLIP-guided Prototype Modulating for Few-shot Action Recognition

Xiang Wang, Shiwei Zhang, Jun Cen, Changxin Gao, Yingya Zhang, Deli Zhao, Nong Sang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the Springer for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00188 2024-10-23 cs.CV 57%

REACT: Recognize Every Action Everywhere All At Once

Naga VS Raviteja Chappa, Pha Nguyen, Page Daniel Dobbs, Khoa Luu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13662 2024-10-18 cs.CV 57%

ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions

Shailaja Keyur Sampat, Yezhou Yang, Chitta Baral

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 15 pages, 3 figures. arXiv admin note: text overlap with arXiv:2004.10796 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13598 2024-10-18 cs.CV 57%

Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding

Jongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae Won Cho, Joon Son Chung

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACMMM 24

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11404 2024-10-16 cs.CV 57%

MoChat: Joints-Grouped Spatio-Temporal Grounding LLM for Multi-Turn Motion Comprehension and Description

Jiawei Mo, Yixuan Chen, Rifen Lin, Yongkang Ni, Min Zeng, Xiping Hu, Min Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14407 2024-10-10 cs.LG cs.CV cs.RO 57%

Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training

Haoran He, Chenjia Bai, Ling Pan, Weinan Zhang, Bin Zhao, Xuelong Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2024. 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04449 2024-10-08 cs.CV 57%

Video Summarization Techniques: A Comprehensive Review

Toqa Alaa, Ahmad Mongy, Assem Bakr, Mariam Diab, Walid Gomaa

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20018 2024-10-03 cs.CV 57%

Visual Context Window Extension: A New Perspective for Long Video Understanding

Hongchen Wei, Zhenzhong Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13947 2024-10-02 cs.HC cs.AI 57%

BlendScape: Enabling End-User Customization of Video-Conferencing Environments through Generative AI

Shwetha Rajaram, Nels Numan, Balasaravanan Thoravi Kumaravel, Nicolai Marquardt, Andrew D. Wilson

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments ACM UIST 2024

详情

展开后加载摘要…

URL PDF HTML 收藏