arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2311.15769 2023-11-28 cs.CV 57%

Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning

Huanjin Yao, Wenhao Wu, Zhiheng Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13946 2023-11-27 cs.MM 57%

Weakly-Supervised Video Moment Retrieval via Regularized Two-Branch Proposal Networks with Erasing Mechanism

Haoyuan Li, Zhou Zhao, Zhu Zhang, Zhijie Lin

专题命中 视频多模态 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11933 2023-11-21 cs.CV 57%

Fully Transformer-Equipped Architecture for End-to-End Referring Video Object Segmentation

Ping Li, Yu Zhang, Li Yuan, Xianghua Xu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Journal ref Information and Processing Management (IPM'2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08245 2023-11-15 cs.CV 57%

TENT: Connect Language Models with IoT Sensors for Zero-Shot Activity Recognition

Yunjiao Zhou, Jianfei Yang, Han Zou, Lihua Xie

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Preprint manuscript in submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08143 2023-11-15 cs.CL 57%

Sinkhorn Transformations for Single-Query Postprocessing in Text-Video Retrieval

Konstantin Yakovlev, Gregory Polyakov, Ilseyar Alimova, Alexander Podolskiy, Andrey Bout, Sergey Nikolenko, Irina Piontkovskaya

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments SIGIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02559 2023-11-07 cs.CV 57%

Rotation Invariant Transformer for Recognizing Object in UAVs

Shuoyi Chen, Mang Ye, Bo Du

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ACM MM2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19432 2023-10-31 cs.RO cs.AI cs.LG 57%

Explaining the Decisions of Deep Policy Networks for Robotic Manipulations

Seongun Kim, Jaesik Choi

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15946 2023-10-25 cs.CV 57%

ShARc: Shape and Appearance Recognition for Person Identification In-the-wild

Haidong Zhu, Wanrong Zheng, Zhaoheng Zheng, Ram Nevatia

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15568 2023-10-25 cs.CV 57%

I$^2$MD: 3D Action Representation Learning with Inter- and Intra-modal Mutual Distillation

Yunyao Mao, Jiajun Deng, Wengang Zhou, Zhenbo Lu, Wanli Ouyang, Houqiang Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments submitted to IJCV. arXiv admin note: substantial text overlap with arXiv:2208.12448

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12724 2023-10-20 cs.CV 57%

Query-aware Long Video Localization and Relation Discrimination for Deep Video Understanding

Yuanxing Xu, Yuting Wei, Bin Wu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ACM MM 2023 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12296 2023-10-20 cs.CV 57%

Understanding Video Transformers for Segmentation: A Survey of Application and Interpretability

Rezaul Karim, Richard P. Wildes

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14829 2023-10-11 cs.CV eess.SP 57%

Deakin RF-Sensing: Experiments on Correlated Knowledge Distillation for Monitoring Human Postures with Radios

Shiva Raj Pokhrel, Jonathan Kua, Deol Satish, Philip Williams, Arkady Zaslavsky, Seng W. Loke, Jinho Choi

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09756 2023-10-11 cs.CV 57%

Video Action Recognition with Attentive Semantic Units

Yifei Chen, Dapeng Chen, Ruijin Liu, Hao Li, Wei Peng

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2023

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 10170-10180

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05825 2023-10-10 cs.MM cs.HC 57%

Write What You Want: Applying Text-to-video Retrieval to Audiovisual Archives

Yuchen Yang

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04716 2023-10-10 cs.CV 57%

Reinforced UI Instruction Grounding: Towards a Generic UI Task Automation API

Zhizheng Zhang, Wenxuan Xie, Xiaoyi Zhang, Yan Lu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08149 2023-10-10 cs.LG cs.AI 57%

Neural Mixed Effects for Nonlinear Personalized Predictions

Torsten Wörtwein, Nicholas Allen, Lisa B. Sheeber, Randy P. Auerbach, Jeffrey F. Cohn, Louis-Philippe Morency

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08303 2023-10-06 cs.CV 57%

Leveraging Next-Active Objects for Context-Aware Anticipation in Egocentric Videos

Sanket Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted in WACV'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02532 2023-10-05 cs.CV 57%

ShaSTA-Fuse: Camera-LiDAR Sensor Fusion to Model Shape and Spatio-Temporal Affinities for 3D Multi-Object Tracking

Tara Sadjadpour, Rares Ambrus, Jeannette Bohg

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01351 2023-10-03 cs.CV 57%

Streaming Motion Forecasting for Autonomous Driving

Ziqi Pang, Deva Ramanan, Mengtian Li, Yu-Xiong Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments IROS 2023, 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00740 2023-10-03 cs.CV cs.CY cs.LG 57%

Top-down Green-ups: Satellite Sensing and Deep Models to Predict Buffelgrass Phenology

Lucas Rosenblatt, Bin Han, Erin Posthumus, Theresa Crimmins, Bill Howe

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17239 2023-10-02 cs.CV 57%

EGVD: Event-Guided Video Deraining

Yueyi Zhang, Jin Wang, Wenming Weng, Xiaoyan Sun, Zhiwei Xiong

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15596 2023-09-28 cs.RO cs.CV 57%

PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation

Shizhe Chen, Ricardo Garcia, Cordelia Schmid, Ivan Laptev

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to CoRL 2023. Project website: https://www.di.ens.fr/willow/research/polarnet/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13962 2023-09-26 cs.CV eess.IV 57%

Egocentric RGB+Depth Action Recognition in Industry-Like Settings

Jyoti Kini, Sarah Fleischer, Ishan Dave, Mubarak Shah

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11006 2023-09-21 cs.RO cs.CV 57%

STARNet: Sensor Trustworthiness and Anomaly Recognition via Approximated Likelihood Regret for Robust Edge Autonomy

Nastaran Darabi, Sina Tayebati, Sureshkumar S., Sathya Ravi, Theja Tulabandhula, Amit R. Trivedi

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09473 2023-09-19 cs.CV 57%

Self-supervised Multi-view Clustering in Computer Vision: A Survey

Jiatai Wang, Zhiwei Xu, Xuewen Yang, Hailong Li, Bo Li, Xuying Meng

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09421 2023-09-19 cs.MM 57%

Unified Pretraining Target Based Video-music Retrieval With Music Rhythm And Video Optical Flow Information

Tianjun Mao, Shansong Liu, Yunxuan Zhang, Dian Li, Ying Shan

专题命中 视频多模态 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06236 2023-09-13 cs.LG cs.CL 57%

The first step is the hardest: Pitfalls of Representing and Tokenizing Temporal Data for Large Language Models

Dimitris Spathis, Fahim Kawsar

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at the Generative AI for Pervasive Computing Symposium (GenAI4PC) at UbiComp 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06102 2023-09-13 cs.CV 57%

Can we predict the Most Replayed data of video streaming platforms?

Alessandro Duico, Ombretta Strafforello, Jan van Gemert

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted Extended Abstract at ICCV 2023 Workshop on AI for Creative Video Editing and Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04302 2023-09-11 cs.CV 57%

Have We Ever Encountered This Before? Retrieving Out-of-Distribution Road Obstacles from Driving Scenes

Youssef Shoeb, Robin Chan, Gesina Schwalbe, Azarm Nowzard, Fatma Güney, Hanno Gottschalk

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 11 pages, 7 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03473 2023-09-08 cs.CV 57%

Temporal Collection and Distribution for Referring Video Object Segmentation

Jiajin Tang, Ge Zheng, Sibei Yang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023; Project page: https://toneyaya.github.io/tempcd/

详情

展开后加载摘要…

URL PDF HTML 收藏