arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2309.00133 2023-09-04 cs.CV 57%

Distraction-free Embeddings for Robust VQA

Atharvan Dogra, Deeksha Varshney, Ashwin Kalyan, Ameet Deshpande, Neeraj Kumar

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00135 2023-08-30 cs.CV 57%

TeViS:Translating Text Synopses to Video Storyboards

Xu Gu, Yuchong Sun, Feiyue Ni, Shizhe Chen, Xihua Wang, Ruihua Song, Boyuan Li, Xiang Cao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14713 2023-08-29 cs.CV 57%

R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple Cameras

Aron Schmied, Tobias Fischer, Martin Danelljan, Marc Pollefeys, Fisher Yu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2023. Project page is available at https://www.vis.xyz/pub/r3d3/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12199 2023-08-24 cs.CV 57%

Towards Real-Time Analysis of Broadcast Badminton Videos

Nitin Nilesh, Tushar Sharma, Anurag Ghosh, C. V. Jawahar

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11489 2023-08-24 cs.CV 57%

Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition

Qitong Wang, Long Zhao, Liangzhe Yuan, Ting Liu, Xi Peng

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Proceedings of IEEE International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11062 2023-08-23 cs.CV cs.LG 57%

UnLoc: A Unified Framework for Video Localization Tasks

Shen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab, Zhonghao Wang, Weina Ge, David Ross, Cordelia Schmid

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05463 2023-08-22 cs.CV 57%

EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone

Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, Pengchuan Zhang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Published in ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18969 2023-08-22 cs.CV 57%

MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction

Jing Wang, Aixin Sun, Hao Zhang, Xiaoli Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACL 2023

Journal ref ACL 2023 long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01203 2023-08-22 cs.CV 57%

Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus

Xiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li, Bhiksha Raj, Yan Lu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments iccv 2023, https://github.com/lxa9867/R2VOS

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09475 2023-08-21 cs.CV 57%

Video-Instrument Synergistic Network for Referring Video Instrument Segmentation in Robotic Surgery

Hongqiu Wang, Lei Zhu, Guang Yang, Yike Guo, Shichen Zhang, Bo Xu, Yueming Jin

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06979 2023-08-21 cs.HC cs.LG cs.MM 57%

A Weakly Supervised Approach to Emotion-change Prediction and Improved Mood Inference

Soujanya Narayana, Ibrahim Radwan, Ravikiran Parameshwara, Iman Abbasnejad, Akshay Asthana, Ramanathan Subramanian, Roland Goecke

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

Comments 9 pages, 3 figures, 6 tables, published in IEEE International Conference on Affective Computing and Intelligent Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03903 2023-08-14 cs.CV 57%

Adversarial Self-Attack Defense and Spatial-Temporal Relation Mining for Visible-Infrared Video Person Re-Identification

Huafeng Li, Le Xu, Yafei Zhang, Dapeng Tao, Zhengtao Yu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 11 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12364 2023-08-14 cs.LG cs.AI 57%

ExBEHRT: Extended Transformer for Electronic Health Records to Predict Disease Subtypes & Progressions

Maurice Rupp, Oriane Peter, Thirupathi Pattipaka

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments ICLR 2023 Workshop on Trustworthy Machine Learning for Healthcare (Website: accepted-papers" target="_blank" rel="noopener">https://sites.google.com/view/tml4h2023/accepted-papers )

Journal ref Lecture Notes in Computer Science, vol 13932. Springer, Cham 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04828 2023-08-10 cs.CV 57%

Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts Learning

Qiang Wang, Junlong Du, Ke Yan, Shouhong Ding

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01217 2023-08-03 cs.CV 57%

TeachCLIP: Multi-Grained Teaching for Efficient Text-to-Video Retrieval

Kaibin Tian, Ruixiang Zhao, Hu Hu, Runquan Xie, Fengzong Lian, Zhanhui Kang, Xirong Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04224 2023-08-02 cs.CV 57%

Visual Causal Scene Refinement for Video Question Answering

Yushen Wei, Yang Liu, Hong Yan, Guanbin Li, Liang Lin

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14392 2023-07-28 cs.CV 57%

Human-centric Scene Understanding for 3D Large-scale Scenarios

Yiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen, Yuenan Hou, Xinge Zhu, Xuming He, Jingyi Yu, Yuexin Ma

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06524 2023-07-26 cs.CV 57%

SST: Real-time End-to-end Monocular 3D Reconstruction via Sparse Spatial-Temporal Guidance

Chenyangguang Zhang, Zhiqiang Lou, Yan Di, Federico Tombari, Xiangyang Ji

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments ICME 2023 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12058 2023-07-25 cs.CV 57%

Discovering Spatio-Temporal Rationales for Video Question Answering

Yicong Li, Junbin Xiao, Chun Feng, Xiang Wang, Tat-Seng Chua

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08914 2023-07-25 cs.CV 57%

MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge

Wei Lin, Leonid Karlinsky, Nina Shvetsova, Horst Possegger, Mateusz Kozinski, Rameswar Panda, Rogerio Feris, Hilde Kuehne, Horst Bischof

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09972 2023-07-20 cs.CV 57%

Fine-grained Text-Video Retrieval with Frozen Image Encoders

Zuozhuo Dai, Fangtao Shao, Qingkun Su, Zilong Dong, Siyu Zhu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02307 2023-07-20 cs.CV 57%

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09356 2023-07-19 cs.CV 57%

OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation

Dongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang, Jianbing Shen

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2023. The code is at https://github.com/wudongming97/OnlineRefer

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02906 2023-07-07 cs.LG cs.CV eess.SP 57%

A Real-time Human Pose Estimation Approach for Optimal Sensor Placement in Sensor-based Human Activity Recognition

Orhan Konak, Alexander Wischmann, Robin van de Water, Bert Arnrich

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13285 2023-06-26 cs.CV eess.IV 57%

Learning Scene Flow With Skeleton Guidance For 3D Action Recognition

Vasileios Magoulianitis, Athanasios Psaltis

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 18 pages, 3 figures, 3 tables, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12953 2023-06-26 cs.CV 57%

Enhancing Next Active Object-based Egocentric Action Anticipation with Guided Attention

Sanket Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to IEEE ICIP 2023, see project page here : https://sanketsans.github.io/guided-attention-egocentric.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06423 2023-06-13 cs.RO cs.AI 57%

Bayesian and Neural Inference on LSTM-based Object Recognition from Tactile and Kinesthetic Information

Francisco Pastor, Jorge García-González, Juan M. Gandarias, Daniel Medina, Pau Closas, Alfonso J. García-Cerezo, Jesús M. Gómez-de-Gabriel

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05241 2023-06-09 cs.MM 57%

Two Heads Are Better Than One: Improving Fake News Video Detection by Correlating with Neighbors

Peng Qi, Yuyang Zhao, Yufeng Shen, Wei Ji, Juan Cao, Tat-Seng Chua

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

Comments To appear in ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13883 2023-06-08 cs.AI 57%

MLink: Linking Black-Box Models from Multiple Domains for Collaborative Inference

Mu Yuan, Lan Zhang, Zimu Zheng, Yi-Nan Zhang, Xiang-Yang Li

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01656 2023-06-05 cs.CV cs.HC 57%

Backchannel Detection and Agreement Estimation from Video with Transformer Networks

Ahmed Amer, Chirag Bhuvaneshwara, Gowtham K. Addluri, Mohammed M. Shaik, Vedant Bonde, Philipp Müller

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted at IEEE IJCNN'23

详情

展开后加载摘要…

URL PDF HTML 收藏