arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2402.11403 2024-03-05 cs.AI 79%

An Empirical Evaluation of Neural and Neuro-symbolic Approaches to Real-time Multimodal Complex Event Detection

Liying Han, Mani B. Srivastava

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00366 2024-03-04 cs.CY cs.CV 79%

Exploring the dynamic interplay of cognitive load and emotional arousal by using multimodal measurements: Correlation of pupil diameter and emotional arousal in emotionally engaging tasks

C. Kosel, S. Michel, T. Seidel, M. Foerster

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments The first two authors contributed equally to the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13270 2024-02-22 physics.ao-ph cs.AI cs.LG physics.data-an 79%

Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model

Xinyu Wang, Kang Chen, Lei Liu, Tao Han, Bin Li, Lei Bai

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13146 2024-02-21 cs.CV 79%

OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog

Adnen Abdessaied, Manuel von Hochmeister, Andreas Bulling

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12845 2024-02-21 cs.AI cs.GT 79%

MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces

Tianyu Zheng, Ge Zhang, Xingwei Qu, Ming Kuang, Stephen W. Huang, Zhaofeng He

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11875 2024-02-20 cs.CL 79%

M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation

Hongcheng Liu, Pingjie Wang, Yu Wang, Yanfeng Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08607 2024-02-16 cs.CY cs.CV cs.LG 79%

Monitoring of Urban Changes with multi-modal Sentinel 1 and 2 Data in Mariupol, Ukraine, in 2022/23

Georg Zitzlsberger, Michal Podhoranyi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted for publication in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00965 2024-02-05 cs.LG cs.CV eess.SP 79%

Multi-Modal Machine Learning Framework for Automated Seizure Detection in Laboratory Rats

Aaron Mullen, Samuel E. Armstrong, Jasmine Perdeh, Bjorn Bauer, Jeffrey Talbert, V. K. Cody Bumgardner

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09057 2024-01-30 cs.CV 79%

CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding

Yunze Liu, Changxi Chen, Zifan Wang, Li Yi

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Journal ref ICRA2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11649 2024-01-23 cs.CV 79%

M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition

Mengmeng Wang, Jiazheng Xing, Boyuan Jiang, Jun Chen, Jianbiao Mei, Xingxing Zuo, Guang Dai, Jingdong Wang, Yong Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01061 2024-01-22 cs.CV 79%

Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation

Chen Liang, Yu Wu, Tianfei Zhou, Wenguan Wang, Zongxin Yang, Yunchao Wei, Yi Yang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Champion solution in YouTube-VOS 2021 Track 3. Extended version published in https://ieeexplore.ieee.org/abstract/document/10083244

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08345 2024-01-17 cs.CV 79%

Multi-view Distillation based on Multi-modal Fusion for Few-shot Action Recognition(CLIP-$\mathrm{M^2}$DF)

Fei Guo, YiKang Wang, Han Qi, WenPing Jin, Li Zhu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08043 2024-01-17 cs.RO cs.CV 79%

Cross-Modal Semi-Dense 6-DoF Tracking of an Event Camera in Challenging Conditions

Yi-Fan Zuo, Wanting Xu, Xia Wang, Yifu Wang, Laurent Kneip

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments accepted by IEEE Transactions on Robotics (T-RO). arXiv admin note: text overlap with arXiv:2202.02556

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07218 2024-01-17 cs.CV 79%

Self-supervised Event-based Monocular Depth Estimation using Cross-modal Consistency

Junyu Zhu, Lina Liu, Bofeng Jiang, Feng Wen, Hongbo Zhang, Wanlong Li, Yong Liu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by IROS2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06942 2024-01-05 cs.CV 79%

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, Yu Qiao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Data and Code: https://github.com/OpenGVLab/InternVideo/tree/main/Data/InternVid

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00988 2024-01-03 cs.CV 79%

Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models

Xinpeng Ding, Jinahua Han, Hang Xu, Xiaodan Liang, Wei Zhang, Xiaomeng Li

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15646 2023-12-29 cs.CY cs.AI econ.GN q-fin.EC 79%

A graph-based multimodal framework to predict gentrification

Javad Eshtiyagh, Baotong Zhang, Yujing Sun, Linhui Wu, Zhao Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref International Conference on Urban Informatics 2023 - Best Paper Award 3rd Place

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13633 2023-12-22 cs.CV 79%

Multi-Modal Domain Adaptation Across Video Scenes for Temporal Video Grounding

Haifeng Huang, Yang Zhao, Zehan Wang, Yan Xia, Zhou Zhao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11935 2023-12-20 cs.AI 79%

Parameterized Decision-making with Multi-modal Perception for Autonomous Driving

Yuyang Xia, Shuncheng Liu, Quanlin Yu, Liwei Deng, You Zhang, Han Su, Kai Zheng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments IEEE International Conference on Data Engineering (ICDE2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03056 2023-12-20 cs.CV 79%

MOISST: Multimodal Optimization of Implicit Scene for SpatioTemporal calibration

Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at IROS2023 Project site: https://qherau.github.io/MOISST/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07378 2023-12-13 cs.CV 79%

X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-modal Knowledge Transfer

Linglin Jing, Ying Xue, Xu Yan, Chaoda Zheng, Dong Wang, Ruimao Zhang, Zhigang Wang, Hui Fang, Bin Zhao, Zhen Li

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19062 2023-11-28 cs.RO cs.AI 79%

A multi-modal table tennis robot system

Andreas Ziegler, Thomas Gossard, Karl Vetter, Jonas Tebbe, Andreas Zell

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments Accepted for RoboLetics: Workshop on Robot Learning in Athletics @CoRL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12344 2023-11-22 cs.CV 79%

Modality Mixer Exploiting Complementary Information for Multi-modal Action Recognition

Sumin Lee, Sangmin Woo, Muhammad Adi Nugroho, Changick Kim

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2208.11314

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07608 2023-11-15 cs.LG cs.AI 79%

MuST: Multimodal Spatiotemporal Graph-Transformer for Hospital Readmission Prediction

Yan Miao, Lequan Yu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20201 2023-11-01 cs.CL 79%

Video-Helpful Multimodal Machine Translation

Yihang Li, Shuichiro Shimizu, Chenhui Chu, Sadao Kurohashi, Wei Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2023 Main Conference (long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16590 2023-10-26 cs.CV 79%

$\mathbb{VD}$-$\mathbb{GR}$: Boosting $\mathbb{V}$isual $\mathbb{D}$ialog with Cascaded Spatial-Temporal Multi-Modal $\mathbb{GR}$aphs

Adnen Abdessaied, Lei Shi, Andreas Bulling

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15670 2023-10-25 cs.CV 79%

Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection

Linyan Huang, Zhiqi Li, Chonghao Sima, Wenhai Wang, Jingdong Wang, Yu Qiao, Hongyang Li

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11169 2023-10-18 cs.LG cs.AI 79%

MST-GAT: A Multimodal Spatial-Temporal Graph Attention Network for Time Series Anomaly Detection

Chaoyue Ding, Shiliang Sun, Jing Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Information Fusion 2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09222 2023-10-13 cs.CV cs.LG 79%

MMTSA: Multimodal Temporal Segment Attention Network for Efficient Human Activity Recognition

Ziqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing, Shwetak Patel, Xin Liu, Yuanchun Shi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04991 2023-10-12 cs.CV 79%

Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling

Haogeng Liu, Qihang Fan, Tingkai Liu, Linjie Yang, Yunzhe Tao, Huaibo Huang, Ran He, Hongxia Yang

专题命中 视频多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏