arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4726 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4726 篇

2504.10921 2025-04-28 cs.IR 78%

MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender Systems

Yibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang, Xing Xu, Yang Yang

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14996 2025-04-08 eess.SY cs.RO cs.SY 78%

EDRF: Enhanced Driving Risk Field Based on Multimodal Trajectory Prediction and Its Applications

Junkai Jiang, Zeyu Han, Yuning Wang, Mengchi Cai, Qingwen Meng, Qing Xu, Jianqiang Wang

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03423 2025-04-07 cs.LG cs.RO 78%

DML-RAM: Deep Multimodal Learning Framework for Robotic Arm Manipulation using Pre-trained Models

Sathish Kumar, Swaroop Damodaran, Naveen Kumar Kuruba, Sumit Jha, Arvind Ramanathan

专题命中 视频多模态 :multimodal(title,abstract)

Comments 7 pages , 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13683 2025-03-12 cs.RO 78%

PrefMMT: Modeling Human Preferences in Preference-based Reinforcement Learning with Multimodal Transformers

Dezhong Zhao, Ruiqi Wang, Dayoon Suh, Taehyeon Kim, Ziqin Yuan, Byung-Cheol Min, Guohua Chen

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11144 2025-02-19 cs.HC 78%

CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding

Xingyu "Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang 'Anthony' Chen, Amy Pavel

专题命中 视频多模态 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13464 2025-01-24 eess.SP 78%

Deep Multi-modal Neural Receiver for 6G Vehicular Communication

Osama Saleem, Mohammed Alfaqawi, Pierre Merdrignac, Abdelaziz Bensrhair, Soheyb Ribouh

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08486 2025-01-23 cs.HC 78%

Can AI Prompt Humans? Multimodal Agents Prompt Players' Game Actions and Show Consequences to Raise Sustainability Awareness

Qinshi Zhang, Ruoyu Wen, Latisha Besariani Hendra, Zijian Ding, Ray LC

专题命中 视频多模态 :multimodal(title,abstract)

Comments 25 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00771 2025-01-13 cs.ET 78%

Resistive memory-based zero-shot liquid state machine for multimodal event data learning

Ning Lin, Shaocong Wang, Yi Li, Bo Wang, Shuhui Shi, Yangu He, Woyu Zhang, Yifei Yu, Yue Zhang, Xinyuan Zhang, Kwunhang Wong, Songqi Wang, Xiaoming Chen, Hao Jiang, Xumeng Zhang, Peng Lin, Xiaoxin Xu, Xiaojuan Qi, Zhongrui Wang, Dashan Shang, Qi Liu, Ming Liu

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.12410 2024-11-26 cs.LG eess.SP stat.ML 78%

Deep sr-DDL: Deep Structurally Regularized Dynamic Dictionary Learning to Integrate Multimodal and Dynamic Functional Connectomics data for Multidimensional Clinical Characterizations

Niharika Shimona D'Souza, Mary Beth Nebel, Deana Crocetti, Nicholas Wymbs, Joshua Robinson, Stewart H. Mostofsky, Archana Venkataraman

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.01931 2024-11-25 cs.LG eess.SP stat.ML 78%

A Deep-Generative Hybrid Model to Integrate Multimodal and Dynamic Connectivity for Predicting Spectrum-Level Deficits in Autism

Niharika Shimona D'Souza, Mary Beth Nebel, Deana Crocetti, Nicholas Wymbs, Joshua Robinson, Stewart Mostofsky, Archana Venkataraman

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09577 2024-11-19 cs.HC 78%

SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas

Yu-Kai Hung, Yun-Chien Huang, Ting-Yu Su, Yen-Ting Lin, Lung-Pan Cheng, Bryan Wang, Shao-Hua Sun

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19760 2024-10-29 cs.CV cs.AI cs.MM eess.IV 78%

Movie Trailer Genre Classification Using Multimodal Pretrained Features

Serkan Sulun, Paula Viana, Matthew E. P. Davies

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

Journal ref Expert Systems with Applications 258 (2024) 125209

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08933 2024-10-28 cs.LG physics.ao-ph physics.comp-ph 78%

Multi-Modal Learning-based Reconstruction of High-Resolution Spatial Wind Speed Fields

Matteo Zambra, Nicolas Farrugia, Dorian Cazau, Alexandre Gensse, Ronan Fablet

专题命中 视频多模态 :multi-modal(title,abstract)

Comments 22 pages, 13 figures. This work is to be submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07238 2024-10-11 cs.HC 78%

vailá: Versatile Anarcho Integrated Liberation Ánalysis in Multimodal Toolbox

Paulo Roberto Pereira Santiago, Abel Gonçalves Chinaglia, Kira Flanagan, Bruno L. S. Bedo, Ligia Yumi Mochida, Juan Aceros, Aline Bononi, Guilherme Manna Cesar

专题命中 视频多模态 :multimodal(title,abstract)

Comments 21 pages, 13 figures, submitted to arXiv under cs.SE (Software Engineering)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13394 2024-08-27 cs.RO 78%

Towards Robust Perception for Assistive Robotics: An RGB-Event-LiDAR Dataset and Multi-Modal Detection Pipeline

Adam Scicluna, Cedric Le Gentil, Sheila Sutjipto, Gavin Paul

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Accepted to the 2024 IEEE International Conference on Automation Science and Engineering (CASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07791 2024-08-16 cs.MM cs.AI cs.CV cs.LG 78%

An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture

Tiancheng Shi, Yuanchen Wei, John R. Kender

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07488 2024-08-15 cs.HC 78%

Towards Enhanced Context Awareness with Vision-based Multimodal Interfaces

Yongquan Hu, Wen Hu, Aaron Quigley

专题命中 视频多模态 :multimodal(title,abstract)

Comments 3 pages, MOBILEHCI Adjunct '24 26th International Conference on Mobile Human-Computer Interaction, September 30-October 3, 2024, Melbourne, VIC, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05255 2024-08-07 cs.HC 78%

Emolysis: A Multimodal Open-Source Group Emotion Analysis and Visualization Toolkit

Shreya Ghosh, Zhixi Cai, Parul Gupta, Garima Sharma, Abhinav Dhall, Munawar Hayat, Tom Gedeon

专题命中 视频多模态 :multimodal(title,abstract)

Comments Accepted by ACII Demo 2024. Both Shreya Ghosh and Zhixi Cai contributed equally to this research

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17590 2024-06-26 cs.MM cs.AI cs.CV 78%

Multimodal Chaptering for Long-Form TV Newscast Video

Khalil Guetari, Yannis Tevissen, Frédéric Petitpont

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13807 2024-06-24 cs.CV cs.AI cs.CL 78%

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Alessandro Suglia, Claudio Greco, Katie Baker, Jose L. Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, Oliver Lemon

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments Code available https://github.com/alanaai/EVUD

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18924 2024-06-07 cs.LG 78%

Remaining useful life prediction of Lithium-ion batteries using spatio-temporal multimodal attention networks

Sungho Suh, Dhruv Aditya Mittal, Hymalai Bello, Bo Zhou, Mayank Shekhar Jha, Paul Lukowicz

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02761 2024-06-06 cs.CV cs.AI cs.LG cs.MM 78%

Multi-layer Learnable Attention Mask for Multimodal Tasks

Wayner Barrios, SouYoung Jin

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19373 2024-05-31 eess.SP cs.LG 78%

Multi-modal Mood Reader: Pre-trained Model Empowers Cross-Subject Emotion Recognition

Yihang Dong, Xuhang Chen, Yanyan Shen, Michael Kwok-Po Ng, Tao Qian, Shuqiang Wang

专题命中 视频多模态 :multi-modal(title);multimodal(abstract)

Comments Accepted by International Conference on Neural Computing for Advanced Applications, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04012 2024-04-03 cs.LG 78%

Temporal Cross-Attention for Dynamic Embedding and Tokenization of Multimodal Electronic Health Records

Yingbo Ma, Suraj Kolla, Dhruv Kaliraman, Victoria Nolan, Zhenhong Hu, Ziyuan Guan, Yuanfang Ren, Brooke Armfield, Tezcan Ozrazgat-Baslanti, Tyler J. Loftus, Parisa Rashidi, Azra Bihorac, Benjamin Shickel

专题命中 视频多模态 :multimodal(title,abstract)

Comments ICLR 2024 Workshop on Learning From Time Series for Health. 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06888 2024-03-14 physics.data-an cs.LG physics.app-ph 78%

Process signature-driven high spatio-temporal resolution alignment of multimodal data

Abhishek Hanchate, Himanshu Balhara, Vishal S. Chindepalli, Satish T. S. Bukkapatnam

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17788 2024-02-29 eess.SP cs.LG 78%

Multimodal Sleep Apnea Detection with Missing or Noisy Modalities

Hamed Fayyaz, Abigail Strang, Niharika S. D'Souza, Rahmatollah Beheshti

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14199 2024-02-07 cs.LG econ.GN q-fin.EC q-fin.TR 78%

MTRGL:Effective Temporal Correlation Discerning through Multi-modal Temporal Relational Graph Learning

Junwei Su, Shan Wu, Jinhui Li

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08359 2024-01-26 cs.RO 78%

Haptics-Enabled Forceps with Multi-Modal Force Sensing: Towards Task-Autonomous Surgery

Tangyou Liu, Tinghua Zhang, Jay Katupitiya, Jiaole Wang, Liao Wu

专题命中 视频多模态 :multi-modal(title,abstract)

Comments 12 pages, 9 figures, accepted by T-MECH

Journal ref IEEE/ASME Transactions on Mechatronics. 2023:1-12

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09388 2024-01-18 cs.RO 78%

CognitiveDog: Large Multimodal Model Based System to Translate Vision and Language into Action of Quadruped Robot

Artem Lykov, Mikhail Litvinov, Mikhail Konenkov, Rinat Prochii, Nikita Burtsev, Ali Alridha Abdulkarim, Artem Bazhenov, Vladimir Berman, Dzmitry Tsetserukou

专题命中 视频多模态 :multimodal(title);multi-modal(abstract)

Comments This paper has been accepted for publication at the HRI2024 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07601 2023-12-14 eess.SP cs.LG 78%

Non-contact Multimodal Indoor Human Monitoring Systems: A Survey

Le Ngu Nguyen, Praneeth Susarla, Anirban Mukherjee, Manuel Lage Cañellas, Constantino Álvarez Casado, Xiaoting Wu, Olli~Silvén, Dinesh Babu Jayagopi, Miguel Bordallo López

专题命中 视频多模态 :multimodal(title,abstract)

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏