arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4726 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4726 篇

2509.26636 2025-10-01 cs.LG 78%

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying, Haoyu Zhao, Runing Yang, Ming Jin, Boyi Li, Marco Pavone, Serena Yeung-Levy, Jun Wang, Dawn Song, Costas Spanos

机构 * UC Berkeley(伯克利大学) Stanford(斯坦福大学) UCL(伦敦大学学院) Virginia Tech(弗吉尼亚理工学院) Nvidia(英伟达公司)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21346 2025-09-29 cs.NE cs.LG q-bio.BM 78%

Spiking Neural Networks for Mental Workload Classification with a Multimodal Approach

Jiahui An, Sara Irina Fabrikant, Giacomo Indiveri, Elisa Donati

机构 * Institute of Neuroinformatics, University of Zurich(神经信息学研究所,苏黎世大学) ETH Zurich(苏黎世联邦理工学院) Digital Society Initiative, University of Zurich(数字社会倡议,苏黎世大学) Department of Geography(地理系)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18457 2025-09-24 cs.LG 78%

GluMind: Multimodal Parallel Attention and Knowledge Retention for Robust Cross-Population Blood Glucose Forecasting

Ebrahim Farahmand, Reza Rahimi Azghan, Nooshin Taheri Chatrudi, Velarie Yaa Ansu-Baidoo, Eric Kim, Gautham Krishna Gudur, Mohit Malu, Owen Krueger, Edison Thomaz, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh

机构 * Arizona State University(亚利桑那州立大学) University of Miami(迈阿密大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18020 2025-09-23 cs.HC 78%

ClassMind: Scaling Classroom Observation and Instructional Feedback with Multimodal AI

Ao Qu, Yuxi Wen, Jiayi Zhang, Yunge Wen, Yibo Zhao, Alok Prakash, Andrés F. Salazar-Gómez, Paul Pu Liang, Jinhua Zhao

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17532 2025-09-23 cs.DC 78%

TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation

Guanxiong Sun, Majid Mirmehdi, Zahraa Abdallah, Raul Santos-Rodriguez, Ian Craddock, Telmo de Menezes e Silva Filho

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04603 2025-09-23 cs.RO 78%

Diffusion-Based Approximate MPC: Fast and Consistent Imitation of Multi-Modal Action Distributions

Pau Marquez Julbe, Julian Nubert, Henrik Hose, Sebastian Trimpe, Katherine J. Kuchenbecker

机构 * Max Planck ETH CLS(马克斯·普朗克-ETH CLS) German Research Foundation(德国研究基金会) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ETH Zürich(苏黎世联邦理工学院) Institute for Data Science in Mechanical Engineering (DSME)(机械工程数据科学研究所) RWTH Aachen University(亚琛工业大学)

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08578 2025-09-22 cs.LG q-bio.PE q-bio.QM 78%

Multi-modal Adaptive Estimation for Temporal Respiratory Disease Outbreak

Hong Liu, Kerui Cen, Yanxing Chen, Zige Liu, Dong Chen, Zifeng Yang, Chitin Hon

机构 * Respiratory Disease AI Laboratory in Epidemic Intelligence and Applications of Medical Big Data Instruments, Macau University of Science and Technology(呼吸疾病人工智能实验室(流行病智能与医学大数据应用)) Faculty of Innovation Engineering, Macau University of Science and Technology(创新工程学院) Institute of Systems Engineering, Macau University of Science and Technology(系统工程研究所) School of Business, Macau University of Science and Technology(商学院) State Key Laboratory of Respiratory Disease, National Clinical Research Center for Respiratory Disease, Guangzhou Institute of Respiratory Health, The First Affiliated Hospital of Guangzhou Medical University(呼吸疾病国家重点实验室、呼吸疾病临床研究中心、广州呼吸健康研究院、广州医学院第一附属医院) Guangzhou National Laboratory(广州国家实验室)

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15233 2025-09-22 cs.MM cs.CL cs.CV 78%

Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents

Xueqiao Zhang, Chao Zhang, Jingtao Xu, Yifan Zhu, Xin Shi, Yi Yang, Yawei Luo

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at EMNLP2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15864 2025-09-18 cs.RO 78%

FlowAct: A Proactive Multimodal Human-robot Interaction System with Continuous Flow of Perception and Modular Action Sub-systems

Timothée Dhaussy, Bassam Jabaian, Fabrice Lefèvre

机构 * Laboratoire Informatique d'Avignon, Avignon University, France(阿维尼翁信息实验室,阿维尼翁大学,法国)

专题命中 视频多模态 :multimodal(title,abstract)

Comments Paper accepted at ICPRAM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05898 2025-09-09 cs.HC 78%

Attention, Action, and Memory: How Multi-modal Interfaces and Cognitive Load Alter Information Retention

Omar Elgohary, Zhu-Tien

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04714 2025-09-08 cs.SI 78%

ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings

Wajiha Naveed, Zartash Afzal Uzmi, Zafar Ayyub Qazi

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04330 2025-09-05 cs.IR 78%

Temporal Interest-Driven Multimodal Personalized Content Generation

Tian Miao

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04210 2025-09-05 cs.CE cs.LG 78%

COBRA: Multimodal Sensing Deep Learning Framework for Remote Chronic Obesity Management via Wrist-Worn Activity Monitoring

Zhengyang Shen, Bo Gao, Mayue Shi

机构 * Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ, UK(帝国理工学院电子与电气工程系) Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford, Oxford OX3 7DQ, UK(牛津大学生物医学工程研究所)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 19 pages, 4 figures. *Correspondence: m.shi16@imperial.ac.uk. Accepted by the IUPESM World Congress on Medical Physics and Biomedical Engineering 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19346 2025-09-05 cs.LG 78%

Short-Form Video Recommendations with Multimodal Embeddings: Addressing Cold-Start and Bias Challenges

Andrii Dzhoha, Katya Mirylenka, Egor Malykh, Marco-Andrea Buchmann, Francesca Catino

机构 * Zalando SE Berlin Germany(泽尔安多德国分公司) Zalando Switzerland AG Zürich Switzerland(泽尔安多瑞士分公司)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19567 2025-08-28 cs.LG 78%

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning

Sheryl Mathew, N Harshit

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11537 2025-08-18 cs.RO 78%

MultiPark: Multimodal Parking Transformer with Next-Segment Prediction

Han Zheng, Zikang Zhou, Guli Zhang, Zhepei Wang, Kaixuan Wang, Peiliang Li, Shaojie Shen, Ming Yang, Tong Qin

机构 * Shanghai Jiao Tong University(上海交通大学) Zhuoyu Technology, Co., Ltd.(珠海宇科技有限公司) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(香港理工大学电子与计算机工程系)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08541 2025-08-13 physics.app-ph 78%

Multimodal learning enables instant ionizing radiation alerts on unmodified mobile phones for real-world emergency response

Yanfeng Xie, Xingzhi Cheng

专题命中 视频多模态 :multimodal(title,abstract)

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05695 2025-08-11 cs.CR cs.LG 78%

MambaITD: An Efficient Cross-Modal Mamba Network for Insider Threat Detection

Kaichuan Kong, Dongjie Liu, Xiaobo Jin, Zhiying Li, Guanggang Geng, Jian Weng

机构 * College of Cyber Security(网络安全学院) Jinan University(济南大学) Department of Electrical and Electronic Engineering(电子与电气工程系) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 视频多模态 :cross-modal(title,abstract)

Comments Submitted to the 2025 IEEE International Conference on Data Mining (ICDM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00210 2025-08-04 stat.CO stat.ME 78%

Efficient rare event estimation for multimodal and high-dimensional system reliability via subset adaptive importance sampling

Sara Helal, Victor Elvira

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17252 2025-07-24 cs.LG math.OC 78%

A Coalition Game for On-demand Multi-modal 3D Automated Delivery System

Farzan Moosavi, Bilal Farooq

机构 * Laboratory of Innovations in Transportation (LiTrans), Toronto Metropolitan University, Toronto, Canada(创新交通实验室(LiTrans),多伦多 Metropolitan 大学,多伦多,加拿大)

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14163 2025-07-22 eess.SP cs.LG stat.ML 78%

UniPhyNet: A Unified Network For Multimodal Physiological Raw Signal Classification

Renxiang Qiu, Raghavendra Selvan

机构 * Department of Computer Science, University of Copenhagen(计算机科学系,哥本哈根大学)

专题命中 视频多模态 :multimodal(title,abstract)

Comments Accepted to be presented at the 35th IEEE International Workshop on Machine Learning for Signal Processing (IEEE MLSP 2025). Source code available at https://github.com/HughYau/UniPhyNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14146 2025-07-22 eess.SP 78%

Estimating Markers of Driving Stress through Multimodal Physiological Monitoring

Kleanthis Avramidis, Emily Zhou, Tiantian Feng, Hossein Hamidi Shishavan, Frederico Marcolino Quintao Severgnini, Danny J. Lohan, Paul Schmalenberg, Ercan M. Dede, Shrikanth Narayanan

专题命中 视频多模态 :multimodal(title,abstract)

Comments 11 pages, 7 figures, 3 tables. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06444 2025-07-17 cs.CE 78%

Eyes on the Road, Mind Beyond Vision: Context-Aware Multi-modal Enhanced Risk Anticipation

Jiaxun Zhang, Haicheng Liao, Yumu Xie, Chengyue Wang, Yanchen Guan, Bin Rao, Zhenning Li

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07804 2025-07-11 cs.LG 78%

Deep Survival Analysis in Multimodal Medical Data: A Parametric and Probabilistic Approach with Competing Risks

Alba Garrido, Alejandro Almodóvar, Patricia A. Apellániz, Juan Parras, Santiago Zazo

机构 * Information Processing and Telecommunications Center, ETSI Telecomunicación, Universidad Politécnica de Madrid, Spain(信息处理与电信中心,电信工程学院,马德里理工大学,西班牙)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 29 pages, 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18585 2025-06-27 gr-qc 78%

Multimodal signatures of asymptotic (A)dS Kalb-Ramond black holes: Constraints through the shadow, weak deflection angle, and topological photon spheres

Reggie C. Pantig, Ali Övgün

专题命中 视频多模态 :multimodal(title,abstract)

Comments 15 pages, 4 figures

Journal ref Annals of Physics 480 (2025) 170104

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18538 2025-05-27 eess.IV cs.LG 78%

Mind Your Vision: Multimodal Estimation of Refractive Disorders Using Electrooculography and Eye Tracking

Xin Wei, Huakun Liu, Yutaro Hirao, Monica Perusquia-Hernandez, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa

机构 * Nara Institute of Science and Technology(奈良科学技術大學) Kyushu University(九州大學)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18815 2025-05-27 cs.LG 78%

MissionGNN: Hierarchical Multimodal GNN-based Weakly Supervised Video Anomaly Recognition with Mission-Specific Knowledge Graph Generation

Sanggeon Yun, Ryozo Masukawa, Minhyoung Na, Mohsen Imani

专题命中 视频多模态 :multimodal(title,abstract)

Comments Accepted to WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11214 2025-05-19 cs.RO 78%

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

Wei Zhao, Gongsheng Li, Zhefei Gong, Pengxiang Ding, Han Zhao, Donglin Wang

机构 * Westlake University(西湖大学) Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18438 2025-05-08 cs.HC 78%

Adaptive Gen-AI Guidance in Virtual Reality: A Multimodal Exploration of Engagement in Neapolitan Pizza-Making

Ka Hei Carrie Lau, Sema Sen, Philipp Stark, Efe Bozkir, Enkelejda Kasneci

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01989 2025-05-06 cs.DS 78%

Exact Set Packing in Multimodal Transportation with Ridesharing System for First/Last Mile

Qian-Ping Gu, Jiajian Leo Liang

专题命中 视频多模态 :multimodal(title,abstract)

Comments 29 pages, 9 tables, 2 figures, and

详情

展开后加载摘要…

URL PDF HTML 收藏