arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2507.02626 2025-07-04 cs.MM 79%

VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning

Siran Chen, Boyu Chen, Chenyun Yu, Yuxiao Luo, Ouyang Yi, Lei Cheng, Chengxiang Zhuo, Zang Li, Yali Wang

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14171 2025-07-04 cs.CV 79%

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, Saining Xie

机构 * New York University(纽约大学) Yale University(耶鲁大学) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://vision-x-nyu.github.io/thinking-in-space.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17457 2025-06-24 cs.CV 79%

When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid Network

Dong Xiao, Guangyao Chen, Peixi Peng, Yangru Huang, Yifan Zhao, Yongxing Dai, Yonghong Tian

机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) Department of Software and Microelectronics, Peking University(软件与微电子系,北京大学) School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University(虚拟现实技术与系统国家重点实验室,北航软件学院) Peng Cheng Laboratory(鹏城实验室) Baidu Inc(百度公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICML 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16401 2025-06-23 cs.CY cs.CV 79%

TrajSceneLLM: A Multimodal Perspective on Semantic GPS Trajectory Analysis

Chunhou Ji, Qiumeng Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Under review for ACM SIGSPATIAL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12456 2025-06-23 cs.CV 79%

Demographics-Informed Neural Network for Multi-Modal Spatiotemporal forecasting of Urban Growth and Travel Patterns Using Satellite Imagery

Eugene Kofi Okrah Denteh, Andrews Danyo, Joshua Kofi Asamoah, Blessing Agyei Kyem, Armstrong Aboah

机构 * Department of Civil, Construction and Environmental Engineering, North Dakota State University(土木、建筑与环境工程系,北达科塔州立大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23623 2025-06-23 cs.CV 79%

On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection

Xiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li, Lei Bai, Xiaoming Liu, Guangtao Zhai, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) Michigan State University(密歇根州立大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 9 figures, published in NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00631 2025-06-17 cs.LG cs.AI 79%

TrialBench: Multi-Modal Artificial Intelligence-Ready Clinical Trial Datasets

Jintai Chen, Yaojun Hu, Mingchen Cai, Yingzhou Lu, Yue Wang, Xu Cao, Miao Lin, Hongxia Xu, Jian Wu, Cao Xiao, Jimeng Sun, Yuqiang Li, Lucas Glass, Kexin Huang, Marinka Zitnik, Tianfan Fu

机构 * AI Thrust, Information Hub, HKUST(GZ)(香港科技大学(广州)人工智能 thrust 与信息中心) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Medicine, Stanford University(斯坦福大学医学院) Computer Science Department, UIUC(伊利诺伊大学厄巴纳-香槟分校计算机科学系) Medical Big Data Center, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University(广东省人民医院(广东省医学科学院)医学大数据中心,南方医科大学) Innovation Institute for Artificial Intelligence in Medicine of Zhejiang University, College of Pharmaceutical Sciences, Zhejiang University(浙江大学人工智能医学创新研究院,浙江大学药学院) The Second Affiliated Hospital, Zhejiang University School of Medicine(浙江大学医学院附属第二医院) GE HealthCare(通用电气医疗) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) IQVIA Computer Science Department, Stanford University(斯坦福大学计算机科学系) Informatics, Harvard Medical School, Harvard University(哈佛医学院,哈佛大学informatics部门) State Key Laboratory for Novel Software Technology at Nanjing University, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室,南京大学计算机科学学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments accepted by Nature Scientific Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16998 2025-06-12 cs.CV 79%

Understanding Long Videos with Multimodal Language Models

Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya, Michael S. Ryoo

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages (main paper), 7 pages appendix. ICLR 2025 conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09345 2025-06-12 cs.CV 79%

An Effective End-to-End Solution for Multimodal Action Recognition

Songping Wang, Xiantao Hu, Yueming Lyu, Caifeng Shan

机构 * School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学) PCA-Lab, School of Computer Science and Engineering, Nanjing University of Science and Technology, China(PCA实验室,计算机科学与工程学院,南京理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18253 2025-06-10 cs.LG cs.AI q-bio.QM 79%

Multimodal Integration of Longitudinal Noninvasive Diagnostics for Survival Prediction in Immunotherapy Using Deep Learning

Melda Yeghaian, Zuhir Bodalal, Daan van den Broek, John B A G Haanen, Regina G H Beets-Tan, Stefano Trebeschi, Marcel A J van Gerven

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref Journal of the American Medical Informatics Association, 2025;, ocaf074

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06205 2025-06-09 cs.RO cs.AI 79%

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning

Sheng Chen, Peiyu He, Jiaxin Hu, Ziyang Liu, Yansheng Wang, Tao Xu, Chi Zhang, Chongchong Zhang, Chao An, Shiyu Cai, Duo Cao, Kangping Chen, Shuai Chu, Tianwei Chu, Mingdi Dan, Min Du, Weiwei Fang, Pengyou Fu, Junkai Hu, Xiaowei Jiang, Zhaodi Jiang, Fuxuan Li, Jun Li, Minghui Li, Mingyao Li, Yanchang Li, Zhibin Li, Guangming Liu, Kairui Liu, Lihao Liu, Weizhi Liu, Xiaoshun Liu, Yufei Liu, Yunfei Liu, Qiang Lu, Yuanfei Luo, Xiang Lv, Hongying Ma, Sai Ma, Lingxian Mi, Sha Sa, Hongxiang Shu, Lei Tian, Chengzhi Wang, Jiayu Wang, Kaijie Wang, Qingyi Wang, Renwen Wang, Tao Wang, Wei Wang, Xirui Wang, Chao Wei, Xuguang Wei, Zijun Xia, Zhaohao Xiao, Tingshuai Yan, Liyan Yang, Yifan Yang, Zhikai Yang, Zhong Yin, Li Yuan, Liuchun Yuan, Chi Zhang, Jinyang Zhang, Junhui Zhang, Linge Zhang, Zhenyi Zhang, Zheyu Zhang, Dongjie Zhu, Hang Li, Yangang Zhang

机构 * Full author list in Contributions(完整作者列表)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Astra Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02382 2025-06-04 cs.CV cs.LG 79%

Multi-level and Multi-modal Action Anticipation

Seulgi Kim, Ghazal Kaviani, Mohit Prabhushankar, Ghassan AlRegib

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in 2025 IEEE International Conference on Image Processing (ICIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13541 2025-06-04 cs.CV cs.LG cs.NE 79%

Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition

Jingyao Wang, Wenwen Qiang, Changwen Zheng, Fuchun Sun

机构 * University of Chinese Academy of Sciences(中国科学院大学) National Key Laboratory of Space Integrated Information System(空间信息集成系统国家重点实验室) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) Department of Computer Science and Technology(计算机科学与技术系)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17587 2025-06-03 cs.LG cs.AI 79%

Multimodal Banking Dataset: Understanding Client Needs through Event Sequences

Dzhambulat Mollaev, Alexander Kostin, Maria Postnova, Ivan Karpukhin, Ivan Kireev, Gleb Gusev, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00531 2025-06-03 cs.LG cs.AI 79%

M2WLLM: Multi-Modal Multi-Task Ultra-Short-term Wind Power Prediction Algorithm Based on Large Language Model

Hang Fana, Mingxuan Lib, Zuhan Zhanga, Long Chengc, Yujian Ye, Dunnan Liua

机构 * School of Economics and Management, North China Electric Power University(华北电力大学经济管理学院) Department of Electrical Engineering, Tsinghua University(清华大学电气工程系) School of Control and Computer Engineering, North China Electric Power University(华北电力大学控制与计算机工程学院) School of Electrical Engineering, Southeast University(东南大学电气工程学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09306 2025-06-03 cs.CV 79%

Keypoint-Integrated Instruction-Following Data Generation for Enhanced Human Pose and Action Understanding in Multimodal Models

Dewen Zhang, Wangpeng An, Hayaru Shouno

机构 * Department of Informatics, Graduate School of Informatics and Engineering, The University of Electro-Communications(信息学院、信息与工程研究生院、电通通信大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at the International Conference on Advanced Concepts for Intelligent Vision Systems (ACIVS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21374 2025-05-28 cs.CV 79%

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Junhao Cheng, Yuying Ge, Teng Wang, Yixiao Ge, Jing Liao, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG广告实验室) City University of Hong Kong(香港城市大学)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Homepage: https://github.com/TencentARC/Video-Holmes

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19125 2025-05-27 cs.CV 79%

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models

Yuqi Liu, Qin Jin, Tianyuan Qu, Xuan Liu, Yang Du, Bei Yu, Jiaya Jia

机构 * CUHK(香港中文大学) HKUST(香港理工大学) RUC(中国人民大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12746 2025-05-26 cs.AI 79%

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs

Haruka Asanuma, Naoko Koide-Majima, Ken Nakamura, Takato Horii, Shinji Nishimoto, Masafumi Oizumi

机构 * The University of Tokyo, Graduate School of Arts and Sciences(东京大学艺术与科学研究生院) Center for Information and Neural Networks (CiNet), National Institute of Information and Communications Technology(信息与神经网络中心(CiNet),信息与通信技术国家研究所) The University of Osaka, Graduate School of Frontier Biosciences(大阪大学前沿生命科学研究生院) The University of Tokyo, Faculty of Engineering(东京大学工学部) The University of Osaka, Graduate School of Engineering Science(大阪大学工学研究院) The University of Osaka, Graduate School of Medicine(大阪大学医学研究院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12007 2025-05-23 cs.CV 79%

Multi-modal Collaborative Optimization and Expansion Network for Event-assisted Single-eye Expression Recognition

Runduo Han, Xiuping Liu, Shangxuan Yi, Yi Zhang, Hongchen Tan

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14405 2025-05-21 cs.CV 79%

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

Jiafeng Liang, Shixin Jiang, Xuan Dong, Ning Wang, Zheng Chu, Hui Su, Jinlan Fu, Ming Liu, See-Kiong Ng, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Peng Cheng Laboratory(鹏城实验室) National University of Singapore(新加坡国立大学) Meituan Inc.(美团公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07158 2025-05-21 cs.LG cs.AI 79%

Early Risk Prediction of Pediatric Cardiac Arrest from Electronic Health Records via Multimodal Fused Transformer

Jiaying Lu, Stephanie R. Brown, Songyuan Liu, Shifan Zhao, Kejun Dong, Del Bold, Michael Fundora, Alaa Aljiffry, Alex Fedorov, Jocelyn Grunwell, Xiao Hu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref in Proceedings of 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12667 2025-05-15 cs.RO cs.CV 79%

METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance

Ziang Guo, Xinhao Lin, Zakhar Yagudin, Artem Lykov, Yong Wang, Yanqiang Li, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(斯克尔科沃科学与技术研究所智能空间机器人实验室,数字工程中心) Institute of Automation, Qilu University of Technology (Shandong Academy of Sciences)(齐鲁科技大学自动化研究所,山东科学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICRA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08657 2025-05-14 cs.RO cs.AI 79%

A Comparative Study of Human Activity Recognition: Motion, Tactile, and multi-modal Approaches

Valerio Belcamino, Nhat Minh Dinh Le, Quan Khanh Luu, Alessandro Carfì, Van Anh Ho, Fulvio Mastrogiovanni

机构 * University of Genoa(热力学与系统工程系,基因瓦大学) The University of Danang–University of Science and Technology(丹那大学–科学与技术大学) Japan Advanced Institute of Science and Technology(日本先进科学研究院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06682 2025-05-13 eess.SP cs.AI 79%

A Short Overview of Multi-Modal Wi-Fi Sensing

Zijian Zhao

机构 * Department of Civil and Environmental Engineering(土木及环境工程系) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06637 2025-05-13 cs.AI 79%

Exploring Multimodal Foundation AI and Expert-in-the-Loop for Sustainable Management of Wild Salmon Fisheries in Indigenous Rivers

Chi Xu, Yili Jin, Sami Ma, Rongsheng Qian, Hao Fang, Jiangchuan Liu, Xue Liu, Edith C. H. Ngai, William I. Atlas, Katrina M. Connors, Mark A. Spoljaric

机构 * Simon Fraser University(西蒙弗雷泽大学) McGill University(麦吉尔大学) The University of Hong Kong(香港大学) Wild Salmon Center(野生鲑鱼中心) Pacific Salmon Foundation(太平洋鲑鱼基金会) Haida Fisheries Program(海达渔业计划)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, accepted by IJCAI 2025, AI and Social Good Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02393 2025-05-09 cs.CV 79%

Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection

Sungheon Jeong, Jihong Park, Mohsen Imani

机构 * University of California, Irvine(加州大学尔湾分校)

专题命中 视频多模态 :multimodal(title);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14423 2025-05-09 cs.CV 79%

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Zhuoman Liu, Weicai Ye, Yan Luximon, Pengfei Wan, Di Zhang

机构 * The Hong Kong Polytechnic University(香港理工大学) Kuaishou Technology(快手科技)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2025. Homepage: https://zhuomanliu.github.io/PhysFlow/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01973 2025-05-06 cs.CV 79%

Visual Dominance and Emerging Multimodal Approaches in Distracted Driving Detection: A Review of Machine Learning Techniques

Anthony Dontoh, Stephanie Ivey, Logan Sirbaugh, Andrews Danyo, Armstrong Aboah

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21248 2025-05-01 cs.CV 79%

Multi-modal Transfer Learning for Dynamic Facial Emotion Recognition in the Wild

Ezra Engel, Lishan Li, Chris Hudy, Robert Schleusner

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏