arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2504.19669 2025-04-29 cs.CL 79%

Multimodal Conditioned Diffusive Time Series Forecasting

Chen Su, Yuanhe Tian, Yan Song

机构 * University of Science and Technology of China(中国科学技术大学) University of Washington(华盛顿大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17049 2025-04-29 cs.AI cs.LG 79%

TabulaTime: A Novel Multimodal Deep Learning Framework for Advancing Acute Coronary Syndrome Prediction through Environmental and Clinical Data Integration

Xin Zhang, Liangxiu Han, Stephen White, Saad Hassan, Philip A Kalra, James Ritchie, Carl Diver, Jennie Shorley

机构 * Department of Computing, and Mathematics, Manchester Metropolitan University(计算机与数学系,曼彻斯特 Metropolitan 大学) Blackpool Teaching Hospitals NHS Foundation Trust(布莱克浦教学医院 NHS 基础信托) Northern Care Alliance, NHS Foundation Trust(北方护理联盟,NHS 基础信托) Department of Engineering, Manchester Metropolitan University(工程系,曼彻斯特 Metropolitan 大学) Faculty of Business and Law, Manchester Metropolitan University(商业与法律学院,曼彻斯特 Metropolitan 大学) Faculty of Medical Sciences, Newcastle University(医学科学学院,纽卡斯尔大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18152 2025-04-28 cs.CV 79%

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Yi-Xing Peng, Qize Yang, Yu-Ming Tang, Shenghao Fu, Kun-Yu Lin, Xihan Wei, Wei-Shi Zheng

机构 * Sun Yat-sen University, China(中山大学) Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China(教育部人工智能与先进计算重点实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12542 2025-04-24 cs.CV 79%

ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Peiran Wu, Yunze Liu, Miao Liu, Junxiao Shen

机构 * University of Bristol(布里斯托大学) X-Intelligence Labs(X-智能实验室) Meta

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15599 2025-04-23 cs.CV cs.LG 79%

Multi-Modal Fusion of In-Situ Video Data and Process Parameters for Online Forecasting of Cookie Drying Readiness

Shichen Li, Chenhui Shao

机构 * Department of Mechanical Science and Engineering, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校机械科学与工程系) Department of Mechanical Engineering, University of Michigan(密歇根大学机械工程系)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13710 2025-04-21 cs.CV 79%

Few-Shot Referring Video Single- and Multi-Object Segmentation via Cross-Modal Affinity with Instance Sequence Matching

Heng Liu, Guanghui Li, Mingqi Gao, Xiantong Zhen, Feng Zheng, Yang Wang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11779 2025-04-17 cs.CV 79%

Multimodal Spatio-temporal Graph Learning for Alignment-free RGBT Video Object Detection

Qishun Wang, Zhengzheng Tu, Chenglong Li, Bo Jiang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12922 2025-04-16 cs.CV 79%

Contextual AD Narration with Interleaved Multimodal Sequence

Hanlin Wang, Zhan Tong, Kecheng Zheng, Yujun Shen, Limin Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR25

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10443 2025-04-15 cs.CV cs.AI cs.CL cs.LG cs.MM 79%

Multimodal Long Video Modeling Based on Temporal Dynamic Context

Haoran Hao, Jiaming Han, Yiyuan Zhang, Xiangyu Yue

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10222 2025-04-15 cs.MM 79%

PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search

Pengfei Hu, Zhenrong Zhang, Qikai Chang, Shuhang Liu, Jiefeng Ma, Jun Du, Jianshu Zhang, Quan Liu, Jianqing Gao, Feng Ma, Qingfeng Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12499 2025-04-15 cs.CV 79%

End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting

Yongqi Wang, Xinxiao Wu, Shuo Yang, Jiebo Luo

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08751 2025-04-15 cs.IR cs.AI cs.CR 79%

Research on the Design of a Short Video Recommendation System Based on Multimodal Information and Differential Privacy

Haowei Yang, Lei Fu, Qingyi Lu, Yue Fan, Tianle Zhang, Ruohan Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06272 2025-04-10 cs.IR cs.AI 79%

RAVEN: An Agentic Framework for Multimodal Entity Discovery from Large-Scale Video Collections

Kevin Dela Rosa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Presented at AI Agent for Information Retrieval: Generating and Ranking (Agent4IR) @ AAAI 2025 [https://sites.google.com/view/ai4ir/aaai-2025]

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08424 2025-04-09 cs.RO cs.AI cs.HC 79%

Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task

Hassan Ali, Philipp Allgeuer, Stefan Wermter

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published in the Proceedings of the 16th International Conference on Social Robotics (ICSR) 2024,15 pages,5 figures,2 tables; work was co-funded by Horizon Europe project TERAIS under Grant agreement number 101079338

Journal ref In: Palinko, O., et al. Social Robotics. ICSR + AI 2024. vol 15563. Springer (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05178 2025-04-08 cs.CV 79%

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

Hao Fang, Runmin Cong, Xiankai Lu, Zhiyang Chen, Wei Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02279 2025-04-08 cs.CV 79%

MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition

Trung Thanh Nguyen, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments This is a part of article arXiv:2504.02287

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01328 2025-04-03 cs.CV 79%

Slow-Fast Architecture for Video Multi-Modal Large Language Models

Min Shi, Shihao Wang, Chieh-Yun Chen, Jitesh Jain, Kai Wang, Junjun Xiong, Guilin Liu, Zhiding Yu, Humphrey Shi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12787 2025-04-02 cs.HC cs.AI 79%

GameVibe: A Multimodal Affective Game Corpus

Matthew Barthet, Maria Kaselimi, Kosmas Pinitas, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 12 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22952 2025-04-01 cs.CV 79%

OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu, Dongyan Zhao, Zilong Zheng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments To appear at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14945 2025-03-27 cs.CV 79%

Generating Multimodal Driving Scenes via Next-Scene Prediction

Yanhao Wu, Haoyang Zhang, Tianwei Lin, Lichao Huang, Shujie Luo, Rui Wu, Congpei Qiu, Wei Ke, Tong Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04923 2025-03-26 cs.CV 79%

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao, Eric Xing, Fahad Shahbaz Khan, Salman Khan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Technical Report of VideoGLaMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18933 2025-03-25 cs.CV 79%

SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction

Enrico Pallotta, Sina Mokhtarzadeh Azar, Shuai Li, Olga Zatsarynna, Juergen Gall

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17827 2025-03-25 cs.CV 79%

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

Wenxuan Zhu, Bing Li, Cheng Zheng, Jinjie Mai, Jun Chen, Letian Jiang, Abdullah Hamdi, Sara Rojas Martinez, Chia-Wen Lin, Mohamed Elhoseiny, Bernard Ghanem

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16466 2025-03-24 cs.HC cs.AI 79%

ACE, Action and Control via Explanations: A Proposal for LLMs to Provide Human-Centered Explainability for Multimodal AI Assistants

Elizabeth Anne Watkins, Emanuel Moss, Ramesh Manuvinakurike, Meng Shi, Richard Beckwith, Giuseppe Raffa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at Human-Centered Explainable AI workshop at CHI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15647 2025-03-21 cs.CV cs.LG 79%

Multi-Modal Gesture Recognition from Video and Surgical Tool Pose Information via Motion Invariants

Jumanh Atoum, Garrison L. H. Johnston, Nabil Simaan, Jie Ying Wu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14498 2025-03-19 cs.CV cs.RO 79%

Tracking Meets Large Multimodal Models for Driving Scenario Understanding

Ayesha Ishaq, Jean Lahoud, Fahad Shahbaz Khan, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages, 8 figures, Github: https://github.com/mbzuai-oryx/TrackingMeetsLMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13646 2025-03-19 cs.CV 79%

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Chiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha, Federico Tombari

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2025. Dataset and code are available at https://github.com/google-research-datasets/egotempo.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13016 2025-03-18 cs.CV 79%

Efficient Motion-Aware Video MLLM

Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo, Haoyu Lu, Bingning Wang, Weipeng Chen, Jing Liu

专题命中 视频多模态 :MLLM(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11695 2025-03-18 cs.LG cs.AI 79%

MELON: Multimodal Mixture-of-Experts with Spectral-Temporal Fusion for Long-Term Mobility Estimation in Critical Care

Jiaqing Zhang, Miguel Contreras, Jessica Sena, Andrea Davidson, Yuanfang Ren, Ziyuan Guan, Tezcan Ozrazgat-Baslanti, Tyler J. Loftus, Subhash Nerella, Azra Bihorac, Parisa Rashidi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02063 2025-03-17 cs.CV 79%

V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts

Adnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas Bulling

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏