arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2603.11896 2026-03-13 cs.CV cs.AI cs.CL 82%

Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models

边看边想:多轮视频推理的在线流式段级内存

Lu Wang, Zhuoran Jin, Yupu Hao, Yubo Chen, Kang Liu, Yulong Ao, Jun Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出Think While Watching框架,通过段级内存保持多轮视频推理连续性,提升流式视频理解的准确率和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03099 2026-03-05 cs.RO 82%

Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping

Point2Act: 多模态大语言模型的高效3D蒸馏用于零样本情境感知抓取

Sang Min Kim, Hyeongjun Heo, Junho Kim, Yonghyeon Lee, Young Min Kim

机构 * Seoul National University(首尔国立大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract)

AI总结 Point2Act通过多模态大语言模型高效蒸馏实现零样本情境感知抓取,生成空间定位响应以支持实际操作任务。

Comments Accepted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02075 2026-03-03 cs.DC 82%

Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines

Trident:异构多模态数据流水线的自适应调度

Ding Pan, Zhuangzhuang Zhou, Long Qian, Binhang Yuan

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(abstract)

AI总结 Trident通过自适应调度框架提升异构多模态数据流水线的吞吐量,优化资源分配与配置转换,降低内存不足风险。

Comments 22 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22299 2026-02-27 cs.MM cs.AI cs.CL cs.LG 82%

Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video Ads

解码钩子:一种多模态大语言模型框架用于分析视频广告的钩子阶段

Kunpeng Zhang, Poppy Zhang, Shawndra Hill, Amel Awadelkarim

机构 * University of Marland, College Park(马里兰大学 College Park分校) Meta Platforms, Inc.(Meta平台公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 本研究提出一种多模态大语言模型框架,用于分析视频广告的钩子阶段,通过多策略采样和主题提炼提升广告初期效果分析的准确性与实用性。

Comments 11 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04104 2026-02-06 cs.HC 82%

Making Videos Accessible for Blind and Low Vision Users Using a Multimodal Agent Video Player

通过多模态代理视频播放器使视频对视障和低视力用户更加可访问

Adriana Olmos, Anoop K. Sinha, Renelito Delos Santos, Ruben Rodriguez Rodriguez, James A. Landay, Sam S. Sepah, Philip Nelson, Shaun K. Kane

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract)

AI总结 本文提出一种多模态代理视频播放器,通过多层提示编排为视障和低视力用户带来交互式、可访问的视频体验,增强用户自主性和信任感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19139 2026-01-30 cs.LG cs.DC cs.ET 82%

Native LLM and MLLM Inference at Scale on Apple Silicon

在Apple Silicon上大规模原生LLM和MLLM推理

Wayner Barrios

机构 * Wiqonn Technologies(Wiqonn技术公司)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 vllm-mlx在Apple Silicon上实现高效LLM和MLLM推理,通过原生优化提升文本模型吞吐量并减少多模态延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11558 2026-01-29 cs.CV cs.AI cs.CL 82%

DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs

DaMO:一种数据高效的多模态协调器,用于视频LLM的时序推理

Bo-Cheng Chiu, Jen-Jee Chen, Yu-Chee Tseng, Feng-Chi Chen, An-Zi Yen

机构 * College of Artificial Intelligence, National Yang Ming Chiao Tung University(人工智能学院,阳明交通大学) Institute of Population Health Sciences, National Health Research Institutes(人口健康科学研究所,国家健康研究院) Department of Computer Science, National Yang Ming Chiao Tung University(计算机科学系,阳明交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 DaMO是一种专为视频LLM设计的数据高效多模态协调器,通过时序感知Fuseformer和四阶段训练范式提升时序推理能力,实现更精确的多模态理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06037 2026-01-23 cs.CL cs.AI cs.CV 82%

TeleMem: Building Long-Term and Multimodal Memory for Agentic AI

TeleMem: 构建面向代理AI的长期和多模态记忆

Chunliang Chen, Ming Guan, Xiao Lin, Jiaxu Li, Luxi Lin, Qiyi Wang, Xiangyu Chen, Jixiang Luo, Changzhi Sun, Dell Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 TeleMem通过统一的长期和多模态记忆系统,提升代理AI在长对话和多模态任务中的表现,实现更高的准确率、更低的令牌使用和更快的操作速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01322 2026-01-06 cs.CV cs.AI cs.LG cs.MM eess.IV 82%

LinMU: Multimodal Understanding Made Linear

LinMU: 使多模态理解线性化

Hongjie Wang, Niraj K. Jha

机构 * Princeton University(普林斯顿大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 LinMU通过线性复杂度设计实现多模态理解,无需二次注意力模块,提升视频处理效率。

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17574 2025-12-22 cs.DC cs.LG 82%

Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing

通过GPU内部调度和资源共享实现解耦的多阶段MLLM推理

Lingxiao Zhao, Haoran Zhou, Yuezhi Che, Dazhao Cheng

机构 * Wuhan University(武汉大学)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 通过GPU内部调度和资源共享实现解耦的多阶段MLLM推理,提升吞吐量和延迟性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21908 2025-12-05 cs.LG 82%

Multi-Modal Machine Learning for Early Trust Prediction in Human-AI Interaction Using Face Image and GSR Bio Signals

多模态机器学习在人机交互中早期信任预测中的应用:利用面部图像和GSR生物信号

Hamid Shamszare, Avishek Choudhury

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract)

AI总结 本研究提出多模态机器学习框架,结合面部图像和GSR生物信号,用于预测人机交互中AI或人类推荐的早期信任,通过多模态堆叠集成提升预测性能。

Comments This version contains errors in content presentation and arrangement, so it is being withdrawn until a corrected version is generated

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10068 2025-12-01 cs.CV cs.AI cs.CL 82%

Mavors: Multi-granularity Video Representation for Multimodal Large Language Model

Mavors:多粒度视频表示用于多模态大语言模型

Yang Shi, Jiaheng Liu, Yushuo Guan, Zhenhua Wu, Yuanxing Zhang, Zihao Wang, Weihong Lin, Jingyun Hua, Zekun Wang, Xinlong Chen, Bohan Zeng, Wentao Zhang, Fuzheng Zhang, Wenjing Yang, Di Zhang

机构 * Peking University(北京大学) Kling Team(Kling团队) Nanjing University(南京大学) CASIA(中国科学院自动化研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Mavors通过多粒度视频表示方法,提升多模态大语言模型在长视频理解中的时空模式保留与计算效率平衡能力。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03179 2025-11-20 cs.CV cs.MM cs.SD eess.AS 82%

UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization

Tiantian Geng, Teng Wang, Jinming Duan, Yanfu Zhang, Weili Guan, Feng Zheng, Ling shao

机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学) School of Computer Science, University of Birmingham(计算机科学学院,伯明翰大学) Department of Computer Science, University of Hong Kong(计算机科学系,香港大学) Division of Informatics, Imaging and Data Sciences, University of Manchester(信息学、成像与数据科学系,曼彻斯特大学) William and Mary(威廉与玛丽学院) Harbin Institute of Technology(哈尔滨工业大学) UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学-Terminus AI实验室)

专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Published on IEEE TPAMI

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10280-10294, August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04351 2025-11-18 eess.SP 82%

RCMCL: A Unified Contrastive Learning Framework for Robust Multi-Modal (RGB-D, Skeleton, Point Cloud) Action Understanding

Hasan Akgul, Mari Eplik, Javier Rojas, Akira Yamamoto, Rajesh Kumar, Maya Singh

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

Comments 11 pages, 6 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17637 2025-10-29 cs.LG 82%

Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal Approach

Yuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen, Yunjun Gao

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21445 2025-10-27 cs.CL cs.AI cs.CV cs.LG 82%

REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring

Thanh Cong Ho, Farah Kharrat, Abderrazek Abid, Fakhri Karray

机构 * 2 Department of Electrical Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email 3 College of Computer Information Sciences Prince Sultan University Email

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14203 2025-10-17 cs.CV cs.CL cs.MM 82%

Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition

Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Taiga Yamane, Naotaka Kawata, Satoshi Suzuki, Taichi Katayama

机构 * NTT, Inc.(日本NTT公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at APSIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12164 2025-10-09 cs.CV cs.AI cs.MM 82%

Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward

Yolo Yunlong Tang, Siting Xu, Teng Wang, Qin Lin, Qinglin Lu, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) Tencent Inc.(腾讯公司)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05829 2025-10-08 cs.SD cs.CV cs.LG cs.MM eess.AS 82%

FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders

Riccardo Fosco Gramaccioni, Christian Marinoni, Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello

机构 * Dept. Information Engineering, Electronics and Telecommunications (DIET), Sapienza University of Rome(信息工程、电子与电信系(DIET),罗马萨皮恩扎大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Acepted at IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01513 2025-10-03 cs.CV cs.AI cs.CL cs.IR 82%

From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding

Basem Rizk, Joel Walsh, Mark Core, Benjamin Nye

机构 * University Of Southern California(美国南加州大学)

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22646 2025-10-02 cs.CV cs.AI cs.CL 82%

Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs

Xingyu Fu, Siyi Liu, Yinuo Xu, Pan Lu, Guangqiuse Hu, Tianbo Yang, Taran Anantasagar, Christopher Shen, Yikai Mao, Yuanzhe Liu, Keyush Shah, Chung Un Lee, Yejin Choi, James Zou, Dan Roth, Chris Callison-Burch

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://deeptracereward.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08283 2025-09-23 cs.IR 82%

Serendipitous Recommendation with Multimodal LLM

Haoting Wang, Jianling Wang, Hao Li, Fangjun Yi, Mengyu Fu, Youwei Zhang, Yifan Liu, Liang Liu, Minmin Chen, Ed H. Chi, Lichan Hong, Haokai Lu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted by 2025 Recsys EARL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04254 2025-09-05 cs.HC 82%

MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals

Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, Paolo Burelli

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11092 2025-08-18 cs.LG 82%

Predictive Multimodal Modeling of Diagnoses and Treatments in EHR

Cindy Shih-Ting Huang, Clarence Boon Liang Ng, Marek Rei

机构 * Imperial College London(帝国理工学院伦敦分校)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08882 2025-08-08 cs.MM cs.AI cs.CV 82%

A Novel Multimodal System to Predict Agitation in People with Dementia Within Clinical Settings: A Proof of Concept

Abeer Badawi, Somayya Elmoghazy, Samira Choudhury, Sara Elgazzar, Khalid Elgazzar, Amer Burhan

机构 * IoT Research Laboratory, Ontario Tech University(Ontario Tech 大学物联网研究实验室) Ontario Shores Centre for Mental Health Sciences(Ontario Shores 精神健康科学中心) Temerty Faculty of Medicine, University of Toronto(多伦多大学Temerty医学学院) Faculty of Science, Ontario Tech University(Ontario Tech 大学科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07261 2025-07-11 cs.LG eess.SP 82%

Robust Multimodal Learning Framework For Intake Gesture Detection Using Contactless Radar and Wearable IMU Sensors

Chunzhuo Wang, Hans Hallez, Bart Vanrumste

机构 * e-Media Research Lab(e-Media研究实验室) ESAT-STADIUS Division, KU Leuven(ESAT-STADIUS部门,KU莱顿大学) M-Group, DistriNet, Department of Computer Science, KU Leuven(M组、DistriNet、计算机科学系,KU莱顿大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments This manuscript has been submitted to a peer-reviewed journal and is currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07184 2025-06-10 cs.AI cs.CL cs.CV 82%

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images

Liangliang You, Junchi Yao, Shu Yang, Guimin Hu, Lijie Hu, Di Wang

机构 * Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学) University of Electronic Science and Technology of China(中国电子科学技术大学) University of Copenhagen(哥本哈根大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20811 2025-06-10 cs.CV cs.CL cs.MM 82%

HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models

Xiao Wang, Jingyun Hua, Weihong Lin, Yuanxing Zhang, Fuzheng Zhang, Jianlong Wu, Di Zhang, Liqiang Nie

机构 * Kuaishou Technology(快手科技) Shandong University(山东省大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02430 2025-06-05 cs.CL cs.AI cs.CV cs.LG 82%

Generative Emotion Cause Explanation in Multimodal Conversations

Lin Wang, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, Zhitao Zhang

机构 * Northeastern University(东北大学) Shenyang Women’s and Children’s Hospital(沈阳市妇女儿童医院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01077 2025-06-03 cs.GR cs.HC 82%

TRiMM: Transformer-Based Rich Motion Matching for Real-Time multi-modal Interaction in Digital Humans

Yueqian Guo, Tianzhao Li, Xin Lyu, Jiehaolin Chen, Zhaohan Wang, Sirui Xiao, Yurun Chen, Yezi He, Helin Li, Fan Zhang

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

Comments 24 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏