arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4728 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4728 篇

2509.18187 2025-09-24 cs.CV cs.AI 62%

V-SenseDrive: A Privacy-Preserving Road Video and In-Vehicle Sensor Fusion Framework for Road Safety & Driver Behaviour Modelling

Muhammad Naveed, Nazia Perwaiz, Sidra Sultana, Mohaira Ahmad, Muhammad Moazam Fraz

机构 * School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST)(电气工程与计算机科学学院(SEECS),国立科学与技术大学(NUST))

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18183 2025-09-24 cs.CV cs.AI 62%

VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation

Jinyue Bian, Zhaoxing Zhang, Zhengyu Liang, Shiwei Zheng, Shengtao Zhang, Rong Shen, Chen Yang, Anzhou Hou

机构 * China, Beijing, Li Auto Inc.(中国北京李自动公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17888 2025-09-23 cs.CV cs.AI 62%

Trainee Action Recognition through Interaction Analysis in CCATT Mixed-Reality Training

Divya Mereddy, Marcos Quinones-Grueiro, Ashwin T S, Eduardo Davalos, Gautam Biswas, Kent Etherton, Tyler Davis, Katelyn Kay, Jill Lear, Benjamin Goldberg

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13722 2025-09-18 cs.CV cs.AI 62%

Mitigating Query Selection Bias in Referring Video Object Segmentation

Dingwei Zhang, Dong Zhang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) The Hong Kong University of Science and Technology(香港科学大学) Nanjing Forestry University(南京林业大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12876 2025-09-17 cs.CL cs.MM 62%

Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents

Fuyu Xing, Zimu Wang, Wei Wang, Haiyang Zhang

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL、cs.MM

Comments Accepted at INLG 2025. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11232 2025-09-16 cs.CV cs.AI 62%

MIS-LSTM: Multichannel Image-Sequence LSTM for Sleep Quality and Stress Prediction

Seongwan Park, Jieun Woo, Siheon Yang

机构 * Sungkyunkwan University(釜山大学) Yeungnam University(延世大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICTC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02074 2025-09-10 cs.CV cs.AI 62%

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction and Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09632 2025-09-09 cs.CV cs.AI 62%

Preacher: Paper-to-Video Agentic System

Jingwei Liu, Ling Yang, Hao Luo, Fan Wang, Hongyan Li, Mengdi Wang

机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) DAMO Academy, Alibaba group(阿里巴巴集团大模型研究院) Hupan Lab(虎扑实验室) National Key Laboratory of General Artificial Intelligence, Peking University(北京人工智能 general artificial intelligence 国家重点实验室) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. Code: https://github.com/Gen-Verse/Paper2Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05298 2025-09-09 cs.HC cs.AI cs.MM 62%

Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression

Rui Xi, Xianghan Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments Accepted to the Proceedings of the 2025 International Conference on Artificial Intelligence and Virtual Reality (AIVR 2025). \c{opyright} 2025 Springer. This is the author-accepted manuscript. Rui Xi and Xianghan Wang contributed equally to this work. The final version will be available via SpringerLink

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21496 2025-09-03 cs.CV cs.AI 62%

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu

机构 * Sensetime(秒氏科技)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01383 2025-09-03 cs.CV cs.MM 62%

Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

Long Zhang, Peipei Song, Jianfeng Dong, Kun Li, Xun Yang

机构 * University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发式智能感知与认知实验室) Zhejiang Gongshang University(浙江工商大学) ReLER, CCAI, Zhejiang University(ReLER,中国计算机学会,浙江大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00210 2025-09-03 cs.CV cs.AI 62%

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu, Qinhan Lv, Xiying Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17932 2025-08-26 cs.CV cs.AI 62%

See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops

Zixuan Dong, Baoyun Peng, Yufei Wang, Lin Liu, Xinxin Dong, Yunlong Cao, Xiaodong Wang

机构 * College of Computer, National University of Defense Technology(计算机学院,国防科技大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17160 2025-08-26 cs.CV cs.AI 62%

Beyond Play and Pause: Turning GPT-4o Spatial Weakness into a Strength for In-Depth Interactive Video Learning

Sajad Goudarzi, Samaneh Zamanifard

机构 * Clemson University(克莱姆森大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM 62%

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16291 2025-08-25 cs.CV cs.MM 62%

Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating Assessment

Fengshun Wang, Qiurui Wang, Peilin Zhao

机构 * Capital University of \ Education Shanghai Jiao Tong University Shanghai China Shanghai Jiao Tong University

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15717 2025-08-22 cs.CV cs.AI 62%

StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Yanlai Yang, Zhuokai Zhao, Satya Narayan Shukla, Aashu Singh, Shlok Kumar Mishra, Lizhu Zhang, Mengye Ren

机构 * Meta AI New York University(纽约大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14941 2025-08-22 cs.MM cs.CL 62%

Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.MM

Comments 12 pages, 4 figures, 2 tables. Extends our earlier framework on hierarchical narrative graphs with a semantic normalization module

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01712 2025-08-18 cs.CV cs.AI 62%

HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection

Han Wang, Zhuoran Wang, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08035 2025-08-12 cs.CV cs.AI 62%

LVBench: An Extreme Long Video Understanding Benchmark

Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Xiaotao Gu, Shiyu Huang, Bin Xu, Yuxiao Dong, Ming Ding, Jie Tang

机构 * Zhipu AI(智谱AI) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15384 2025-08-11 cs.CV cs.AI 62%

MetaOcc: Spatio-Temporal Fusion of Surround-View 4D Radar and Camera for 3D Occupancy Prediction with Dual Training Strategies

Long Yang, Lianqing Zheng, Wenjin Ai, Minghao Liu, Sen Li, Qunshu Lin, Shengyu Yan, Jie Bai, Zhixiong Ma, Tao Huang, Xichan Zhu

机构 * School of Automotive Studies, Tongji University(同济大学汽车学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Automobile, Chang'an University(长安大学汽车学院) School of Information and Electrical Engineering, Hangzhou City University(杭州城市学院信息与电气工程学院) College of Science and Engineering, James Cook University(詹姆斯库克大学科学与工程学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03009 2025-08-06 cs.CV cs.AI 62%

Enhancing Long Video Question Answering with Scene-Localized Frame Grouping

Xuyi Yang, Wenhao Zhang, Hongbo Jin, Lin Liu, Hongbo Xu, Yongwei Nie, Fei Yu, Fei Ma

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18688 2025-08-06 cs.CV cs.AI 62%

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

Faraz Waseem, Muhammad Shahzad

机构 * University Of Reading(阅读大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 35 pages, 18 figures, Manuscript submitted to ACM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01781 2025-08-05 cs.CL cs.AI 62%

A comprehensive taxonomy of hallucinations in Large Language Models

Manuel Cossio

机构 * Universitat de Barcelona(巴塞罗那大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 55 pages, 16 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00085 2025-08-04 cs.CV cs.AI 62%

Punching Bag vs. Punching Person: Motion Transferability in Videos

Raiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran, Yogesh Rawat

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Center for Vision Technology, SRI International(视觉技术中心,SRI国际)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02713 2025-08-04 cs.CV cs.CL 62%

LLaVA-Video: Video Instruction Tuning With Synthetic Data

Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, Chunyuan Li

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) BUPT(北京邮电大学) ByteDance(字节跳动)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Project page: https://llava-vl.github.io/blog/2024-09-30-llava-video/; Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01430 2025-08-01 cs.CV cs.AI cs.LG 62%

Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots

Dong Lao, Zhengyang Hu, Francesco Locatello, Yanchao Yang, Stefano Soatto

机构 * UCLA(加州大学洛杉矶分校) HKU(香港大学) ISTA(因斯布鲁克大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20947 2025-08-01 cs.CV cs.MM 62%

Hierarchical Sub-action Tree for Continuous Sign Language Recognition

Dejie Yang, Zhu Xu, Xinjie Gao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University, Beijing, China(王轩计算机技术研究所,北京大学,北京,中国) State Key Laboratory of General Artificial Intelligence, Peking Universitys, Beijing, China(通用人工智能国家重点实验室,北京大学,北京,中国)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Journal ref ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17402 2025-07-29 cs.CV cs.IR cs.MM 62%

HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

Jun Li, Jinpeng Wang, Chaolei Tan, Niu Lian, Long Chen, Yaowei Wang, Min Zhang, Shu-Tao Xia, Bin Chen

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Research Center of Artificial Intelligence, Peng Cheng Laboratory(鹏城实验室人工智能研究中心) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ICCV'25. 13 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17958 2025-07-28 cs.LG cs.AI cs.CV 62%

VIBE: Video-Input Brain Encoder for fMRI Response Modeling

Daniel Carlström Schad, Shrey Dixit, Janis Keck, Viktor Studenyak, Aleksandr Shpilevoi, Andrej Bicanski

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏