arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2508.14442 2025-08-21 cs.HC cs.AI 57%

Detecting Reading-Induced Confusion Using EEG and Eye Tracking

Haojun Zhuang, Dünya Baradari, Nataliya Kosmyna, Arnav Balyan, Constanze Albrecht, Stephanie Chen, Pattie Maes

机构 * University of California, Berkeley(加州大学伯克利分校) MIT Media Lab(MIT媒体实验室) Princeton University(普林斯顿大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07961 2025-08-20 cs.CV 57%

Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, Andrea Vedaldi

机构 * Visual Geometry Group, University of Oxford(视觉几何组,牛津大学) Naver Labs Europe(Naver欧洲实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 17 pages, 6 figures, ICCV 2025 Highlight, Project page: https://geo4d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00838 2025-08-20 cs.CV cs.LG 57%

Spatially-guided Temporal Aggregation for Robust Event-RGB Optical Flow Estimation

Qianang Zhou, Junhui Hou, Meiyi Yang, Yongjian Deng, Youfu Li, Junlin Xiong

机构 * Department of Automation, University of Science and Technology of China(自动化系,中国科学技术大学) Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学) Department of Mechanical Engineering, City University of Hong Kong(机械工程系,香港城市大学) College of Computer Science, Beijing University of Technology(计算机科学学院,北京理工大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 11 pages, 8 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13205 2025-08-20 cs.CV eess.IV 57%

YOLO11-CR: a Lightweight Convolution-and-Attention Framework for Accurate Fatigue Driving Detection

Zhebin Jin, Ligang Dong

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24039 2025-08-19 cs.CV cs.HC 57%

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Shubhabrata Mukherjee, Jack Lang, Obeen Kwon, Iryna Zenyuk, Valerie Brogden, Adam Weber, Daniela Ushizima

机构 * Lawrence Berkeley National Laboratory(伯克利国家实验室) University of California, Irvine(加州大学尔湾分校) University of California, Berkeley(加州大学伯克利分校) Covalent Metrology(协力计量)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments This paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02356 2025-08-19 cs.CV 57%

InterRVOS: Interaction-aware Referring Video Object Segmentation

Woojeong Jin, Seongchan Kim, Jaeho Lee, Seungryong Kim

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07992 2025-08-18 cs.MM 57%

Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos

Haisong Gong, Bolan Su, Xinrong Zhang, Jing Li, Qiang Liu, Shu Wu, Liang Wang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments in submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10784 2025-08-15 q-bio.NC cs.CV 57%

Insights from the Algonauts 2025 Winners

Paul S. Scotti, Mihir Tripathy

机构 * Medical AI Research Center ( MedARC )(医学人工智能研究中心(MedARC)) Baylor College of Medicine(贝勒医学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Perspective piece on Algonauts 2025 Challenge conclusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07171 2025-08-15 cs.CV 57%

EventRR: Event Referential Reasoning for Referring Video Object Segmentation

Huihui Xu, Jiashi Lin, Haoyu Chen, Junjun He, Lei Zhu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Northwestern Polytechnical University(西北工业大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00399 2025-08-15 cs.CV 57%

iSafetyBench: A video-language benchmark for safety in industrial environment

Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to VISION'25 - ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03096 2025-08-15 cs.CV 57%

Scaling Open-Vocabulary Action Detection

Zhen Hao Sia, Yogesh Singh Rawat

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09857 2025-08-14 cs.CV 57%

OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better

Yupeng Zhou, Zhen Li, Ziheng Ouyang, Yuming Chen, Ruoyi Du, Daquan Zhou, Bin Fu, Yihao Liu, Peng Gao, Ming-Ming Cheng, Qibin Hou

机构 * VCIP, School of Computer Science, Nankai University(VCIP,计算机科学学院,南开大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18923 2025-08-14 cs.CV 57%

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Meng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu, Haoran Tang, Haoze Zhao, Chen Wang, Jiahua Dong, Wangbo Yu, Ge Zhang, Jun Song, Xiang Li, Bo Zheng, Ian Reid, Xiaodan Liang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09262 2025-08-14 cs.CV cs.LG 57%

Harnessing Input-Adaptive Inference for Efficient VLN

Dongwoo Kang, Akhil Perincherry, Zachary Coalson, Aiden Gabriel, Stefan Lee, Sanghyun Hong

机构 * Oregon State University(俄勒冈州立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025 [Poster]

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08989 2025-08-13 cs.CV 57%

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08590 2025-08-13 cs.CV cs.HC 57%

QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection

Yuxiao Wang, Wolin Liang, Yu Lei, Weiying Xue, Nan Zhuang, Qi Liu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07989 2025-08-12 cs.CV cs.HC 57%

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

Xiantao Zhang

机构 * Beihang University(北航大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 3 figures, 2 tables. Accepted at CV4A11y, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07626 2025-08-12 cs.CV cs.RO 57%

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

Dejie Yang, Zijing Zhao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07312 2025-08-12 cs.CV 57%

MobileViCLIP: An Efficient Video-Text Model for Mobile Devices

Min Yang, Zihan Jia, Zhilin Dai, Sheng Guo, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) MyBank, Ant Group(蚂蚁集团MyBank) Shanghai AI Lab(上海AI实验室)

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07006 2025-08-12 eess.IV cs.CV 57%

Spatio-Temporal Conditional Diffusion Models for Forecasting Future Multiple Sclerosis Lesion Masks Conditioned on Treatments

Gian Mario Favero, Ge Ya Luo, Nima Fathi, Justin Szeto, Douglas L. Arnold, Brennan Nichyporuk, Chris Pal, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to MICCAI 2025 (LMID Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12524 2025-08-12 cs.CV cs.HC cs.LG eess.IV 57%

Inference-Time Gaze Refinement for Micro-Expression Recognition: Enhancing Event-Based Eye Tracking with Motion-Aware Post-Processing

Nuwan Bandara, Thivya Kandappu, Archan Misra

机构 * School of Computing(计算学院) Information Systems, Singapore Management University, Singapore(信息系统,新加坡管理大学,新加坡)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at 4DMR@IJCAI25: International IJCAI Workshop on 1st Challenge and Workshop for 4D Micro-Expression Recognition for Mind Reading, August 29, 2025, Guangzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04592 2025-08-11 cs.LG cs.AI cs.CE cs.IR 57%

CAMEF: Causal-Augmented Multi-Modality Event-Driven Financial Forecasting by Integrating Time Series Patterns and Salient Macroeconomic Announcements

Yang Zhang, Wenbo Yang, Jun Wang, Qiang Ma, Jie Xiong

机构 * Southwestern University of Finance and Economics(西南财经大学) Kyoto Institute of Technology(京都技术大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted in SIGKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13667 2025-08-11 cs.CV 57%

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

Fu Rong, Meng Lan, Qian Zhang, Lefei Zhang

机构 * National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(国家多媒体软件工程研究中心,计算机学院,武汉大学) Hong Kong University of Science and Technology(香港科学与技术大学) Horizon Robotics

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10683 2025-08-11 cs.RO cs.AI 57%

Learning to Initialize Trajectory Optimization for Vision-Based Autonomous Flight in Unknown Environments

Yicheng Chen, Jinjie Li, Wenyuan Qin, Yongzhao Hua, Xiwang Dong, Qingdong Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted to IROS 2025. Source code available

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05145 2025-08-08 cs.AI 57%

Graph-based Event Log Repair

Sebastiano Dissegna, Chiara Di Francescomarino, Massimiliano Ronzani

机构 * Department of Computer Science Engineering University of Trento(计算机科学工程系 特伦托大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03651 2025-08-06 cs.HC cs.AI 57%

Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired

Ruei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap, Anhong Guo

机构 * University of Michigan(密歇根大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments ACM ASSETS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01546 2025-08-05 cs.CV 57%

E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation

Zeyu Xu, Junkang Zhang, Qiang Wang, Yi Liu

机构 * Zeyu Xu(作者) Junkang Zhang(作者) Qiang Wang(作者) Yi Liu(作者)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21971 2025-07-30 cs.CV 57%

EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation

Zhijiang Li, Haoran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20629 2025-07-29 cs.CV 57%

DAMS:Dual-Branch Adaptive Multiscale Spatiotemporal Framework for Video Anomaly Detection

Dezhi An, Wenqiang Liu, Kefan Wang, Zening Chen, Jun Lu, Shengcai Zhang

机构 * School of Cyberspace Security,Gansu University of Political Science and Law(网络安全学院、政治学科学校)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 13 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20939 2025-07-29 cs.CV 57%

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Search Application Department, Tencent CSIG(腾讯CSIG搜索应用部门) Tencent Hunyuan(腾讯文生视频) Big Data Platform Department, Tencent PCG(腾讯PCG大数据平台部门)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://tencentarc.github.io/posts/arc-video-announcement/

详情

展开后加载摘要…

URL PDF HTML 收藏