arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4728 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4728 篇

2507.16151 2025-07-23 cs.CV cs.AI 62%

SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities

Yasser Ashraf, Ahmed Sharshar, Velibor Bojkovic, Bin Gu

机构 * Department of Machine Learning(机器学习系) Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13820 2025-07-21 cs.CV cs.AI 62%

Team of One: Cracking Complex Video QA with Model Synergy

Jun Xie, Zhaoran Zhao, Xiongjun Guan, Yingjian Zhu, Hongzhu Yi, Xinming Wang, Feng Chen, Zhepeng Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12816 2025-07-18 cs.CV cs.AI 62%

FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering

Ju-Young Oh, Ho-Joong Kim, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments SMC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08411 2025-07-15 cs.LG cs.AI cs.CV stat.AP 62%

BiDepth: A Bidirectional-Depth Neural Network for Spatio-Temporal Prediction

Sina Ehsani, Fenglian Pan, Qingpei Hu, Jian Liu

机构 * University of Arizona(亚利桑那大学) Chinese Academy of Sciences(中国科学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 6 figures. Submitted to ACM TKDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06523 2025-07-10 cs.CV cs.CL cs.GR 62%

FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

Liqiang Jing, Viet Lai, Seunghyun Yoon, Trung Bui, Xinya Du

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04976 2025-07-08 cs.CV cs.CL 62%

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

Eunseop Yoon, Hee Suk Yoon, Mark A. Hasegawa-Johnson, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) University of Illinois at Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13860 2025-07-08 cs.CV cs.AI 62%

Domain Adaptation of VLM for Soccer Video Understanding

Tiancheng Jiang, Henry Wang, Md Sirajus Salekin, Parmida Atighehchian, Shinan Zhang

机构 * Massachusetts Institute of Technology(麻省理工学院) Amazon Web Services(亚马逊网络服务)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures, accepted to the 11th IEEE International Workshop on Computer Vision in Sports (CVSports) at CVPR 2025; supplementary appendix included

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00950 2025-07-02 cs.CV cs.LG cs.MM 62%

MVP: Winning Solution to SMP Challenge 2025 Video Track

Liliang Ye, Yunyao Zhang, Yafeng Wu, Yi-Ping Phoebe Chen, Junqing Yu, Wei Yang, Zikai Song

机构 * Huazhong University of Science and Technology(华中科技大学) La Trobe University(拉特罗布大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09105 2025-07-02 cs.CV cs.AI 62%

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

Chenglin Li, Qianglong Chen, Zhi Li, Feng Tao, Yin Zhang

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21272 2025-06-30 cs.GR cs.CV cs.MM 62%

FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

Jiayi Zheng, Xiaodong Cun

机构 * GVC Lab, Great Bay University(Great Bay大学GVC实验室)

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV、cs.MM

Comments Project Page: https://jayleejia.github.io/FairyGen/ ; Code: https://github.com/GVCLab/FairyGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18071 2025-06-30 cs.CV cs.AI 62%

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

Jisheng Dang, Huilin Song, Junbin Xiao, Bimei Wang, Han Peng, Haoxuan Li, Xun Yang, Meng Wang, Tat-Seng Chua

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21080 2025-06-27 cs.CV cs.AI cs.LG 62%

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan, Calvin Murdock, Ishwarya Ananthabhotla, Yijun Qian, Vamsi Krishna Ithapu, Dinesh Manocha, Ruohan Gao

机构 * University of Maryland, College Park(马里兰大学学院市分校) Meta Reality Labs(Meta现实实验室) Worcester Polytechnic Institute(沃斯特理工学院) University of Toronto(多伦多大学) FAIR, Meta AI(Meta AI)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20342 2025-06-26 cs.CV cs.AI cs.LG 62%

Feature Hallucination for Self-supervised Action Recognition

Lei Wang, Piotr Koniusz

机构 * Griffith University(格里菲斯大学) Data61/CSIRO(Data61/澳大利亚联邦科学与工业研究组织) University of New South Wales(新南威尔士大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in International Journal of Computer Vision (IJCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00526 2025-06-24 cs.CV cs.AI cs.GR 62%

Human Action CLIPs: Detecting AI-generated Human Motion

Matyas Bohacek, Hany Farid

机构 * Google(谷歌) Stanford University(斯坦福大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Journal ref Workshop on Deepfake Detection, Localization and Interpretability @ IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15835 2025-06-23 eess.IV cs.AI cs.CV 62%

MoNetV2: Enhanced Motion Network for Freehand 3D Ultrasound Reconstruction

Mingyuan Luo, Xin Yang, Zhongnuo Yan, Yan Cao, Yuanji Zhang, Xindi Hu, Jin Wang, Haoxuan Ding, Wei Han, Litao Sun, Dong Ni

机构 * National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University, Shenzhen, Guangdong, China(国家级医学超声关键技术研发实验室、生物医学工程学院、深圳大学医学院、深圳大学、深圳、广东、中国) Medical UltraSound Image Computing (MUSIC) Lab, Shenzhen University, Shenzhen, Guangdong, China(医学超声图像计算(MUSIC)实验室、深圳大学、深圳、广东、中国) Shenzhen RayShape Medical Technology Inc.(深圳RayShape医疗科技有限公司) Cancer Center, Department of Ultrasound Medicine, Zhejiang Provincial People’s Hospital, Affiliated People’s Hospital of Hangzhou Medical College, Hangzhou, Zhejiang, China(肿瘤中心、超声医学科、浙江省人民医院、杭州医学院附属人民医院、杭州、浙江、中国) Department of Health Management Center, Qilu Hospital, Cheeloo College of Medicine, Shandong University, Jinan, Shandong, China(健康管理中心、齐鲁医院、山东大学齐鲁医学院、济南、山东、中国)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14144 2025-06-18 cs.CV cs.AI 62%

SceneAware: Scene-Constrained Pedestrian Trajectory Prediction with LLM-Guided Walkability

Juho Bai, Inwook Shim

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13956 2025-06-18 cs.CL cs.AI cs.RO 62%

ASMR: Augmenting Life Scenario using Large Generative Models for Robotic Action Reflection

Shang-Chi Tsai, Seiya Kawano, Angel Garcia Contreras, Koichiro Yoshino, Yun-Nung Chen

机构 * National Taiwan University(国立台湾大学) RIKEN(日本研究机构)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments IWSDS 2024 Best Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13654 2025-06-17 cs.CV cs.AI 62%

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

Shulin Tian, Ruiqi Wang, Hongming Guo, Penghao Wu, Yuhao Dong, Xiuying Wang, Jingkang Yang, Hao Zhang, Hongyuan Zhu, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) A*STAR, Singapore(新加坡A*STAR) Simon Fraser University(西蒙弗雷泽大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://egolife-ai.github.io/Ego-R1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08185 2025-06-17 cs.CV cs.AI 62%

Agentic Surgical AI: Surgeon Style Fingerprinting and Privacy Risk Quantification via Discrete Diffusion in a Vision-Language-Action Framework

Huixin Zhan, Jason H. Moore

机构 * Cedars-Sinai Medical Center(西达萨凡纳医学中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08493 2025-06-11 cs.CV cs.MM 62%

Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization

Qilin Yin, Wei Lu, Xiangyang Luo, Xiaochun Cao

机构 * School of Computer Science and Engineering, MoE Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-sen University(计算机科学与工程学院、信息科技关键实验室、广东信息安全技术重点实验室、中山大学) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室) School of Cyber Science and Technology, Shenzhen Campus, Sun Yat-sen University(网络安全科学与技术学院、深圳校区、中山大学)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08003 2025-06-10 cs.CV cs.AI 62%

Audio-Sync Video Generation with Multi-Stream Temporal Control

Shuchen Weng, Haojie Zheng, Zheng Chang, Si Li, Boxin Shi, Xinlong Wang

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Nat’l Key Lab of General AI, School of Intelligence Science and Technology, Peking University(国家通用人工智能实验室,北京大学智能科学与技术学院) Nat’l Eng. Research Ctr. of Visual Tech., School of Computer Science, Peking University(国家视觉技术工程研究中心,北京大学计算机学院)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06355 2025-06-10 cs.CY cs.CE cs.CL cs.CV 62%

LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment

Lingyao Li, Dawei Li, Zhenhui Ou, Xiaoran Xu, Jingxiao Liu, Zihui Ma, Runlong Yu, Min Deng

机构 * University of South Florida(佛罗里达州立大学) Arizona State University(亚利桑那州立大学) Massachusetts Institute of Technology(麻省理工学院) New York University(纽约大学) University of Alabama(阿拉巴马大学) Texas Tech University(德克萨斯科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21991 2025-06-09 cs.CV cs.AI 62%

A Lightweight Dual-Branch System for Weakly-Supervised Video Anomaly Detection on Consumer Edge Devices

Wen-Dong Jiang, Chih-Yung Chang, Ssu-Chi Kuai, Diptendu Sinha Roy

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This manuscript has been submitted to IEEE TCE and is under consideration for publication, with potential copyright transfer in the future

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03179 2025-06-05 cs.CV cs.AI 62%

Vid-SME: Membership Inference Attacks against Large Video Understanding Models

Qi Li, Runpeng Yu, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19475 2025-06-04 cs.CV cs.AI cs.LG 62%

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Sonia Joseph, Praneet Suresh, Lorenz Hufe, Edward Stevinson, Robert Graham, Yash Vadi, Danilo Bzdok, Sebastian Lapuschkin, Lee Sharkey, Blake Aaron Richards

机构 * Mila Quebec(蒙特利尔大学) McGill University(麦吉尔大学) Meta Université de Montréal(蒙特利尔大学) Imperial College London(伦敦帝国理工学院) Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫 Heinrich Hertz 研究所) Technological University Dublin(都柏林技术大学) Apollo Research(Apollo 研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 4 pages, 3 figures, 9 tables. Oral and Tutorial at the CVPR Mechanistic Interpretability for Vision (MIV) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00928 2025-06-03 cs.CV cs.CL 62%

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

Olga Loginova, Sofía Ortega Loguinova

机构 * University of Trento(特伦托大学) Maastricht University(马斯特里赫特大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19753 2025-06-03 cs.GR cs.AI cs.CV 62%

A Survey on Event-driven 3D Reconstruction: Development under Different Categories

Chuanzhi Xu, Haoxian Zhou, Haodong Chen, Vera Chung, Qiang Qu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments We have decided not to submit this article and plan to withdraw it from public display. The content of this article will be presented in a more comprehensive form in another work

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16372 2025-05-23 cs.CV cs.AI 62%

Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition

Feng Liu, Bingyu Nan, Xuezhong Qian, Xiaolan Fu

机构 * School of Psychology,Shanghai Jiao Tong University(上海交通大学心理学学院) Jiangnan University(江南大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15447 2025-05-22 cs.CV cs.AI 62%

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Ziqiang Xu, Qi Dai, Tian Xie, Yifan Yang, Kai Qiu, DongDong Chen, Zuxuan Wu, Chong Luo

机构 * Fudan University(复旦大学) Microsoft(微软公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17821 2025-05-21 cs.CV cs.CL 62%

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

Xinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu, Wenhan Luo, Yaowei Wang, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏