arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2511.17199 2025-11-24 cs.CV 57%

VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation

VLA-4D: 将4D意识嵌入视觉-语言-动作模型中以实现时空一致的机器人操作

Hanyu Zhou, Chuanhao Ma, Gim Hee Lee

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 VLA-4D通过4D意识增强视觉-语言-动作模型,实现时空一致的机器人操作,提升动作执行的空间平滑性和时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12114 2025-11-24 cs.HC cs.CV eess.IV 57%

Behind the Screens: Uncovering Bias in AI-Driven Video Interview Assessments Using Counterfactuals

屏幕背后:利用反事实揭示AI驱动视频面试评估中的偏见

Dena F. Mujtaba, Nihar R. Mahapatra

机构 * Department of Electrical and Computer Engineering, Michigan State University(电子与计算机工程系,密歇根州立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于反事实的框架,用于评估AI驱动视频面试评估中的偏见,通过生成对抗网络生成反事实数据,揭示不同群体间的差异,提升伦理透明度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20470 2025-11-21 cs.CV 57%

Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理

Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun

机构 * State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 Conan通过多阶段渐进冷启动策略和AIR RLVR框架,实现证据基础的多步视频推理,超越基线模型,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05274 2025-11-21 cs.CV 57%

From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos

从游戏到回放:面向时间精细粒度视频的组合视频检索

Animesh Gupta, Jay Parmar, Ishan Rajendrakumar Dave, Mubarak Shah

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 TF-CoVR提出了一种针对时间精细粒度视频检索的框架,通过预训练视频编码器和对比学习提升检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24709 2025-11-20 cs.CV 57%

IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?

Yang Chen, Minghao Liu, Yufan Shen, Yunwen Li, Tianyuan Huang, Xinyu Fang, Tianyu Zheng, Wenxuan Huang, Cheng Yang, Daocheng Fu, Jianbiao Mei, Rong Wu, Yunfei Zhao, Licheng Wen, Xuemeng Yang, Song Mao, Qunshu Lin, Zhi Yu, Yongliang Shen, Yu Qiao, Botian Shi

机构 * IWR-Bench Team(IWR-Bench团队)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14249 2025-11-19 cs.CL 57%

Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning

Rui Liu, Yuan Zhao, Zhenqi Jia

机构 * Rui Liu \equalcontrib , Yuan Zhao \equalcontrib , Zhenqi Jia(Rui Liu、Yuan Zhao、Zhenqi Jia)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13993 2025-11-19 cs.CV 57%

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025, Project webpage: https://vision.cs.utexas.edu/projects/CrossTrainer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16597 2025-11-19 cs.CV 57%

EventHallusion: Diagnosing Event Hallucinations in Video LLMs

Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Na Zhao, Zhiyu Tan, Hao Li, Xingjun Ma, Jingjing Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13190 2025-11-18 cs.CV 57%

Video Spatial Reasoning with Object-Centric 3D Rollout

Haoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang, Linglong Li, Ge Li, Xiaodan Liang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13035 2025-11-18 cs.LG cs.AI 57%

One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow

Zeyuan Wang, Da Li, Yulin Chen, Ye Shi, Liang Bai, Tianyuan Yu, Yanwei Fu

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted in AAAI 2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12291 2025-11-18 cs.CV 57%

One target to align them all: LiDAR, RGB and event cameras extrinsic calibration for Autonomous Driving

Andrea Bertogalli, Giacomo Boracchi, Luca Magri

机构 * DEIB Politecnico di Milano(都灵理工大学DEIB)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07251 2025-11-18 cs.CV 57%

Understanding Dynamic Scenes in Ego Centric 4D Point Clouds

Junsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang, Gaoang Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted as a poster to AAAI 2026; will be published in the proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02473 2025-11-18 cs.CV 57%

Generative Perception of Shape and Material from Differential Motion

Xinran Nicole Han, Ko Nishino, Todd Zickler

机构 * Harvard University(哈佛大学) Kyoto University(京都大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12154 2025-11-18 cs.LG cs.AI 57%

Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions

Gustavo Polleti, Marlesson Santana, Eduardo Fontes

机构 * Trustly

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12150 2025-11-18 cs.CV 57%

Breaking the Modality Wall: Time-step Mixup for Efficient Spiking Knowledge Transfer from Static to Event Domain

Yuqi Xie, Shuhan Ye, Yi Yu, Chong Wang, Qixin Zhang, Jiazhen Xu, Le Shen, Yuanbin Qian, Jiangbo Qian, Guoqi Li

机构 * Ningbo University(宁波大学) Nanyang Technological University(南洋理工大学) Merchants’ Guild Economics and Cultural(商帮经济与文化) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04953 2025-11-18 cs.CV 57%

APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval

Hong Gao, Yiming Bao, Xuezhen Tu, Bin Zhong, Linan Yue, Minling Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01727 2025-11-17 cs.LG cs.CV 57%

OccamVTS: Distilling Vision Models to 1% Parameters for Time Series Forecasting

Sisuo Lyu, Siru Zhong, Weilin Ruan, Qingxiang Liu, Qingsong Wen, Hui Xiong, Yuxuan Liang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04668 2025-11-17 cs.CV 57%

SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding

Ellis Brown, Arijit Ray, Ranjay Krishna, Ross Girshick, Rob Fergus, Saining Xie

机构 * New York University(纽约大学) Boston University(波士顿大学) AllenAI Vercept

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Project page: https://ellisbrown.github.io/sims-v

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25238 2025-11-14 cs.CV 57%

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

Qianqian Qiao, DanDan Zheng, Yihang Bo, Bao Peng, Heng Huang, Longteng Jiang, Huaye Wang, Jingdong Chen, Jun Zhou, Xin Jin

机构 * Nanjing University(南京大学) Huazhong University of Science and Technology(华中科技大学) Beijing Film Academy(北京电影学院) University of Science and Technology of China(中国科学技术大学) Beijing Electronic Science and Technology Institute(北京电子科技学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09870 2025-11-14 cs.CV 57%

SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection

Jia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong, Guodao Zhang, Zhi Liu, Jiyong Zhang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to 40th AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04369 2025-11-14 cs.CV 57%

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou, Xuchong Zhang, Xin Wei, Ye Yuan, Huayu Zhang, Jinglin Xu, Hao Sun

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03184 2025-11-13 eess.IV cs.CV 57%

EvRWKV: A Continuous Interactive RWKV Framework for Effective Event-Guided Low-Light Image Enhancement

Wenjie Cai, Qingguo Meng, Zhenyu Wang, Xingbo Dong, Zhe Jin

机构 * Anhui Provincial International Joint Research center for Advanced technology in Medical imaging(安徽省国际联合先进医学影像技术研究中心) School of Artificial Intelligence(人工智能学院) Anhui University(安徽大学) State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) Anhui Provincial Key Laboratory of Secure Artificial Intelligence(安徽省安全人工智能重点实验室) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(粤港澳大湾区人工智能与数字经济实验室(深圳))

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07106 2025-11-11 cs.CV 57%

HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving

Zhongyu Xia, Zhiwei Lin, Yongtao Wang, Ming-Hsuan Yang

机构 * University of California, Merced(加州大学默塞德分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Preliminary version, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07049 2025-11-11 cs.CV cs.CR 57%

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

Hui Lu, Yi Yu, Song Xia, Yiming Yang, Deepu Rajan, Boon Poh Ng, Alex Kot, Xudong Jiang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments AAAI 2026 (Oral presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06281 2025-11-11 cs.CV 57%

VideoSSR: Video Self-Supervised Reinforcement Learning

Zefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang, Daizong Liu, Yu Cheng

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Wuhan University(武汉大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05636 2025-11-10 cs.LG cs.AI 57%

Graph Learning

Feng Xia, Ciyuan Peng, Jing Ren, Falih Gozi Febrinanto, Renqiang Luo, Vidya Saikrishna, Shuo Yu, Xiangjie Kong

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 185 pages

Journal ref Foundations and Trends in Signal Processing, Vol. 19, No. 4, pp 371-551. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04281 2025-11-07 cs.CV 57%

DINOv2 Driven Gait Representation Learning for Video-Based Visible-Infrared Person Re-identification

Yujie Yang, Shuang Li, Jun Ye, Neng Dong, Fan Li, Huafeng Li

机构 * Kunming University of Science and Technology(昆明理工大学) Chongqing University of Post and Telecommunications(重庆邮电大学) China University of Mining Technology(中国矿业大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03332 2025-11-06 cs.CV 57%

Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge

Yi Yang, Yiming Xu, Timo Kaiser, Hao Cheng, Bodo Rosenhahn, Michael Ying Yang

机构 * Leibniz University Hannover(莱布尼茨汉诺威大学) University of Twente(特文特大学) University of Bath(巴斯大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18422 2025-11-06 cs.CV 57%

Breaking the Encoder Barrier for Seamless Video-Language Understanding

Handong Li, Yiyuan Zhang, Longteng Guo, Xiangyu Yue, Jing Liu

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MMLab, CUHK(CUHK MMLab) Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 12 pages

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 23167-23176

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13174 2025-11-06 cs.CV 57%

Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models

Hao Cheng, Erjia Xiao, Yichi Wang, Chengyuan Yu, Mengshu Sun, Qiang Zhang, Jiahang Cao, Yijie Guo, Ning Liu, Kaidi Xu, Jize Zhang, Chao Shen, Philip Torr, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Xi’an Jiaotong University(西安交通大学) The Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) Beijing University of Technology(北京理工大学) Duke University(杜克大学) X-Humanoid Project(X-Humanoid 项目)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏