arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2501.15808 2025-06-27 cs.CV 57%

ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring

Xiaopeng Lin, Yulong Huang, Hongwei Ren, Zunchang Liu, Yue Zhou, Haotian Fu, Bojun Cheng

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20601 2025-06-26 cs.CV 57%

Video Perception Models for 3D Scene Synthesis

Rui Huang, Guangyao Zhai, Zuria Bauer, Marc Pollefeys, Federico Tombari, Leonidas Guibas, Gao Huang, Francis Engelmann

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17840 2025-06-24 cs.LG cs.AI 57%

Causal Spherical Hypergraph Networks for Modelling Social Uncertainty

Anoushka Harit, Zhongtian Sun

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17746 2025-06-24 cs.CV 57%

PhysID: Physics-based Interactive Dynamics from a Single-view Image

Sourabh Vasant Gothe, Ayon Chattopadhyay, Gunturi Venkata Sai Phani Kiran, Pratik, Vibhav Agarwal, Jayesh Rajkumar Vachhani, Sourav Ghosh, Parameswaranath VM, Barath Raj KR

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Published in 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Project page: https://physid.github.io/

Journal ref 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025, pp. 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16701 2025-06-23 cs.CV 57%

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

Xiaodan Hu, Chuhang Zou, Suchen Wang, Jaechul Kim, Narendra Ahuja

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon.com LLC(亚马逊公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16201 2025-06-23 cs.RO cs.CV 57%

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

Sen Wang, Le Wang, Sanping Zhou, Jingyi Tian, Jiayi Li, Haowen Sun, Wei Tang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人类-机器混合增强智能重点实验室) National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) Xi’an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08343 2025-06-19 cs.CL 57%

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency

Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna, Tianyi Zhou

机构 * University College London(伦敦大学学院) University of Washington(华盛顿大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13956 2025-06-19 cs.CV 57%

Improving LLM Video Understanding with 16 Frames Per Second

Yixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li, Zejun Ma, Chao Zhang

机构 * Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12811 2025-06-17 cs.LG cs.AI 57%

Flow-Based Policy for Online Reinforcement Learning

Lei Lv, Yunfei Li, Yu Luo, Fuchun Sun, Tao Kong, Jiafeng Xu, Xiao Ma

机构 * Tsinghua University(清华大学) Shanghai Research Institute for Intelligent Autonomous Systems,Tongji University(上海智能自主系统研究所,同济大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03948 2025-06-17 cs.CV 57%

Reading Between the Lanes: Text VideoQA on the Road

George Tom, Minesh Mathew, Sergi Garcia, Dimosthenis Karatzas, C. V. Jawahar

机构 * Center for Visual Information Technology (CVIT)(视觉信息科技中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11040 2025-06-16 cs.LG cs.CL cs.ET 57%

Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

Feifei Shi, Xueyan Yin, Kang Wang, Wanyu Tu, Qifu Sun, Huansheng Ning

机构 * School of Computer and Communication Engineering, University of Science and Technology Beijing(计算机与通信工程学院,北京科技大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09565 2025-06-16 cs.CV 57%

SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields

Qijing Li, Jingxiang Sun, Liang An, Zhaoqi Su, Hongwen Zhang, Yebin Liu

机构 * Beijing Normal University(北京师范大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15513 2025-06-11 cs.CV 57%

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler

Xingjian Zhang, Xi Weng, Yihao Yue, Zhaoxin Fan, Wenjun Wu, Lei Huang

机构 * SKLCCSE, Institute of Artificial Intelligence, Beihang University, Beijing, China(信息与通信工程学院,人工智能研究院,北京航空航天大学,北京,中国) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) Hangzhou International Innovation Institute, Beihang University, Hangzhou, China(杭州国际创新研究院,北京航空航天大学,杭州,中国)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments code and training recipes are available at https://github.com/ZhangXJ199/TinyLLaVA-Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03817 2025-06-11 cs.CV 57%

Markerless Multi-view 3D Human Pose Estimation: a survey

Ana Filipa Rodrigues Nogueira, Hélder P. Oliveira, Luís F. Teixeira

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 26 pages, 10 tables, 6 figures, accepted at Image and Vision Computing (IMAVIS)

Journal ref In: Image and Vision Computing 155 (2025), p. 105437. issn: 0262-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24718 2025-06-10 cs.CV 57%

Reinforcing Video Reasoning with Focused Thinking

Jisheng Dang, Jingze Wu, Teng Wang, Xuanhui Lin, Nannan Zhu, Hongbo Chen, Wei-Shi Zheng, Meng Wang, Tat-Seng Chua

机构 * Sun Yat-sen University(中山大学) Lanzhou University(兰州大学) University of Hong Kong(香港大学) Hefei University of Technology(合肥工业大学) National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16784 2025-06-10 cs.CV 57%

Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles

Jun Xie, Xiongjun Guan, Yingjian Zhu, Zhaoran Zhao, Xinming Wang, Hongzhu Yi, Feng Chen, Zhepeng Wang

机构 * Lenovo Research(联想研究院) Tsinghua University(清华大学) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(CAS)(中国科学院自动化研究所) Zhongguancun Academy(中关村学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06918 2025-06-10 cs.CV cs.RO 57%

Reading in the Dark with Foveated Event Vision

Carl Brander, Giovanni Cioffi, Nico Messikommer, Davide Scaramuzza

机构 * Robotics and Perception Group, University of Zurich, Switzerland(苏黎世大学机器人与感知组)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025 Workshop on Event-based Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06359 2025-06-10 cs.LG cs.AI 57%

From Transformers to Large Language Models: A systematic review of AI applications in the energy sector towards Agentic Digital Twins

Gabriel Antonesi, Tudor Cioara, Ionut Anghel, Vasilis Michalakopoulos, Elissaios Sarmas, Liana Toderean

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16652 2025-06-10 cs.CV cs.LG 57%

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding

Feilong Tang, Chengzhi Liu, Zhongxing Xu, Ming Hu, Zelin Peng, Zhiwei Yang, Jionglong Su, Minquan Lin, Yifan Peng, Xuelian Cheng, Imran Razzak, Zongyuan Ge

机构 * Monash University(蒙纳士大学) MBZUAI XJTLU(西安交通大学) Shanghai Jiaotong University(上海交通大学) Fudan University(复旦大学) University of Minnesota(明尼苏达大学) Cornell University(康奈尔大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Clarification note for the CVPR 2025 paper (FarSight). Prepared by a subset of the original authors; remaining co-authors are acknowledged in the text

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06165 2025-06-09 cs.HC cs.AI 57%

(AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation

Eunhye Grace Ko, Soo Hyoung Joo

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06128 2025-06-09 cs.CV 57%

CCLSTM: Coupled Convolutional Long-Short Term Memory Network for Occupancy Flow Forecasting

Peter Lengyel

机构 * aiMotive Budapest, Hungary(aiMotive布达佩斯,匈牙利)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05302 2025-06-06 cs.CV 57%

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Weifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren, Tingwei Chen, Renrui Zhang, Ziyu Guo, Wentao Zhang, Lei Zhang, Hongsheng Li

机构 * CUHK(香港中文大学) HKU(香港大学) PolyU Peking University(北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 19 pages, 13 figures, Website: https://Perceive-Anything.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05163 2025-06-06 cs.CV 57%

FRED: The Florence RGB-Event Drone Dataset

Gabriele Magrini, Niccolò Marini, Federico Becattini, Lorenzo Berlincioni, Niccolò Biondi, Pietro Pala, Alberto Del Bimbo

机构 * University of Florence(佛罗伦萨大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04983 2025-06-06 cs.CV 57%

TextVidBench: A Benchmark for Long Video Scene Text Understanding

Yangyang Zhong, Ji Qi, Yuan Yao, Pengxin Luo, Yunfeng Yan, Donglian Qi, Zhiyuan Liu, Tat-Seng Chua

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03473 2025-06-05 cs.CV 57%

MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval

Xinru Ying, Jiaqi Mo, Jingyang Lin, Canghong Jin, Fangfang Wang, Lina Wei

机构 * Hangzhou City University(杭州城市大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Hangzhou Normal University(杭州师范大学) Zhejiang Provincial Engineering Research Center for Real-Time SmartTech in Urban Security Governance(浙江省实时智能安防技术工程研究中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03350 2025-06-05 cs.RO cs.AI 57%

Adversarial Attacks on Robotic Vision Language Action Models

Eliot Krzysztof Jones, Alexander Robey, Andy Zou, Zachary Ravichandran, George J. Pappas, Hamed Hassani, Matt Fredrikson, J. Zico Kolter

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02891 2025-06-04 cs.CV 57%

OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis

Jiewen Hu, Leena Mathur, Paul Pu Liang, Louis-Philippe Morency

机构 * Carnegie Mellon University(卡内基梅隆大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments IEEE FG 2025, \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02439 2025-06-04 cs.CV 57%

Video-Level Language-Driven Video-Based Visible-Infrared Person Re-Identification

Shuang Li, Jiaxu Leng, Changjiang Kuang, Mingpi Tan, Xinbo Gao

机构 * School of Computer Science and Technology, Chongqing University of Posts and Telecommunications(重庆邮电大学计算机科学与技术学院) Chongqing Institute for Brain and Intelligence(重庆脑科学与智能技术研究院) Guangyang Bay Laboratory(广阳湾实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by IEEE TIFS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20218 2025-06-04 cs.CV 57%

Video Motion Graphs

Haiyang Liu, Zhan Xu, Fa-Ting Hong, Hsin-Ping Huang, Yi Zhou, Yang Zhou

机构 * The University of Tokyo(东京大学) Adobe Research(Adobe研究)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 14 pages,10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01908 2025-06-03 cs.CV 57%

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Hongyu Li, Songhao Han, Yue Liao, Junfeng Luo, Jialin Gao, Shuicheng Yan, Si Liu

机构 * BUAA(北京航空航天大学) NUS(国立大学新加坡) Meituan(美团)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏