arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4726 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4726 篇

2505.02060 2025-05-09 cs.CV 70%

Transforming faces into video stories -- VideoFace2.0

Branko Brkljač, Vladimir Kalušev, Branislav Popović, Milan Sečujski

机构 * Department of Power, Electronic and Telecommunication Engineering(电力、电子与电信工程系) Faculty of Technical Sciences, University of Novi Sad(技术科学学院,诺维萨德大学) Visual Computing & Perception Group(视觉计算与感知组) The Institute for Artificial Intelligence Research and Development of Serbia(塞尔维亚人工智能研究与开发研究所)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments 4 Pages, 2 Figures, 1 Table, 1 Algorithm; Associated VideoFace2.0 code, test videos and results visualizations are available at https://github.com/brkljac/VideoFace2.0 ; Preprint accepted for publication at the 14th Mediterranean Conference on Embedded Computing (MECO), 10-14 June 2025, Budva, Montenegro

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21189 2025-05-01 cs.LG cs.AI cs.ET 70%

Artificial Intelligence for Personalized Prediction of Alzheimer's Disease Progression: A Survey of Methods, Data Challenges, and Future Directions

Gulsah Hancerliogullari Koksalmis, Bulent Soykan, Laura J. Brattain, Hsin-Hsiung Huang

机构 * Department of Industrial Engineering and Management Systems University of Central Florida(工业工程与管理系统系 佛罗里达中央大学) Institute for Simulation and Training University of Central Florida(模拟与培训研究所 佛罗里达中央大学) Department of Internal Medicine University of Central Florida(内科医学系 佛罗里达中央大学) Department of Statistics and Data Science University of Central Florida(统计学与数据科学系 佛罗里达中央大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

Comments 25 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11949 2025-04-17 cs.CV 70%

Flow Intelligence: Robust Feature Matching via Temporal Signature Correlation

Jie Wang, Chen Ye Gan, Caoqi Wei, Jiangtao Wen, Yuxing Han

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00672 2025-04-15 cs.CV 70%

ExpertAF: Expert Actionable Feedback from Video

Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani, Kristen Grauman

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07962 2025-04-11 cs.CV 70%

GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation

Lang Lin, Xueyang Yu, Ziqi Pang, Yu-Xiong Wang

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07519 2025-04-11 cs.CV 70%

VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

Henghao Zhao, Ge-Peng Ji, Rui Yan, Huan Xiong, Zechao Li

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00513 2025-03-19 cs.CV cs.IR cs.LG 70%

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17053 2025-03-18 cs.CV 70%

Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding

Akash Kumar, Zsolt Kira, Yogesh Singh Rawat

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments ICLR'25 Main Conference. Project Page: https://akash2907.github.io/cospal_webpage

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11412 2025-03-17 cs.CV 70%

MTV-Inpaint: Multi-Task Long Video Inpainting

Shiyuan Yang, Zheng Gu, Liang Hou, Xin Tao, Pengfei Wan, Xiaodong Chen, Jing Liao

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07653 2025-03-12 cs.LG cs.CL cs.SI 70%

Early Detection of Mental Health Issues Using Social Media Posts

Qasim Bin Saeed, Ijaz Ahmed

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09367 2025-03-10 cs.CV 70%

Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, Jing Liu

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13056 2025-03-07 cs.CV 70%

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

Bing Cao, Quanhao Lu, Jiekang Feng, Qilong Wang, Qinghua Hu, Pengfei Zhu

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICLR25

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18143 2025-02-26 cs.CV 70%

LightFC-X: Lightweight Convolutional Tracker for RGB-X Tracking

Yunfeng Li, Bo Wang, Ye Li

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08234 2025-02-13 cs.CV 70%

Learning Human Skill Generators at Key-Step Levels

Yilu Wu, Chenhui Zhu, Shuai Wang, Hanlin Wang, Jing Wang, Zhaoxiang Zhang, Limin Wang

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18940 2025-02-03 cs.CV 70%

TV-Dialogue: Crafting Theme-Aware Video Dialogues with Immersive Interaction

Sai Wang, Fan Ma, Xinyi Li, Hehe Fan, Yu Wu

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05108 2025-01-10 cs.CV 70%

Optimizing Multitask Industrial Processes with Predictive Action Guidance

Naval Kishore Mehta, Arvind, Shyam Sunder Prasad, Sumeet Saurav, Sanjay Singh

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01245 2025-01-03 cs.CV cs.LG 70%

SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization

Yongle Huang, Haodong Chen, Zhenbang Xu, Zihan Jia, Haozhou Sun, Dian Shao

专题命中 视频多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments AAAI 2025; Code: https://github.com/KyleHuang9/SeFAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20964 2025-01-01 cs.CV 70%

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

Peng Jin, Hao Li, Li Yuan, Shuicheng Yan, Jie Chen

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). arXiv admin note: substantial text overlap with arXiv:2303.14369

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20613 2024-12-31 cs.CV 70%

Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study

Yulin Fei, Yuhui Gao, Xingyuan Xian, Xiaojin Zhang, Tao Wu, Wei Chen

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by CoLing 2025 (The 31st International Conference on Computational Linguistics)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06878 2024-12-11 cs.CV cs.LG 70%

SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations

Zhaorun Chen, Francesco Pinto, Minzhou Pan, Bo Li

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 43 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07945 2024-11-13 cs.CV 70%

SimBase: A Simple Baseline for Temporal Video Grounding

Peijun Bao, Alex C. Kot

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05774 2024-11-12 cs.CV 70%

ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition

Mohammadreza Salehi, Jae Sung Park, Tanush Yadav, Aditya Kusupati, Ranjay Krishna, Yejin Choi, Hannaneh Hajishirzi, Ali Farhadi

专题命中 视频多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Journal ref NeurIPS 2024 Track Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21958 2024-10-30 cs.CV 70%

Spatio-temporal Transformers for Action Unit Classification with Event Cameras

Luca Cultrera, Federico Becattini, Lorenzo Berlincioni, Claudio Ferrari, Alberto Del Bimbo

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Under review at CVIU. arXiv admin note: substantial text overlap with arXiv:2409.10213

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17434 2024-10-24 cs.CV 70%

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, Zhuang Liu, Hu Xu, Hyunwoo J. Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, Vikas Chandra

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Project page: https://vision-cair.github.io/LongVU

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03038 2024-10-08 cs.LG cs.CV 70%

CPFD: Confidence-aware Privileged Feature Distillation for Short Video Classification

Jinghao Shi, Xiang Shen, Kaili Zhao, Xuedong Wang, Vera Wen, Zixuan Wang, Yifan Wu, Zhixin Zhang

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Camera ready for CIKM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01144 2024-10-03 cs.CV 70%

Uncertainty-Guided Enhancement on Driving Perception System via Foundation Models

Yunhao Yang, Yuxin Hu, Mao Ye, Zaiwei Zhang, Zhichao Lu, Yi Xu, Ufuk Topcu, Ben Snyder

专题命中 视频多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11432 2024-08-22 cs.CV 70%

T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval

Yili Li, Jing Yu, Keke Gai, Bang Liu, Gang Xiong, Qi Wu

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02336 2024-08-07 cs.CV cs.LG 70%

Infusing Environmental Captions for Long-Form Video Language Grounding

Hyogun Lee, Soyeon Hong, Mujeen Sung, Jinwoo Choi

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19981 2024-07-30 cs.CV 70%

Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter

Chao Liu, Xin Liu, Zitong Yu, Yonghong Hou, Huanjing Yue, Jingyu Yang

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by IJCB 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06628 2024-07-10 cs.CV 70%

Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition

Mingfang Zhang, Yifei Huang, Ruicong Liu, Yoichi Sato

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏