arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2507.20763 2025-07-29 cs.CV 57%

KASportsFormer: Kinematic Anatomy Enhanced Transformer for 3D Human Pose Estimation on Short Sports Scene Video

Zhuoer Yin, Calvin Yeung, Tomohiro Suzuki, Ryota Tanaka, Keisuke Fujii

机构 * Nagoya University(名古屋大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19948 2025-07-29 cs.CV 57%

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

Luoxi Jing, Dianxi Shi, Zhe Liu, Songchang Jin, Chunping Qiu, Ziteng Qiao, Yuxian Li, Jianqiang Xia

机构 * School of Computer Science, Peking University(北京大学计算机科学系) Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室) College of Computer, National University of Defense Technology(国防科技大学计算机学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IJCAI 2025 (International Joint Conference on Artificial Intelligence)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19863 2025-07-29 cs.MM 57%

Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and Expansion

Chia-Ming Lee, Bo-Cheng Qiu, Cheng-Jun Kang, Yi-Hsuan Wu, Jun-Lin Chen, Yu-Fan Lin, Yi-Shiuan Chou, Chih-Chung Hsu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01301 2025-07-29 cs.RO cs.AI 57%

Bi-LAT: Bilateral Control-Based Imitation Learning via Natural Language and Action Chunking with Transformers

Takumi Kobayashi, Masato Kobayashi, Thanpimon Buamanee, Yuki Uranishi

机构 * Graduate School of Information Science and Technology, The University of Osaka(信息科学与技术研究生学校,大阪大学) D3 Center, The University of Osaka(大阪大学D3中心) Graduate School of Maritime Sciences, Kobe University(海洋科学研究生学校, Kobe大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19599 2025-07-29 cs.CV 57%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03492 2025-07-29 cs.CV 57%

Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation

Suhwan Cho, Seunghoon Lee, Minhyeok Lee, Jungho Lee, Sangyoun Lee

机构 * GenGenAI Yonsei University(延世大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ICCVW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18634 2025-07-25 cs.CV 57%

Captain Cinema: Towards Short Movie Generation

Junfei Xiao, Ceyuan Yang, Lvmin Zhang, Shengqu Cai, Yang Zhao, Yuwei Guo, Gordon Wetzstein, Maneesh Agrawala, Alan Yuille, Lu Jiang

机构 * Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Under review. Project page: https://thecinema.ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18342 2025-07-25 cs.CV 57%

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Yuping He, Yifei Huang, Guo Chen, Baoqi Pei, Jilan Xu, Tong Lu, Jiangmiao Pang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Tokyo(东京大学) Zhejiang University(浙江大学) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00566 2025-07-25 cs.CV 57%

Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment

Kai Zhou, Shuhai Zhang, Zeng You, Jinwu Hu, Mingkui Tan, Fei Liu

机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) South China University of Technology(南方科技大学) Pazhou Lab(琶洲实验室) School of Future Technology, South China University of Technology(未来技术学院) Peng Cheng Laboratory(鹏城实验室) Key Laboratory of Big Data and Intelligent Robot (South China University of Technology), Ministry of Education(大数据与智能机器人重点实验室)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments This paper is accepted by IEEE TIP 2025 (The journal version is available at https://doi.org/10.1109/TIP.2025.3586487). Code is publicly available at https://github.com/kaai520/PGFA

Journal ref IEEE Transactions on Image Processing 34 (2025) 4602-4617

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16363 2025-07-23 cs.LG cs.MM 57%

Bipartite Patient-Modality Graph Learning with Event-Conditional Modelling of Censoring for Cancer Survival Prediction

Hailin Yue, Hulin Kuang, Jin Liu, Junjian Li, Lanlan Wang, Mengshen He, Jianxin Wang

机构 * Hunan Provincial Key Lab on Bioinformatics, School of Computer Science and Engineering, Central South University(湖南省级生物信息学重点实验室,计算机科学与工程学院,中南大学) Xinjiang Engineering Research Center of Big Data and Intelligent Software, School of Software, Xinjiang University(新疆大数据与智能软件工程研究中心,软件学院,新疆大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22139 2025-07-23 cs.CV 57%

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Shaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo, Jian Luan

机构 * MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14301 2025-07-22 cs.IR cs.CV cs.DB 57%

LOVO: Efficient Complex Object Query in Large-Scale Video Datasets

Yuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu, Xinyu Lu, Jiong Lou, Chentao Wu, Wei Zhao, Jie Li

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments @inproceedings{liu2025lovo,title={LOVO: Efficient Complex Object Query in Large-Scale Video Datasets},author={Liu, Yuxin and Peng, Yuezhang and Zhou, Hefeng and Liu, Hongze and Lu, Xinyu and Lou, Jiong and Wu, Chentao and Zhao, Wei and Li, Jie},booktitle={2025 IEEE 41st International Conference on Data Engineering (ICDE)},pages={1938--1951},year={2025},organization={IEEE Computer Society}}

Journal ref 2025 IEEE 41st International Conference on Data Engineering (ICDE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12628 2025-07-18 cs.CV 57%

Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection

Sandipan Sarma, Agney Talwarr, Arijit Sur

机构 * Department of Computer Science and Engineering, Indian Institute of Technology, Guwahati, Assam(计算机科学与工程系,印度理工学院,古瓦哈蒂,阿萨姆)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23765 2025-07-18 cs.CV 57%

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Yun Li, Yiming Zhang, Tao Lin, Xiangrui Liu, Wenxiao Cai, Zheng Liu, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) China University of Geosciences(中国地质大学) Nanyang Technological University(南洋理工大学) BAAI(百度人工智能研究院) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11579 2025-07-17 cs.CV 57%

Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers

Weiming Ren, Wentao Ma, Huan Yang, Cong Wei, Ge Zhang, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) University of Toronto(多伦多大学) Kuaishou Technology(快手科技) Vector Institute(向量研究所) M-A-P

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025 Camera Ready Version. Project Page: https://tiger-ai-lab.github.io/Vamba/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10745 2025-07-17 cs.CV 57%

Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition

Jeonghyeok Do, Munchurl Kim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025 (camera-ready version). Please visit our project page at https://kaist-viclab.github.io/TDSM_site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20254 2025-07-16 cs.CV 57%

Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement

Kun Yuan, Tingxuan Chen, Shi Li, Joel L. Lavanchy, Christian Heiliger, Ege Özsoy, Yiming Huang, Long Bai, Nassir Navab, Vinkle Srivastav, Hongliang Ren, Nicolas Padoy

机构 * University of Strasbourg(斯特拉斯堡大学) CNRS(法国国家科学研究中心) INSERM(法国国家健康与医学研究院) ICube(ICube研究中心) UMR7357(法国大学-研究中心7357) IHU Strasbourg(斯特拉斯堡IHU医院) University Digestive Health Care Center – Clarunis(大学消化健康中心 – Clarunis) Ludwig Maximilian University of Munich(慕尼黑路易斯·马克西米利安大学) Chinese University of Hong Kong(香港中文大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09403 2025-07-15 cs.IR cs.MM 57%

Balancing Semantic Relevance and Engagement in Related Video Recommendations

Amit Jaspal, Feng Zhang, Wei Chang, Sumit Kumar, Yubo Wang, Roni Mittleman, Qifan Wang, Weize Mao

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05357 2025-07-15 eess.IV cs.CV 57%

Unmixing Optical Signals from Undersampled Volumetric Measurements by Filtering the Pixel Latent Variables

Catherine Bouchard, Andréanne Deschênes, Vincent Boulanger, Jean-Michel Bellavance, Julia Chabbert, Alexy Pelletier-Rioux, Flavie Lavoie-Cardinal, Christian Gagné

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 42 pages, 9 figures (main paper) + 22 pages, 15 figures (supplementary material)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17224 2025-07-14 cs.CV 57%

Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition

Zhifeng Wang, Qixuan Zhang, Peter Zhang, Wenjia Niu, Kaihao Zhang, Ramesh Sankaranarayana, Sabrina Caldwell, Tom Gedeon

机构 * School of Computing, Australian National University (ANU)(澳大利亚国立大学计算机学院) Quriosity Pty Ltd Human-Centric Advancements Chair in AI, Curtin University and Australian National University(Curtin大学人工智能人本发展主席职位,澳大利亚国立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07415 2025-07-11 cs.CV 57%

EPIC: Efficient Prompt Interaction for Text-Image Classification

Xinyao Yu, Hao Sun, Zeyu Ling, Ziwei Niu, Zhenjia Bai, Rui Qin, Yen-Wei Chen, Lanfen Lin

机构 * College of Computer Science and Technology, Zhejiang University, Hangzhou, China(浙江大学计算机科学与技术学院) College of Information Science, Ritsumeikan University, Shiga, Japan(立命馆大学信息科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2401.14856

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07262 2025-07-11 cs.CV 57%

DisenQ: Disentangling Q-Former for Activity-Biometrics

Shehreen Azad, Yogesh S Rawat

机构 * Center for Research in Computer Vision(计算机视觉研究中心) University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06531 2025-07-10 cs.CV 57%

ILNet: Trajectory Prediction with Inverse Learning Attention for Enhancing Intention Capture

Mingjin Zeng, Nan Ouyang, Wenkang Wan, Lei Ao, Qing Cai, Kai Sheng

机构 * Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University(协同智能系统重点实验室,教育部,西安电子科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05336 2025-07-08 cs.CV 57%

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani, Abdelrahman Shaker, Zhiqiang Shen, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Linköping University(林奈大学) Australian National University(澳大利亚国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13805 2025-07-08 cs.CV 57%

mmEgoHand: Egocentric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMU

Yizhe Lv, Tingting Zhang, Zhijian Wang, Yunpeng Song, Han Ding, Jinsong Han, Fei Wang

机构 * School of Software Engineering, Xi’an Jiaotong University(软件工程学院,西安交通大学) School of Cyber Science and Engineering, Xi’an Jiaotong University(网络科学与工程学院,西安交通大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 11 pages, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01492 2025-07-03 cs.CV 57%

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

Jiyang Tang, Hengyi Li, Yifan Du, Wayne Xin Zhao

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) College of Artificial Intelligence, Nankai University(南开大学人工智能学院) School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14607 2025-07-01 cs.CV 57%

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang, Wei-Shi Zheng, Jian-Fang Hu

机构 * Sun Yat-sen University(中山大学) Southern University of Science and Technology(南方科技大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Project page: \url{https://isee-laboratory.github.io/ReferDINO}

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17690 2025-07-01 cs.CV 57%

CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model

Ziyu Yao, Xuxin Cheng, Zhiqi Huang, Lei Li

机构 * Peking University(北京大学) University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10432 2025-06-30 cs.LG cs.CL 57%

BeamLLM: Vision-Empowered mmWave Beam Prediction with Large Language Models

Can Zheng, Jiguang He, Guofa Cai, Zitong Yu, Chung G. Kang

机构 * School of Electrical Engineering, Korea University(韩国大学电气工程学院) School of Computing and Information Technology, Great Bay University(大湾大学计算与信息科技学院) School of Information Engineering, Guangdong University of Technology(广东技术大学信息工程学院)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL

Comments 6 pages, 7 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21317 2025-06-27 cs.CV 57%

LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning

Dewen Zhang, Tahir Hussain, Wangpeng An, Hayaru Shouno

机构 * Department of Informatics, Graduate School of Informatics and Engineering, The University of Electro-Communications(信息学院、信息工程研究生院、东京电波通信大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2409.09306

详情

展开后加载摘要…

URL PDF HTML 收藏