arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2601.11220 2026-01-19 cs.CL 57%

MultiCaption: Detecting disinformation using multilingual visual claims

多 caption:利用多语言视觉主张检测虚假信息

Rafael Martins Frade, Rrubaa Panchendrarajan, Arkaitz Zubiaga

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

AI总结 MultiCaption通过多语言视觉主张数据集提升多模态虚假信息检测性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02937 2026-01-19 cs.LG cs.AI 57%

Towards Explainable Traffic Flow Prediction with Large Language Models

面向大语言模型的可解释交通流预测

Xusen Guo, Qiming Zhang, Junyue Jiang, Mingxing Peng, Meixin Zhu, Hao, Yang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)) Johns Hopkins University(约翰霍普金斯大学) Department of Civil and System Engineering(土木与系统工程系)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出xTP-LLM模型,利用大语言模型生成可解释的交通流预测,首次将LLM应用于交通预测的可解释性研究。

Comments 31pages, 16 figures

Journal ref Communications in Transportation Research, vol. 4, 100150, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09430 2026-01-15 cs.CV 57%

Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs

Video-MSR: 多跳空间推理能力的MLLMs基准测试

Rui Zhu, Xin Shen, Shuchen Wu, Chenxi Miao, Xin Yu, Yang Li, Weikang Li, Deguo Xia, Jizhou Huang

机构 * Baidu Inc.(百度公司) Nanjing University(南京大学) The University of Queensland(昆士兰大学) Peking University(北京大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 Video-MSR通过四个任务评估MLLMs的多跳空间推理能力,发现模型在复杂推理中存在显著不足,并通过微调提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09350 2026-01-15 cs.CV 57%

See More, Store Less: Memory-Efficient Resolution for Video Moment Retrieval

看清更多,存储更少:视频时刻检索的内存高效解决方案

Mingyu Jeon, Sungjin Han, Jinkwon Hwang, Minchol Kwon, Jonghee Kim, Junyeong Kim

机构 * Department of Artificial Intelligence, Chung-Ang University(人工智能系, Chung-Ang 大学) Electronics and Telecommunications Research Institute (ETRI)(电子与电信研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 SMORE通过查询引导的描述和自适应压缩技术,在视频时刻检索中实现高信息分辨率与内存效率的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09031 2026-01-15 cs.RO cs.AI 57%

Generalizable Geometric Prior and Recurrent Spiking Feature Learning for Humanoid Robot Manipulation

可推广的几何先验与递归脉冲特征学习用于双足机器人操作

Xuetao Li, Wenke Huang, Mang Ye, Jifeng Xuan, Bo Du, Sheng Liu, Miao Li

机构 * School of Computer Science, School of Robotics, Institute of Technological Sciences, Wuhan University(计算机学院、机器人学院、技术科学研究院、武汉大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出RGMP-S方法,通过几何先验和递归脉冲网络提升双足机器人操作的泛化能力与数据效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08587 2026-01-15 cs.CV 57%

MoCha:End-to-End Video Character Replacement without Structural Guidance

MoCha:端到端视频人物替换无需结构引导

Zhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng, Jun Liang, Jing Li

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 MoCha提出了一种无需结构引导的端到端视频人物替换方法,通过单帧掩码实现高效替换,并设计了三种数据集以克服数据稀缺问题。

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07366 2026-01-13 cs.CV 57%

HiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression

HiVid-Narrator:基于场景优先的ASR锚定压缩的分层视频叙事生成

Haoxuan Li, Mengyan Li, Junjun Zheng

机构 * Taobao & Tmall Group of Alibaba(淘宝与天猫集团)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 HiVid-Narrator通过分阶段构建和SPA-Compressor压缩技术,在减少输入标记的同时生成高质量的电子商务视频叙事。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07290 2026-01-13 cs.CV 57%

VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding

VideoLoom:一种用于联合空间-时间理解的视频大语言模型

Jiapeng Shi, Junke Wang, Zuyao You, Bo He, Zuxuan Wu

机构 * Fudan University(复旦大学) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 VideoLoom是一种用于联合空间-时间理解的视频大语言模型,通过LoomData-8.7k数据集和LoomBench基准测试,在多个视频理解任务中取得优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04982 2026-01-09 cs.RO cs.AI 57%

When to Act: Calibrated Confidence for Reliable Human Intention Prediction in Assistive Robotics

何时行动:校准信心以在辅助机器人中实现可靠的意图预测

Johannes A. Gaus, Winfried Ilg, Daniel Haeufle

机构 * Hertie Institute for Clinical Brain Research & Center for Integrative Neuroscience, University of Tübingen(海德堡临床脑研究所及整合神经科学中心,图宾根大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于校准概率的安全触发框架,用于提升辅助机器人中意图预测的可靠性,通过校准信心减少误校准并提高安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04891 2026-01-09 cs.CV cs.LG 57%

Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform

在工业生成式AI平台上扩展视觉语言模型用于制药长格式视频推理

Suyash Mishra, Qiang Li, Srikanth Patil, Satyanarayan Pati, Baddu Narendra

机构 * Roche(罗氏) Accenture(埃森哲) Involead

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出工业GenAI框架,解决制药领域长格式视频推理问题,通过多模态架构和实证分析提升效率并揭示现有VLMs的限制。

Comments Submitted to the Industry Track of Top Tier Conference; currently under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17229 2026-01-06 cs.CV 57%

Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos

视频侦探:通过反复寻找关键线索来回答长视频中的问题

Henghui Du, Chunjie Zhang, Xi Chen, Chang Zhou, Di Hu

机构 * Gaoling School of Artificial Intelligence(北京中国人民大学人工智能学院) Renmin University of China(中国人民大学) AI Technology Center Online Video Business Unit(人工智能技术中心在线视频业务部) Tencent PCG(腾讯PCG)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 VideoDetective通过问题感知的记忆机制,使大语言模型能高效处理长视频问答任务,减少计算资源消耗并提升关键信息提取能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14960 2026-01-05 cs.CV 57%

Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition

具有交叉注意力的体-手模态专家化网络用于细粒度骨骼动作识别

Seungyeon Cho, Tae-Kyun Kim

机构 * School of Computing, KAIST(韩国成均馆大学计算机学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出BHaRNet,通过体-手专家化网络和交叉注意力机制,提升细粒度骨骼动作识别的准确率,同时降低计算成本。

Comments 7 figures, 8 pages. Accepted to IROS 2025, project page: https://github.com/VinnyCSY/BHaRNet

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hangzhou, China, 2025, pp. 11614-11621

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24834 2026-01-01 cs.AI 57%

GenZ: Foundational models as latent variable generators within traditional statistical models

GenZ:在传统统计模型中通过可解释的语义特征作为潜在变量生成器的基础模型

Marko Jojic, Nebojsa Jojic

机构 * Arizona State University(亚利桑那州立大学) Microsoft Research(微软研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 GenZ通过可解释的语义特征连接基础模型与统计模型,实现更精准的预测,其在房地产价格预测和电影推荐中均取得显著成效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13415 2025-12-30 cs.CV 57%

USTM: Unified Spatial and Temporal Modeling for Continuous Sign Language Recognition

USTM:统一的空间和时间建模用于连续手语识别

Ahmed Abul Hasanaath, Hamzah Luqman

机构 * College of Information and Computer Science(信息与计算机科学学院) King Fahd University of Petroleum and Minerals(国王法赫德石油矿物大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 USTM通过结合Swin Transformer与轻量级时间适配器,实现高效的空间-时间建模,提升连续手语识别的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19537 2025-12-23 cs.CL 57%

Event Extraction in Large Language Model

大规模语言模型中的事件提取

Bobo Li, Xudong Han, Jiang Liu, Yuzhe Ding, Liqiang Jing, Zhaoqi Zhang, Jinheng Li, Xinya Du, Fei Li, Meishan Zhang, Min Zhang, Aixin Sun, Philip S. Yu, Hao Fei

机构 * National University of Singapore(新加坡国立大学) University of Sussex(苏塞克斯大学) Wuhan University(武汉大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) Nanyang Technological University(南洋理工大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文探讨了大规模语言模型中事件提取的系统组件,提出通过事件方案、结构和存储提升事件处理的可靠性与智能体准备性。

Comments 38 pages, 9 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19048 2025-12-23 cs.CV 57%

WaTeRFlow: Watermark Temporal Robustness via Flow Consistency

WaTeRFlow:通过流一致性实现水印时间鲁棒性

Utae Jeong, Sumin In, Hyunju Ryu, Jaewan Choi, Feng Yang, Jongheon Jeong, Seungryong Kim, Sangpil Kim

机构 * Korea University(韩国大学) Google DeepMind(谷歌DeepMind) KAIST AI(韩国科学技术院人工智能)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

AI总结 WaTeRFlow通过流一致性提升I2V场景下的水印鲁棒性,结合FUSE引擎、时间一致性损失和语义保持损失,实现跨模态水印恢复。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24511 2025-12-23 cs.LG cs.AI 57%

Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting

慢思考LLM能否在时间上进行推理?时间序列预测的实证研究

Mingyue Cheng, Jiahao Wang, Daoyu Wang, Xiaoyu Tao, Qi Liu, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨慢思考LLM能否在时间序列预测中进行推理,通过实验证明其零样本预测能力,揭示LLM在时间动态上的潜力与局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17020 2025-12-23 cs.CV 57%

CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms

CrossLMM: 通过双交叉注意力机制解耦长视频序列与LMMs

Shilin Yan, Jiaming Han, Joey Tsai, Hongwei Xue, Rongyao Fang, Lingyi Hong, Ziyu Guo, Ray Zhang

机构 * Accio Team, Alibaba Group(阿里集团阿奇奥团队) CUHK MMLab(CUHK多模态实验室) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 CrossLMM通过双交叉注意力机制解耦长视频序列,有效减少视觉token数量并保持性能,适用于多种视频多模态模型基准测试。

Comments Project page: https://github.com/shilinyan99/CrossLMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17958 2025-12-23 cs.RO cs.AI 57%

Real-Time Human-Robot Interaction Intent Detection Using RGB-based Pose and Emotion Cues with Cross-Camera Model Generalization

基于RGB姿态和情绪线索的实时人机交互意图检测及跨摄像头模型泛化

Farida Mohsen, Ali Safa

机构 * College of Science and Engineering, Hamad Bin Khalifa University(科学与工程学院,哈马德·本·卡西姆大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于RGB姿态和情绪线索的实时人机交互意图检测方法,通过MINT-RVAE生成意图序列,实现跨摄像头和环境的强泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16924 2025-12-19 cs.CV 57%

The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

世界是你的画布:通过参考图像、轨迹和文本绘画可提示事件

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Yue Yu, Yihao Meng, Wen Wang, Ka Leong Cheng, Shuailei Ma, Qingyan Bai, Yixuan Li, Cheng Chen, Yanhong Zeng, Xing Zhu, Yujun Shen, Qifeng Chen

机构 * HKUST(香港科技大学) Ant Group(蚂蚁集团) ZJU(浙江大学) NEU(南京大学) CUHK(香港中文大学) NTU(南洋理工大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 WorldCanvas通过结合文本、轨迹和参考图像,实现可提示的多代理交互事件生成,提升世界模型的交互性和可控性。

Comments Project page and code: https://worldcanvas.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16461 2025-12-19 cs.CV cs.RO 57%

SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning

SNOW:基于世界知识的时空场景理解用于开放世界具身推理

Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Esslingen University of Applied Sciences(埃斯林根应用科学大学) Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业联合会) University of Michigan(密歇根大学) Voxel51 Inc.(Voxel51公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 SNOW通过整合视觉语言模型与点云几何,实现统一的4D场景理解,提升开放世界具身推理的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13238 2025-12-16 cs.CV 57%

Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance

Ego-EXTRA:视频语言眼动数据集用于专家-学员协助

Francesco Ragusa, Michele Mazzamuto, Rosario Forte, Irene D'Ambra, James Fort, Jakob Engel, Antonino Furnari, Giovanni Maria Farinella

机构 * Department of Mathematics and Computer Science - University of Catania(数学与计算机科学系 - 卡塔尼亚大学) Next Vision s.r.l. - Spinoff of the University of Catania(Next Vision公司 - 卡塔尼亚大学衍生机构) Meta Reality Labs Research(Meta现实实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 Ego-EXTRA数据集通过专家-学员双向对话评估多模态大语言模型,揭示其在专家级帮助中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13015 2025-12-16 cs.CV 57%

What Happens Next? Next Scene Prediction with a Unified Video Model

接下来会发生什么?一种统一的视频模型用于下一场景预测

Xinjie Li, Zhimin Chen, Rui Zhao, Florian Schiffers, Zhenyu Liao, Vimal Bhat

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出Next Scene Prediction任务,通过联合框架结合Qwen-VL和LTX,利用预训练、监督微调和强化学习提升视频模型的时间推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12462 2025-12-16 cs.LG cs.AI q-bio.NC 57%

Dynamical modeling of nonlinear latent factors in multiscale neural activity with real-time inference

多尺度神经活动中非线性潜在因子的动力学建模

Eray Erturk, Maryam M. Shanechi

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) University of Southern California(南加州大学) Department of Computer Science(计算机科学系) Department of Biomedical Engineering(生物医学工程系) Neuroscience Graduate Program(神经科学研究生项目)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本研究提出了一种多尺度动态框架,用于实时解码多模态神经活动,通过非线性聚合不同时间尺度和分布的数据以提升解码性能。

Comments Published at the 39th Annual Conference on Neural Information Processing Systems 2025. Code is available at https://github.com/ShanechiLab/mrine

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06080 2025-12-16 cs.CV 57%

CAST-Phys: Contactless Affective States Through Physiological signals Database

CAST-Phys: 通过生理信号进行无接触情绪状态的数据库

Joaquim Comas, Alexander Joel Vera, Xavier Vives, Eleonora De Filippi, Alexandre Pereda, Federico Sukno

机构 * Department of Information and Communication Technologies, Pompeu Fabra University(信息与通信技术系,庞培法布拉大学) Eurecat Centre Tecnològic(埃鲁卡特技术中心)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 CAST-Phys数据库通过多模态远程生理信号实现无接触情绪识别,提供高质量的生理和面部数据以提升情绪识别技术。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00696 2025-12-16 q-bio.MN cs.AI cs.ET 57%

Hierarchical Molecular Language Models (HMLMs)

分层分子语言模型(HMLMs)

Hasi Hays, Yue Yu, William J. Richardson

机构 * Department of Chemical Engineering, University of Arkansas(化学工程系,亚拉巴马大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 HMLMs通过构建分子语言模型,将细胞信号传递建模为分子语言,利用分层注意力机制和跨尺度操作符整合多模态数据,提升对复杂信号网络的时间动态预测能力,推动精准医学发展。

Comments The current version includes minor revisions to the preprint v2 (arXiv preprint arXiv:2512.00696), Added the Supplementary materials section

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11510 2025-12-15 cs.CV 57%

Reconstruction as a Bridge for Event-Based Visual Question Answering

重建作为事件驱动视觉问答的桥梁

Hanyue Lou, Jiayi Zhou, Yang Zhang, Boyu Li, Yi Wang, Guangnan Ye, Boxin Shi

机构 * Peking University(北京大学) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于重建的框架,通过FRT和ART方法提升事件驱动视觉问答性能,并引入EvQA基准验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11189 2025-12-15 cs.CV 57%

Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization

多任务学习与扩展时间移位模块用于时间动作定位

Anh-Kiet Duong, Petra Gomez-Krämer

机构 * L3i Laboratory, La Rochelle University(拉罗谢尔大学L3i实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出基于扩展时间移位模块的多任务学习方法,用于多视角多模态视频中的时间动作定位,通过引入背景类别和加权集成策略提升了预测的鲁棒性和一致性。

Comments BinEgo360@ICCV25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23158 2025-12-15 cs.LG cs.AI 57%

Deep Learning-Based Detection of Cognitive Impairment from Passive Smartphone Sensing with Routine-Aware Augmentation and Demographic Personalization

基于深度学习的被动智能手机传感认知障碍检测:具有常规感知增强和人口统计学个性化

Yufei Shen, Ji Hwan Park, Minchao Huang, Jared F. Benge, Justin F. Rousseau, Rosemary A. Lester-Smith, Edison Thomaz

机构 * Department of Electrical and Computer Engineering, Cockrell School of Engineering, The University of Texas at Austin(电气与计算机工程系,Cockrell工程学院,德克萨斯大学奥斯汀分校) Department of Neurology, Dell Medical School, The University of Texas at Austin(神经病学系,德克萨斯医学学院,德克萨斯大学奥斯汀分校) Department of Neurology, University of Texas Southwestern Medical Center(神经病学系,德克萨斯西南医学中心) Peter O’Donnell Jr. Brain Institute, University of Texas Southwestern Medical Center(彼得·奥·唐纳德·杰罗姆脑研究所,德克萨斯西南医学中心) Department of Speech, Language, and Hearing Sciences, Moody College of Communication, The University of Texas at Austin(语言病理学系,摩依学院,德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出基于深度学习的被动智能手机传感方法,通过常规感知增强和人口统计学个性化技术,提高模型在老年人认知障碍检测中的泛化能力。

Comments Accepted at 2025 IEEE EMBS International Conference on Biomedical and Health Informatics (IEEE BHI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09354 2025-12-11 cs.CV 57%

Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding

Video-QTR: 一种基于查询的轻量级视频理解时序推理框架

Xinkui Zhao, Zuxin Wang, Yifan Zhang, Guanjie Cheng, Yueshen Xu, Shuiguang Deng, Chang Liu, Naibo Wang, Jianwei Yin

机构 * School of Software Technology, Zhejiang University, Hangzhou, China(浙江大学软件技术学院) School of Computer Science, Zhejiang University, Hangzhou, China(浙江大学计算机科学学院) School of Software Engineering, Xidian University, Xi’an, China(西安电子科技大学软件工程学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 Video-QTR通过查询驱动的时序推理框架,实现轻量级视频理解,显著降低计算成本并提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏