VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted by The 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted by The 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM
Comments CVPR2020. Project page: https://anyirao.com/projects/SceneSeg.html
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Comments IEEE International Conference on Computer Vision (ICCV) Workshop on Large Scale Holistic Video Understanding. The datasets and code are available at https://github.com/vivoutlaw/tcbp
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Comments Chapter 3 in A.Gadomski (ed.): Multiscale Locomotion: Its Active-Matter Addressing Physical Principles; UTP University of Science & Technology
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM
Journal ref MediaEval 2019
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 7 pages, 6 figures, submitted to 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments MINOS is a simulator designed to support research on end-to-end navigation
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Comments 4 pages, 3 figures, arXiv:1605.06778, arXiv:1512.03385
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Journal ref ACM - ICMI 2017, Nov 2017, Glasgow, United Kingdom
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 8 pages, 6 figures, accepted by ICRA'17
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Comments 5 pages, 4 figures, ICASSP 2016 accepted
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
Comments Preprint of accepted manuscript for the Elsevier Image and Vision Computing Journal (IMAVIS). The paper will be published by IMAVIS under DOI 10.1016/j.imavis.2015.12.004
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM
Comments Multimodal Survey for Wearable Sensor-based Human Action Recognition
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 22nd International Conference on Image Analysis and Processing Workshops - Multimodal Action Recognition on the MECCANO Dataset, 2023
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Accepted at ICCV 2023; Codebase released at https://github.com/gorjanradevski/multimodal-distillation
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments CVPR 2023 Workshop on Multimodal Learning for Earth and Environment (MultiEarth)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments 8-page extended version of our challenge paper in ACM MM 2021. It presents the overview of grand challenge "Multi-modal Ads Video Understanding" in ACM MM 2021. Our grand challenge is also the Tencent Advertising Algorithm Competition (TAAC) 2021
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments 8 pages, 4 figures, accepted to AAAI 2020, code available at: https://github.com/ztangent/multimodal-dmm
专题命中 视频多模态 :multi-modal(title);multimodal(abstract,comments);分类 cs.CV
Comments This is the accepted version of the following article: Kragh M, Underwood J. Multimodal obstacle detection in unstructured environments with conditional random fields. J Field Robotics. 2019, 1-20., which has been published in final form at https://doi.org/10.1002/rob.21866
专题命中 视频多模态 :multimodal(title,comments);multi-modal(abstract);分类 cs.CV
Comments Keywords Robot-Assisted Surgery, Surgical Gesture Classification, Multi-task Learning, Multimodal Learning, Long Short-term Recurrent Neural Networks, Convolutional Neural Networks
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Multimodal Machine Learning Workshop at NIPS 2015
NARU:用于理解日语超长视频中叙事演变与文化细微差别的基准
机构 * The University of Tokyo(东京大学) ; Kyushu University(九州大学) ; Macau University of Science and Technology(澳门科技大学) ; Infinimind Japan Inc.(Infinimind日本公司) ; University of Alberta(阿尔伯塔大学)
专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI、cs.MM
AI总结 该研究推出NARU基准,涵盖155个146.8小时日语视频的1481个问题,评估模型在长程叙事整合与文化推理上的局限,为MLLM开发提供测试平台。
Comments Yuheng Huang and Jianlang Chen contributed equally to this work. More details available on the project's website https://ma-labo.github.io/naru/ and https://infinimind.io/en/company/news/2026/narubench-release
超越试错:面向图像到视频一致性的智能体优化
机构 * Google Cloud(谷歌云) ; Google DeepMind(谷歌DeepMind)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 针对图像到视频模型试错效率低的问题,提出Agentic Self-Improvement框架,通过两阶段优化提升视频与文本一致性,生成视频胜率达69%,为视频生成模型提供实用可控的优化方法。
HFS: 为高效视频推理的全局查询感知帧选择
机构 * The Hong Kong Polytechnic University(香港理工大学)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
AI总结 HFS提出一种端到端可训练的帧选择框架,通过任务自适应方法提升视频推理效率。
Comments Accepted to the Main Track of ACM Multimedia 2026 (ACM MM '26)
CARVE:用于高效3D医学体积理解的视觉证据跨切片各向异性重分配
专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.CL、cs.AI
AI总结 针对3D医学体积理解中切片式MLLM的视觉令牌冗余问题,提出无需训练的CARVE框架,通过跨切片各向异性重分配压缩80%令牌,在AMOS-MM等基准上性能优于现有方法。