VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
机构 * Nanjing University(南京大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Honor Device Co., Ltd(荣誉设备有限公司)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Nanjing University(南京大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Honor Device Co., Ltd(荣誉设备有限公司)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
机构 * The College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; WeChat Vision, Tencent Inc(腾讯公司微信视觉团队)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
机构 * Kunlun Inc.(昆仑公司)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Lenovo Research(联想研究院) ; University of Chinese Academy of Sciences(中国科学院大学) ; Tsinghua University(清华大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) ; Baidu Inc.(百度公司) ; University of Sydney(悉尼大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Australian National University(澳大利亚国立大学) ; GVC Lab, Great Bay University(大湾大学GVC实验室) ; Intellindust AI Lab(Intellindust AI实验室)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Meituan Project(美团项目)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments 11 pages, 4 figures
机构 * Qingdao Institute of Software, College of Computer Science and Technology, China University of Petroleum (East China)(青岛软件研究所,计算机科学与技术学院,中国石油大学(华东)) ; Thrust of AI, Hong Kong University of Science and Technology (GuangZhou)(人工智能研究组,香港科学与技术大学(广州)) ; Shandong Key Laboratory of Intelligent Oil & Gas Industrial Software(山东省智能油气工业软件重点实验室) ; Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学) ; School of Automation, Southeast University(自动化学院,东南大学) ; School of Computer Science, Beijing Institute of Technology(计算机科学学院,北京理工大学) ; School of Artificial Intelligence and Computer Science, Nantong University(人工智能与计算机科学学院,南通大学) ; Thrust of Intelligent Transportation, Hong Kong University of Science and Technology (GuangZhou)(智能交通研究组,香港科学与技术大学(广州)) ; Department of Electrical Engineering and Electronics, University of Liverpool(电子工程与电子学院,利物浦大学)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
Comments 27 pages, 12 figures
机构 * The University of Hong Kong(香港大学) ; Meta
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments 18 pages, 10 figures
Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments Accepted at ICML 2025
机构 * Center for Systems Science and Engineering(系统科学与工程中心) ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CL
Comments Last revised 13 Feb 2025. Under review in Nature portfolio
机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机学院,北京大学) ; Nanyang Technological University(南洋理工大学) ; WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * University of Waterloo(滑铁卢大学) ; University of Toronto(多伦多大学) ; Vector Institute(向量研究所) ; Shanghai University(上海大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments Dataset: https://huggingface.co/datasets/TIGER-Lab/VideoEval-Pro, Project Webpage: https://tiger-ai-lab.github.io/VideoEval-Pro
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai Jiaotong University(上海交通大学) ; Merchants Union Consumer Finance Company Limited
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * ETH Zürich(苏黎世联邦理工学院) ; Sony Semiconductor Solutions Europe, Sony Europe B.V.(索尼半导体欧洲解决方案,索尼欧洲有限公司)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments arXiv admin note: text overlap with arXiv:2504.10400
机构 * College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
机构 * National University of Singapore(新加坡国立大学) ; Southern University of Science and Technology(南方科技大学) ; University of Oxford(牛津大学)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
Comments Early accepted by MICCAI 2025
机构 * University of Macau(澳门大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中 视频多模态 :multimodal(abstract);分类 cs.AI
Journal ref IJCAI 2025
机构 * Artificial Intelligence Institute University of South Carolina(人工智能研究所南卡罗来纳大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.AI
Comments 8 pages, 8 figures, 4 tables, IEEE Conference on Artificial Intelligence (IEEE CAI) 2025
机构 * Department of Traffic Information and Control Engineering, Jilin University(交通信息与控制工程系,吉林大学) ; BIT-Barcelona Innovative Transportation Research Group, Civil Engineering School, UPC Barcelona Tech(巴塞罗那创新交通研究组,土木工程学院,UPC巴塞罗那技术学院) ; Computer Vision Center (CVC), Computer Science Department, Universitat Autònoma de Barcelona (UAB)(计算机视觉中心(CVC),计算机科学系,巴塞罗那自治大学(UAB))
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments IEEE SPL
机构 * University of California, Irvine(加州大学伊万斯分校) ; Pusan National University(釜山国立大学)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
机构 * Georgia Institute of Technology(佐治亚理工学院) ; SRI International(SRI国际)
专题命中 视频多模态 :multimodal(abstract);分类 cs.AI
机构 * Google DeepMind(谷歌DeepMind) ; Columbia University(哥伦比亚大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * The University of Hong Kong, Pok Fu Lam, Hong Kong(香港大学) ; Kuaishou Technology, Shenzhen, China(快手科技) ; The Hong Kong University of Science and Technology (HKUST), Hong Kong(香港科技大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Panasonic Connect Co., Ltd.(松下电器(中国)有限公司) ; Stanford University(斯坦福大学) ; Panasonic R&D Company of America(松下美国研发公司) ; Panasonic Holdings Corporation(松下控股公司)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * University of Stuttgart Germany(斯图加特大学) ; University of Tuebingen Germany(图宾根大学) ; The Center for Bionic Intelligence Tuebingen Stuttgart Germany(图宾根-斯图加特生物智能中心)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
Comments Accepted at SIGGRAPH 2025, link: https://zhiminghu.net/hu25_hoigaze.html
机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; University of Trento(特伦多大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Clinical Hospital of Chengdu Brain Science Institute(成都脑科学研究院临床医院) ; MOE Key Lab for NeuroInformation(教育部神经信息关键实验室) ; China-Cuba Belt and Road Joint Laboratory on Neurotechnology and Brain-Apparatus Communication(中 cuba 带路神经技术与脑-装置通信联合实验室) ; School of Life Science and Technology, University of Electronic Science and Technology of China(电子科技大学生命科学与技术学院) ; School of Artificial Intelligence, Chongqing University of Education(重庆教育学院人工智能学院) ; Sichuan Academy of Medical Sciences and Sichuan Provincial People’s Hospital(四川省医学科学院和四川省人民医院) ; Research Unit of NeuroInformation (2019RU035), Chinese Academy of Medical Sciences(神经信息研究单位(2019RU035),中国医学科学院)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
Comments 13 pages, 6 figures, 6 tables; This work has been submitted to Neural Networks for possible publication