Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
机构 * School of Computer Science University of Bristol(计算机科学学院英国布里斯托尔大学)
专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV
Comments Accepted at BMVC 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * School of Computer Science University of Bristol(计算机科学学院英国布里斯托尔大学)
专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV
Comments Accepted at BMVC 2025
机构 * Departments of Mechanical and Electrical Engineering, Southern Methodist University(机械与电气工程系,南方 Methodist 大学)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
Comments Submitted to IEEE
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 23 pages, 5 figures
机构 * Electronic Information School, Wuhan University, Wuhan 430072, China(武汉大学电子信息学院)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted by ICCV 2025
机构 * The University of Hong Kong(香港大学) ; The Chinese University of Hong Kong(香港中文大学) ; Google(谷歌)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments Accepted to UIST 2025. Project website: https://zhaorunning.github.io/NoteIt/
机构 * Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, China(中国科学院智能信息处理重点实验室(中国科学院)、计算技术研究所、中国)
专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV
机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) ; School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) ; School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) ; Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 20 pages,6 figures,survey
机构 * Amazon(亚马逊) ; School of Psychology and Neuroscience, University of Glasgow(心理学与神经科学学院,格拉斯哥大学) ; Ben-Gurion University of the Negev(内盖夫本·古里安大学) ; University of Cambridge(剑桥大学) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) ; Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学教育部长江网络与网络安全重点实验室) ; College of Computer and Mathematics, Central South University of Forestry and Technology(中南林业科技大学计算机与数学学院) ; Department of Computer Science, State University of New York(纽约州立大学新帕尔茨分校计算机科学系)
专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL
Comments 13 pages, 6 figures
机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
机构 * Nanjing University(南京大学) ; Fudan University(复旦大学) ; S-Lab, Nanyang Technological University(南洋理工大学S实验室) ; NVIDIA(NVIDIA公司) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV
Comments Project page: https://vchitect.github.io/LongVie-project/
机构 * School of Computer Science and Engineering, MoE Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-sen University(计算机科学与工程学院、信息科技教育部重点实验室、广东省信息安全技术重点实验室、中山大学) ; State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 13 pages,4 figures. arXiv admin note: text overlap with arXiv:2507.16596
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM
机构 * Beijing University of Technology(北京理工大学) ; Chinese Academy of Sciences(中国科学院) ; Capital Medical University(首都医科大学)
专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV
机构 * University of Central Florida(佛罗里达中央大学) ; Case Western Reserve University(凯斯西储大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL
Comments Accepted to the 24th International Semantic Web Conference Resource Track (ISWC 2025)
机构 * Beijing Jiaotong University(北京交通大学) ; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) ; Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)
专题命中 视频多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV
机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) ; Shituoyun (Nanjing) Technology Co., Ltd(石图云(南京)科技有限公司) ; School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI
Comments 12 pages, 4 figures, submitted to IEEE Transactions on Knowledge and Data Engineering
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
Comments 10 pages
机构 * Beijing University of Technology(北京理工大学) ; Shanghai Jiao Tong University(上海交通大学) ; Chinese Academy of Sciences(中国科学院) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Accepted by ICCV 2025 (Poster)
机构 * Data Science Institute University of Chicago(芝加哥大学数据科学研究所) ; Department of Psychology, Neuroscience Institute University of Chicago(芝加哥大学心理学系、神经科学研究所)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
机构 * Intel(英特尔)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Accepted to CVAM Workshop at ICCV 2025
机构 * Department of Computer Science & Engineering University of Minnesota(计算机科学与工程系明尼苏达大学) ; Department of Mechanical Engineering University of Minnesota(机械工程系明尼苏达大学) ; Department of Neurosurgery University of Minnesota(神经外科系明尼苏达大学)
专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV
机构 * University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校) ; Tietoevry(蒂奥维瑞) ; Tampere University(塔尔库大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments Published at Reinforcement Learning and Video Games workshop https://sites.google.com/view/rlvg-workshop-2025/home
机构 * Intelligent Editing Team(智能编辑团队) ; Intelligent Creation, ByteDance Inc.(智能创作,字节跳动公司)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
机构 * Vrije Universiteit Amsterdam(范·艾克大学阿姆斯特丹)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM
机构 * Department of Computer Science, Chongqing University, China(重庆大学计算机科学系) ; National Elite Institute of Engineering, Chongqing University, China(重庆大学工程精英研究院) ; College of Computer Science and Technology, National University of Deffense Technology, China(国防科技大学计算机科学与技术学院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Journal ref ICCV 2025
机构 * TeleAI, China Telecom(TeleAI,中国电信) ; Computer Vision Lab, CAIDAS & IFI, University of Wurzburg(计算机视觉实验室,CAIDAS与IFI,乌尔姆大学) ; INSAIT, Sofia University(INSAIT,索菲亚大学) ; ShanghaiTech University(上海科技大学) ; University of Copenhagen(哥本哈根大学) ; AI Institute, Shanghai Jiao Tong University(人工智能研究院,上海交通大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments ICCV2025 accepted