Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
机构 * Florida State University(佛罗里达州立大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 12 pages, 5 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Florida State University(佛罗里达州立大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 12 pages, 5 figures
机构 * CASIA(中国科学院自动化研究所) ; UCAS(中国科学技术大学) ; ZGCA(北京智感科技有限公司) ; HKU(香港大学) ; HKUST(香港科技大学) ; NTU(国立台湾大学) ; PKU(北京大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Preprint, Under review
机构 * VUNO Inc.(VUNO公司) ; KAIST(韩国科学技术院)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI
Comments 38 pages, 17 figures, preprint
机构 * Korea University(韩国大学) ; University of Seoul(首尔大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 23 pages, 17 figures
机构 * Fudan University(复旦大学) ; University of Southern California(南加州大学) ; ByteDance(字节跳动)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL
机构 * University of California, San Diego(加州大学圣地亚哥分校) ; ByteDance(字节跳动) ; The University of Queensland(昆士兰大学) ; University of Southern California(南加州大学) ; University at Buffalo(布法罗大学) ; University of California, Merced(加州大学默塞德分校)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments Accepted to COLM 2025
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Qianfan Team, Baidu AI Cloud(百度AI云团队)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 12 pages
机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 19 pages, 12 figures, 6 tables
机构 * Santa Clara University(圣克拉拉大学) ; DOCOMO Innovations, Inc.(DOCOMO创新公司) ; Rochester Institute of Technology(罗切斯特理工学院)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL
Comments EMNLP Findings
机构 * Embia, Computer Science Department, Faculty of Science, Université de Moncton(Embia计算机科学系,科学学院,蒙特龙大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 9 pages, 3 figures, IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL
Comments EMNLP 2025
机构 * NC AI
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments 19 pages, 1 figure, 14 tables. Technical report for VARCO-VISION-2.0, a Korean-English bilingual VLM in 14B and 1.7B variants. Key features: multi-image understanding, OCR with text localization, improved Korean capabilities
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL
机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理实验室,计算机科学学院,复旦大学) ; The Chinese University of Hong Kong, Shatin, Hong Kong(香港中文大学,沙田,香港) ; The University of Hong Kong, Pokfulam, Hong Kong(香港大学,薄扶林,香港)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments This work has been submitted to the IEEE for possible publication
机构 * DFKI Niedersachsen(德克萨斯联合研究所(北莱茵威斯特法伦)) ; Cooperative and Autonomous Systems, DFKI Niedersachsen(合作与自主系统,DFKI北莱茵威斯特法伦) ; German Research Center for Artificial Intelligence(德国人工智能研究中心) ; Semantic Information Systems, Osnabrück University(语义信息系统,奥斯纳布吕克大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted at 24th International Conference on Machine Learning and Applications (ICMLA'25)
机构 * Shanghai Ocean University(上海海洋大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments 72 Pages, 11 Figures
机构 * School of Computer Science, Carnegie Mellon University(计算机科学系,卡内基梅隆大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted by ICLR 2025. Project page: https://zhangce01.github.io/DeGF/
机构 * University of Amsterdam(阿姆斯特丹大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI
机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet), Norway(Simula数字工程中心(SimulaMet)) ; Oslo Metropolitan University (OsloMet), Norway(奥斯陆 Metropolitan 大学(OsloMet)) ; Simula Research Laboratory, Norway(Simula研究实验室)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted as a full paper at the 38th IEEE International Symposium on Computer-Based Medical Systems (CBMS) 2025
机构 * University of Manchester(曼彻斯特大学) ; Microsoft Research(微软研究院)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted to EMNLP 2025 Main Conference
机构 * Institute for Systems and Robotics(系统与机器人研究所) ; University of Lisbon(里斯本大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) ; Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) ; Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) ; Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference
机构 * Department of Computer Science University of California, Irvine(计算机科学系加州大学伊文斯顿分校)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, 200032, China(复旦大学基础医学学院数字医学研究中心) ; Shanghai Key Laboratory of Medical Imaging Computing and Computer Assisted Intervention, Shanghai, 200032, China(上海市医疗影像计算与计算机辅助干预重点实验室) ; Department of Urology, Qilu Hospital of Shandong University, Jinan, Shandong, 250012, China(山东大学齐鲁医院泌尿科) ; Microsoft Research Asia, Shanghai, 200232, China(微软亚洲研究院) ; Department of Radiology, The First Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, 310006, China(浙江大学医学院附属第一医院放射科) ; Center of Health data science, Linyi People’s Hospital, Shandong, 276003, China(临沂人民医院健康数据科学中心) ; Shandong Open Laboratory of Data Innovation Application, Shandong, 276003, China(山东省数据创新应用开放实验室) ; Department of Radiology, the First People’s Hospital of Lianyungang, Lianyungang, 222002, China(连云港第一人民医院放射科) ; Department of Urology, Zhangye People’s Hospital affiliated to Hexi University, Zhangye, 734000, China(张掖人民医院(河西大学附属)泌尿科) ; Department of Urology, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, 200025, China(上海交通大学附属瑞金医院泌尿科) ; Department of Urology, Linyi People’s Hospital, Shandong, 276003, China(临沂人民医院泌尿科) ; Department of Urology, Zhongshan Hospital, Fudan University, Shanghai, 200032, China(复旦大学中山医院泌尿科)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI
机构 * Shanghai University(上海大学) ; Beijing University of Chemical Technology(北京化工大学) ; University of Minnesota(明尼苏达大学)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted by ICCV 2025
机构 * TUDelft(代尔夫特理工大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL
Comments 16 pages (including references), 5 figures and 6 tables
机构 * Independent Researcher(独立研究者)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM
机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) ; National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) ; Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) ; Xi’an Jiaotong University(西安交通大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Osaka University(大阪大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted by IEEE Transactions on Information Forensics & Security