LLMs Can Compensate for Deficiencies in Visual Representations
机构 * Fujitsu Limited(富士通有限公司) ; MBZUAI ; Tohoku University(东北大学) ; RIKEN(日本研究机构)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2025 Findings
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Fujitsu Limited(富士通有限公司) ; MBZUAI ; Tohoku University(东北大学) ; RIKEN(日本研究机构)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2025 Findings
机构 * POSTECH ; University of California, Berkeley(加州大学伯克利分校)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2025 Main Conference
机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(东南大学新一代人工智能技术及其交叉应用关键实验室) ; Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) ; Wuhan University of Technology(武汉理工大学) ; National University of Singapore(新加坡国立大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to Findings of EMNLP 2025
机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) ; School of Computer Science, Wuhan University(武汉大学计算机学院) ; School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) ; Cognitive AI Lab, Shanghai Huawei Technologies, China(上海华为技术有限公司认知人工智能实验室)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by EMNLP 2025
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; University of Washington(华盛顿大学) ; Shenzhen Institutes of Advanced Technology(深圳先进技术研究所) ; Chinese Academy of Sciences(中国科学院) ; State Key Laboratory of Ophthalmology(眼科学国家重点实验室) ; Zhongshan Ophthalmic Center(中山眼科中心) ; Sun Yat-sen University(中山大学) ; Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science(广东省眼科学与视觉科学重点实验室) ; Guangdong Provincial Clinical Research Center for Ocular Diseases(广东省眼科临床研究中心) ; Department of Bioengineering(生物工程系) ; Department of Radiology(放射科)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by IEEE TPAMI, 14 pages, 15 tables, 4 figures with Appendix
机构 * AITRICS Seoul(AITRICS首尔) ; Chung-Ang University(Chung-ang 大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments ICCV 2025 Workshop on Curated Data for Efficient Learning (CDEL)
机构 * Brown University(布朗大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments All code and resources are available at: https://github.com/ethahtz/vlm_conflicting_info_processing
机构 * College of Big Data and Internet, Shenzhen Technology University, China(大数据与互联网学院,深圳科技大学,中国) ; Sino-German College of Intelligent Manufacturing, Shenzhen Technology University, China(中德智能制造学院,深圳科技大学,中国) ; Robotics Department, Institute of Smart Systems and Artificial Intelligence (ISSAI), Nazarbayev University, Kazakhstan(机器人系,智能系统与人工智能研究所(ISSAI),纳扎尔巴耶夫大学,哈萨克斯坦) ; Foundation Robotics Labs, Germany(基础机器人实验室,德国)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract)
Comments This paper has been accepted by the 2025 International Conference on Climbing and Walking Robots (CLAWAR). These authors contributed equally to this work: Zexiang Guo, Hengxiang Chen, Xinheng Mai
机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) ; School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) ; School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算与信息系统学院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 23 pages, 10 figures
机构 * Purdue University(普渡大学)
专题命中 图文多模态 :multimodal(abstract);image-text(abstract)
机构 * Hewlett Packard Enterprise (Hewlett Packard Labs)(惠普企业公司(惠普实验室))
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted: IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2025
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 31 pages, 15 figures, 6 tables
Journal ref Computer Science Review, Vol. 58, pp. 100766, 2025
机构 * Microsoft Research(微软研究院) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室) ; Nanjing University(南京大学)
专题命中 图文多模态 :multimodal(abstract);MLLM(abstract)
机构 * Stanford University(斯坦福大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ACL 2025 Main. Project Website: https://s-vco.github.io/
机构 * Tencent AI Lab(腾讯AI实验室) ; Institute of Automation, Chinese Academy of Sciences(自动化研究所)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS
Comments Accepted by INTERSPEECH 2025
机构 * Artificial Intelligence and Intelligent Information Systems, University of Trier(人工智能与智能信息系统,特里尔大学) ; German Research Center for Artificial Intelligence (DFKI), Trier(德国人工智能研究中心(DFKI),特里尔)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
机构 * National Taiwan University(台湾大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ACL 2025 main
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Preprint
机构 * Robotics Program KAIST(韩国釜山科学技术院机器人计划) ; Electrical Engineering KAIST(韩国釜山科学技术院电子工程)
专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract)
Comments 5 pages, ICASSP 2025. The first two authors are equally contributed
Journal ref ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
专题命中 图文多模态 :multi-modal(abstract);image-text(abstract)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 18 pages, 10 figures, 6 tables, preprint. arXiv admin note: substantial text overlap with arXiv:2309.17002
机构 * TIB – Leibniz Information Centre for Science and Technology(蒂宾根-莱比锡信息科学与技术研究中心) ; Hochschule Hannover – Data(汉诺威高等学院-数据) ; H Institute for Applied Data Science(汉诺威应用数据科学研究所) ; L3S Research Center – Leibniz University Hannover(L3S研究中心-汉诺威莱比锡大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments 12 pages (excluding references), 8 tables, 1 equation
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments 22 pages, 11 figures, submitted as a preprint. ArXiv preprint only, not submitted to a journal yet
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments (v2) Clarified fine-tuning process, updated appendix
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted to CVPR 2025
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by WACV2025
Journal ref https://openaccess.thecvf.com/content/WACV2025/papers/Huang_Who_Brings_the_Frisbee_Probing_Hidden_Hallucination_Factors_in_Large_WACV_2025_paper.pdf
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI