F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
专题命中 视觉推理 :vision language model(title);vision-language model(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :vision language model(title);vision-language model(abstract);分类 cs.CV
机构 * Colorado School of Mines(科罗拉多矿业学院)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 8 pages, 8 figures
机构 * University of Electronic Science and Technology of China(电子科学与技术大学) ; Southwestern University of Finance and Economics(西南财经大学) ; Tongji University(同济大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments 19 pages, 11 figures. Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)
机构 * School of Interdisciplinary Science, Beijing Institute of Technology(交叉科学学院,北京理工大学) ; College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) ; National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing, Beijing Institute of Technology(空间智能信息处理国家重点实验室,北京理工大学) ; School of Optics and Photonics, Beijing Institute of Technology(光学与 photonics 学院,北京理工大学) ; State Key Laboratory of Explosion Science and Safety Protection, Beijing(爆炸科学与安全防护国家重点实验室,北京)
专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV
机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学) ; Tongyi Fun Team, Alibaba Group(阿里云团队) ; Zhejiang University(浙江大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025 Main
机构 * Tredence Inc.(特伦德公司) ; Indian Institute of Technology Kharagpur(印度理工学院克拉格浦尔分校) ; Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所) ; Rajiv Gandhi University(拉吉夫·甘地大学) ; Tsinghua University(清华大学) ; Palmer Research Laboratories(帕勒姆研究实验室)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 7 pages, 5 figures, 4 tables
机构 * Sun Yat-sen University(中山大学) ; Peng Cheng Laboratory(鹏城实验室) ; Jinan University(暨南大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.LG
Comments NeurIPS 2025 Datasets and Benchmarks Track poster
机构 * National University of Singapore(新加坡国立大学) ; Sea AI Lab(Sea AI实验室)
专题命中 视觉推理 :visual reasoning(title);vision-language model(abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * NVIDIA ; NIH/NCI ; CHOP/UPenn ; Basaksehir Cam and Sakura City Hospital
专题命中 视觉推理 :visual language model(title);vision-language model(abstract);分类 cs.CV
Comments NV-Reason-CXR-3B
机构 * Rochester Institute of Technology(罗切斯特技术研究所) ; Snap Inc.(Snap公司) ; University of Rochester(罗切斯特大学) ; DEVCOM Army Research Laboratory(陆军研究实验室)
专题命中 视觉推理 :visual reasoning(title);vision-language model(abstract);分类 cs.AI
Comments NeurIPS 2025
机构 * Tsinghua University(清华大学) ; Pengcheng Laboratory(鹏城实验室) ; Sun Yat-sen University(中山大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
机构 * School of Computing, Macquarie University, Australia(计算机学院,麦考瑞大学,澳大利亚)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
机构 * Department of Computer Science and Informatics, Emory University(计算机科学与信息学系,埃默里大学) ; Department of Computer Science and Department of Electrical and Computer Engineering, University of Southern California(计算机科学系和电气与计算机工程系,南加州大学) ; Department of Computer Science, University of Tokyo(计算机科学系,东京大学) ; Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学) ; Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(生物医学工程系,佐治亚理工学院和埃默里大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.LG
Comments EMNLP 2025. Project Website: https://vehme.github.io/
机构 * Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室) ; Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
Comments Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables
机构 * City University of Hong Kong(香港城市大学) ; Tencent(腾讯) ; Zhejiang University(浙江大学)
专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) ; State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) ; Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
Journal ref Neurocomputing, Volume 659, 2026, 131217
机构 * The University of Melbourne(墨尔本大学) ; China University of Petroleum (East China)(中国石油大学(华东))
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments Accepted by ACM Multimedia 2025
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments Work in progress
机构 * Computer and Information Sciences Department, Cornell University(康奈尔大学计算机与信息科学系) ; Waymo LLC(Waymo公司) ; UC Berkeley(伯克利大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments In proceedings of IROS 2025
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) ; Xiamen University(厦门大学) ; The Hong Kong University of Science and Technology(香港理工大学) ; Nanyang Technological University(南洋理工大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
机构 * Turing Inc.(图灵公司)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 2nd Place Winner, ICCV 2025 2COOOL Competition
机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
Comments EMNLP 2025 Wordplay (Spotlight)
机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT
机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) ; Tencent AI Seattle Lab(腾讯AI西雅图实验室) ; University of Edinburgh(爱丁堡大学) ; NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted as a spotlight at NeurIPS 2025
机构 * McGill University(麦吉尔大学) ; Massachusetts Institute of Technology(麻省理工学院) ; Shanghai Jiao Tong University(上海交通大学) ; The Hong Kong Polytechnic University(香港理工大学) ; University of Florida(佛罗里达大学) ; The University of Hong Kong(香港大学)
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.CV
机构 * Xiaomi Inc.(小米公司)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
机构 * Department of Computer Science, University of Applied Sciences and Arts Dortmund(应用科学与艺术大学多特蒙德计算机科学系) ; Institute for Medical Informatics, Biometry and Epidemiology (IMIBE)(医学信息学、生物统计与流行病学研究所) ; Institute for Artificial Intelligence in Medicine (IKIM), University Hospital Essen(医学人工智能研究所,埃森大学医院)
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI
Comments Accepted at ICDAR 2025
机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动学院) ; McGill University(麦吉尔大学) ; Automotive and Robotics, Xiaomi Corporation(小米公司汽车与机器人部门) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 19 pages, 8 figures
Journal ref EMNLP2025 Fundings