VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.LG
Comments EMNLP 2025. Project Website: https://vehme.github.io/
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.LG
Comments EMNLP 2025. Project Website: https://vehme.github.io/
机构 * Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室) ; Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
Comments Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables
机构 * City University of Hong Kong(香港城市大学) ; Tencent(腾讯) ; Zhejiang University(浙江大学)
专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) ; State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) ; Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
Journal ref Neurocomputing, Volume 659, 2026, 131217
机构 * The University of Melbourne(墨尔本大学) ; China University of Petroleum (East China)(中国石油大学(华东))
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments Accepted by ACM Multimedia 2025
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments Work in progress
机构 * Computer and Information Sciences Department, Cornell University(康奈尔大学计算机与信息科学系) ; Waymo LLC(Waymo公司) ; UC Berkeley(伯克利大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments In proceedings of IROS 2025
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) ; Xiamen University(厦门大学) ; The Hong Kong University of Science and Technology(香港理工大学) ; Nanyang Technological University(南洋理工大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
机构 * Turing Inc.(图灵公司)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 2nd Place Winner, ICCV 2025 2COOOL Competition
机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
Comments EMNLP 2025 Wordplay (Spotlight)
机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT
机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) ; Tencent AI Seattle Lab(腾讯AI西雅图实验室) ; University of Edinburgh(爱丁堡大学) ; NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted as a spotlight at NeurIPS 2025
机构 * McGill University(麦吉尔大学) ; Massachusetts Institute of Technology(麻省理工学院) ; Shanghai Jiao Tong University(上海交通大学) ; The Hong Kong Polytechnic University(香港理工大学) ; University of Florida(佛罗里达大学) ; The University of Hong Kong(香港大学)
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.CV
机构 * Xiaomi Inc.(小米公司)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
机构 * Department of Computer Science, University of Applied Sciences and Arts Dortmund(应用科学与艺术大学多特蒙德计算机科学系) ; Institute for Medical Informatics, Biometry and Epidemiology (IMIBE)(医学信息学、生物统计与流行病学研究所) ; Institute for Artificial Intelligence in Medicine (IKIM), University Hospital Essen(医学人工智能研究所,埃森大学医院)
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI
Comments Accepted at ICDAR 2025
机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动学院) ; McGill University(麦吉尔大学) ; Automotive and Robotics, Xiaomi Corporation(小米公司汽车与机器人部门) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 19 pages, 8 figures
Journal ref EMNLP2025 Fundings
机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) ; National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV
Comments Work in progress
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
机构 * The Pennsylvania State University(宾夕法尼亚州立大学) ; Pacific Northwest National Laboratory(太平洋西北国家实验室)
专题命中 视觉推理 :VLM(title);vision language model(abstract);分类 cs.CV
机构 * Johns Hopkins University(约翰霍普金斯大学) ; StepFun ; BUPT(北京邮电大学) ; UCAS(中国科学院大学) ; THU(清华大学) ; HUST(华中科技大学)
专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * MIT(麻省理工学院) ; NVIDIA(英伟达)
专题命中 视觉推理 :vision language model(title);vision-language model(abstract);分类 cs.CV
Comments Project Website: https://www.anjiecheng.me/sr3d
机构 * Xiaomi Corporation Beijing, China(小米公司北京)
专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV
机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG
Comments Accepted at IEEE GLOBECOM 2025
机构 * Department of Computer Science, University of Southern California(计算机科学系,南加州大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
Comments Conference on Robot Learning (CoRL) 2025. 50 pages and 30 figures. v2 is the camera-ready and includes a few more new experiments compared to v1
机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) ; Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) ; Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) ; Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) ; Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)
专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted to AAAI Fall Symposium 2025 on AI Trustworthiness and Risk Assessment for Challenging Contexts (ATRACC)
机构 * Shandong University, School of Control Science ; Rutgers University, Department of Computer Science 110 Frelinghuysen Road Piscataway New Brunswick NJ USA 08854 ; University of California, Santa Barbara 1210 Cheadle Hall Santa Barbara California USA 93106 ; University of Michigan, School of Information 500 S State St Ann Arbor Michigan USA 48109 ; The Chinese University of Hong Kong, Department of Computer Science ; The College of William \& Mary, School of Computing, Data Sciences \& Physics 200 Stadium Dr Williamsburg Virginia USA 23185 ; The Chinese University of Hong Kong, Department of Psychiatry 9 Chuen On Rd Tai Po New Territories Hong Kong SAR, China 999077 ; Rutgers University, Department of Computer Science ; University of California, Santa Barbara ; University of Michigan, School of Information ; The College of William \& Mary, School of Computing, Data Sciences \& Physics ; The Chinese University of Hong Kong, Department of Psychiatry
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 25 pages, 9 figures, 2 tables