Do large language vision models understand 3D shapes?
专题命中 其他VLM :vision language model(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 其他VLM :vision language model(abstract);分类 cs.CV
机构 * Harvard AI and Robotics Lab, Harvard University(哈佛人工智能与机器人实验室,哈佛大学) ; Broad Institute(博德研究所) ; Hong Kong Polytechnic University(香港理工大学) ; School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) ; City University of Hong Kong(香港城市大学) ; Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(自然与人工智能研究学院,哈佛大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments 24 pages, 14 figures, 16 tables
Journal ref ICCV 2025
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Alibaba Group(阿里巴巴集团) ; Peking University(北京大学) ; Luoyang Institute for Robot and Intelligent Equipment(洛阳机器人与智能装备研究所)
专题命中 其他VLM :MLLM(abstract);分类 cs.CV
Comments Accepted by ICCV 2025
机构 * Chair of Robotics, Artificial Intelligence and Real-Time Systems(机器人、人工智能与实时系统教授席)
专题命中 其他VLM :vision language model(abstract);分类 cs.AI
Comments Conference paper accepted for GACLM 2025
机构 * Computer Engineering, Texas A\&M University, College Station, TX 77843, USA ; College of Science ; Engineering, Hamad Bin Khalifa University, Doha, Qatar ; Dept. of Electrical \& Computer Engineering, Texas A\&M University at Qatar, Doha, Qatar ; Orthotics, Ankara University, Ankara, Turkey
专题命中 其他VLM :vision-language model(abstract);分类 cs.LG
Comments Presented at ICML 2025 Workshop on DataWorld
机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) ; China University of Geosciences(中国地质大学) ; Nanyang Technological University(南洋理工大学) ; BAAI(百度人工智能研究院) ; Stanford University(斯坦福大学)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
机构 * Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong, China(人工智能与机器人研究中心,香港科学与创新研究所,中国科学院,香港,中国) ; School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科技大学) ; Surrey Institute for People-Centred Artificial Intelligence, CVSSP, University of Surrey(以人为中心的人工智能研究所,CVSSP, Surrey大学) ; School of Electronics and Information Engineering, Shenzhen University(电子与信息工程学院,深圳大学) ; University of Rochester(罗切斯特大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.LG
Comments Under Review
机构 * Core Contributors(核心贡献者) ; Tiger-AI-Lab(虎鲸人工智能实验室)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments ICLR 2025 camera-ready version. Project page: https://tiger-ai-lab.github.io/MEGA-Bench/
机构 * Texas A&M University(德克萨斯大学) ; Stanford University(斯坦福大学) ; Snap Inc.(Snap公司) ; CU Boulder(科罗拉多大学博尔德分校) ; UT Austin(得克萨斯大学奥斯汀分校) ; California Institute of Technology(加州理工学院) ; Topaz Labs(Topaz实验室) ; UC Merced(加州大学默塞德分校)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Project page: https://4kagent.github.io
机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) ; Sun Yat-sen University(孙中山大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted to CVPR 2025
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
机构 * Tsinghua University(清华大学) ; Tencent Hunyuan X(腾讯混元实验室)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
Comments Accepted to ICCV 2025
机构 * University of Science(科学大学) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; CRIPAC, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所CRIPAC) ; VIS, Baidu Inc.(百度公司VIS)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Code is available at https://github.com/Yueming6568/DeltaEdit
专题命中 其他VLM :vision-language model(abstract);分类 cs.AI
Comments 14 pages
机构 * Wuhan University(武汉大学)
专题命中 其他VLM :MLLM(abstract);分类 cs.AI
机构 * Zhejiang Gongshang University(浙江工商大学) ; Zhejiang Yuexiu University(浙江越秀大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted to ICCV 2025
机构 * Department of Information Engineering, Harbin Institute of Technology(信息工程系,哈尔滨工业大学) ; National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences(空间信息集成国家重点实验室,中国科学院软件研究所) ; Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) ; Hubei Luojia Laboratory, National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(湖北珞珈实验室,国家多媒体软件工程技术研究中心,武汉大学计算机学院)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments 17 pages, 18 figures, Accepted by IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
机构 * Washington University in St. Louis(圣路易斯华盛顿大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted at ICCV 2025
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中 其他VLM :vision-language model(abstract);分类 cs.AI
机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) ; Laboratory of Art and Archaeology Image(艺术与考古图像实验室)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
机构 * Media and Data Science Research Lab, Adobe(Adobe媒体与数据科学研究实验室)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
Comments Accepted in CVPR 2025 Workshop on What's Next in Multimodal Foundational Models
机构 * ÉTS Montréal(ÉTS蒙特利尔) ; Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM)(蒙特利尔大学中心医院研究中心(CRCHUM))
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments MICCAI 2025. Code: https://github.com/jusiro/SS-Text
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments CVPR 2025
机构 * Kempelen Institute of Intelligent Technologies(智能技术研究所) ; Language Technology, University of Helsinki(语言技术,赫尔辛基大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.LG
机构 * Signals and Interactive Systems Lab, University of Trento, Italy(信号与交互系统实验室,特伦托大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
机构 * Institute of Mathematics and Statistics, University of São Paulo(数学与统计学研究所,圣保罗大学) ; Department of Information and Computing Sciences, Utrecht University(信息与计算科学系,乌得勒支大学)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.LG
机构 * Tsinghua University(清华大学)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
机构 * School of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感与信息工程学院) ; School of Data Science, Lingnan University(岭南大学数据科学学院)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV
机构 * Institut de Robòtica i Informàtica Industrial, CSIC-UPC(工业机器人与计算机研究所,CSIC-UPC)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted at ICRA 2025. Project page at https://barbany.github.io/bifold/