MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
MechVQA:在综合机械图纸理解上基准测试与增强多模态大语言模型
机构 * Beijing Academy of Artificial Intelligence (BAAI), China(北京人工智能研究院) ; Institute of Information Engineering, Chinese Academy of Sciences, China(信息工程研究所) ; Beijing University of Technology, China(北京理工大学)
专题命中 视觉问答 :MLLM(abstract,abstract_cn);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 针对多模态大语言模型在机械工程图纸理解上的不足,提出首个综合机械图纸理解数据集MechVQA,并开发MechVL模型,通过多阶段训练显著提升性能。
Comments accept by iclm2026, add github link
Journal ref Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), Seoul, South Korea, PMLR 306 (2026)