A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
基于MLLM的视觉丰富文档理解综述:方法、挑战与新兴趋势
机构 * The University of Western Australia(西澳大学) ; The University of Melbourne(墨尔本大学) ; Weill Cornell Medicine(韦尔·柯尔医学中心)
专题命中 多模态训练与对齐 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本文综述了基于MLLM的视觉丰富文档理解最新进展,探讨了文本、视觉和布局特征的表示与整合技术,以及预训练、指令微调等训练方法,分析了数据稀缺、多页文档处理等挑战及新兴趋势。
Comments Accepted at ACL 2026 Findings