Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
解锁多模态文档智能:从当前成就到视觉文档检索的未来前沿
机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Alibaba Cloud Computing(阿里云计算) ; Hong Kong University of Science and Technology(香港科技大学) ; University of Illinois Chicago(伊利诺伊大学芝加哥分校)
专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract)
AI总结 本文综述了视觉文档检索领域,探讨了多模态大语言模型时代下的方法演进与挑战,提出未来发展方向。
Comments Under review. This version updates the relevant works released before 15 March, 2026