Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
Doc-V*:多页文档视觉问答中的粗到细交互推理
机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) ; MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus团队) ; School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) ; School of Data Science, Fudan University(复旦大学数据科学学院)
AI总结 Doc-V*提出一种无OCR的智能框架,通过序列证据聚合解决多页文档VQA问题,结合语义检索和目标页面获取,提升回答准确性和证据收集效率,在五个基准测试中优于开源基线,提升领域外性能达47.9%。
Comments Accepted by Association for Computational Linguistics (ACL)