SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
SCoPE VLM:面向视觉语言模型高效文档导航的 selective context processing
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Konkuk University(韩国康克伦大学)
专题命中 GUI与网页智能体 :agentic(abstract);分类 cs.CL
AI总结 SCoPE VLM通过引入滚动链机制和定制强化学习方法,实现高效文档导航,提升视觉语言模型在多页文档问答中的代理阅读能力。