Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
Deep-Reporter:基于多模态长形式生成的深度研究
机构 * National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学) ; University of Edinburgh(爱丁堡大学) ; Beijing Institute of Technology(北京理工大学)
专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI
AI总结 本文提出Deep-Reporter框架,通过多模态搜索、检查清单引导合成和上下文管理,解决多模态长形式生成中的整合与选择难题,通过8K高质量数据集和M2LongBench测试集验证其有效性。
Comments 41 pages, 6 figures, 8 tables. Code available at https://github.com/fangda-ye/Deep-Report. v2: corrected typos and updated experimental results