ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
Jovana Kondic, Pengyuan Li, Dhiraj Joshi, Zexue He, Shafiq Abedin, Jennifer Sun, Ben Wiesel, Eli Schwartz, Ahmed Nassar, Bo Wu, Assaf Arbelle, Aude Oliva, Dan Gutfreund, Leonid Karlinsky, Rogerio Feris
机构
*
MIT(麻省理工学院)
;
MIT-IBM Watson AI Labs(麻省理工-IBM沃森人工智能实验室)
;
IBM Research(IBM研究院)
机构
*
School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)
;
Labs, Huawei Technologies Co., LTD(华为技术有限公司2012实验室)
;
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)
专题命中
文档图表理解
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
Large Language Model Department, Tencent(腾讯大语言模型部)
;
Nankai University(南开大学)
机构
*
Johns Hopkins University(约翰霍普金斯大学)
;
Zhejiang University(浙江大学)
;
Independant Researcher(独立研究者)
;
Tsinghua University(清华大学)
;
University of Central Florida(佛罗里达中央大学)
;
City University of Hong Kong(香港城市大学)
;
Adobe Research(Adobe研究)
LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement
LightSTAR: 通过视觉自适应精炼的轻量级选择实现高效视觉文档检索
Tongkun Guan, Haocheng Wang, Wei Shen, Xiaokang Yang
机构
*
MoE Key Lab of Artificial Intelligence, AI Institute, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院人工智能研究院教育部人工智能重点实验室)
PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing
PereStruct: 面向鲁棒历史文档解析的多模态语义组装
Maksim Shandybo, Ivan Bespalov, Daniil Yefimov, Marina Kosheleva, Alexander Loukianov
机构
*
IGIC RAS(俄罗斯科学院信息传输问题研究所)
;
Yandex Cloud
;
National University of Science and Technology MISIS(莫斯科国立钢铁合金学院)
;
Nekrasov Central Universal Scientific Library(涅克拉索夫中央综合科学图书馆)