OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
OmniParser V2:结构化思维点用于统一的视觉文本解析及其在多模态大语言模型中的通用性
Wenwen Yu, Zhibo Yang, Jianqiang Wan, Sibo Song, Jun Tang, Wenqing Cheng, Yuliang Liu, Xiang Bai
机构
*
School of Information Science and Engineering, East China University of Science and Technology(东华大学信息科学与工程学院)
;
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
;
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)
;
Alibaba Group(阿里巴巴集团)
专题命中
知识编辑与模型理解
:large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL
Comments8 real pages + AI content; proceedings of the workshop "Recent Progress in Computational String Geometry,'' held at the Chennai Mathematical Institute in India in January 2026; comments welcome
GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations
GAIR:具有地理对齐隐式表示的定位感知自监督对比预训练
Zeping Liu, Ni Lao, Zhangyu Wang, Junfeng Jiao, Gengchen Mai
机构
*
SEAI Lab, Department of Geography and the Environment, The University of Texas at Austin(地理与环境系SEAI实验室,德克萨斯大学奥斯汀分校)
;
Google LLC, Mountain View, CA, USA(谷歌公司,山景城,加利福尼亚州,美国)
;
SIT Lab, School of Computing and Information Science, The University of Maine(计算与信息科学系SIT实验室,缅因大学)
;
Urban Information Lab, School of Architecture, The University of Texas at Austin(城市信息实验室,建筑系,德克萨斯大学奥斯汀分校)