SINA: A Fully Automated Circuit Schematic Image to Netlist Generator Using Artificial Intelligence
SINA: 一种使用人工智能的全自动电路原理图图像到网表生成器
Saoud Aldowaish, Yashwanth Karumanchi, Kai-Chen Chiang, Mohammed Ayman Habib, Finn Murphy, Rishen Cao, Morteza Fayazi
机构
*
The Department of Electrical and Computer Engineering, University of Utah(犹他大学电气与计算机工程系)
;
The Department of Electrical, Computer & Energy Engineering, University of Colorado Boulder(科罗拉多大学博尔德分校电气、计算机与能源工程系)
TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs
TABVERSE:大语言模型与视觉语言模型中跨格式表格理解的基准测试
Momina Ahsan, Sarfraz Ahmad, Ming Shan Hee, Roy Ka-Wei Lee, Preslav Nakov
机构
*
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)
机构
*
School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院)
;
Melbourne School of Psychological Sciences, The University of Melbourne(墨尔本大学墨尔本心理科学学院)
;
LILT
专题命中
文档图表理解
:vision-language model(abstract);vision language model(abstract);VLM(abstract_cn)
ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation
ClinOCR-Bench:用于光学字符识别模型评估的综合临床扫描文档数据集
Enshuo Hsu, Jin Zhou, Kirk Roberts
机构
*
McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston(德克萨斯大学健康科学中心休斯顿分校麦威廉斯生物医学信息学学院)
;
Enterprise Development & Integration, University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心企业开发与集成)
专题命中
文档图表理解
:vision language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.AI
MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework
MM-Matryoshka:通过二维多模态套娃训练框架实现预算弹性视觉文档检索
Haowen Xiang, Yibo Yan, Jiahao Huo, Yu Huang, Yi Cao, Mingdong Ou, Xuming Hu
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Alibaba Cloud Computing(阿里云计算)
;
Hong Kong University of Science and Technology(香港科技大学)
LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement
LightSTAR: 通过视觉自适应精炼的轻量级选择实现高效视觉文档检索
Tongkun Guan, Haocheng Wang, Wei Shen, Xiaokang Yang
机构
*
MoE Key Lab of Artificial Intelligence, AI Institute, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院人工智能研究院教育部人工智能重点实验室)
PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing
PereStruct: 面向鲁棒历史文档解析的多模态语义组装
Maksim Shandybo, Ivan Bespalov, Daniil Yefimov, Marina Kosheleva, Alexander Loukianov
机构
*
IGIC RAS(俄罗斯科学院信息传输问题研究所)
;
Yandex Cloud
;
National University of Science and Technology MISIS(莫斯科国立钢铁合金学院)
;
Nekrasov Central Universal Scientific Library(涅克拉索夫中央综合科学图书馆)
Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation
解耦语义与逻辑:一种无需训练的从粗到精的视频检索增强生成流水线
Jiaxin Dai, Zehang Wei, Jiamin Yan, Xiang Xiang
机构
*
School of Computer Science & Tech, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
;
School of AI and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)