Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R
双路Top-K检索与1v1 VLM重排序用于CoVR-R
Yuyang Sun, Yongliang Wu, Xingyu Zhu, Yuxia Chen, Zhenxiang Jiang, Yangguang Ji, Wenbo Zhu, Yanxi Shi, Jay Wu, Shuo Wang, Xu Yang
机构
*
Southeast University(东南大学)
;
National University of Singapore(新加坡国立大学)
;
Independent Researcher(独立研究者)
;
Opus AI Research(Opus AI研究)
;
University of Science and Technology of China(中国科学技术大学)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
LookWise: 知道何时何地关注多模态大语言模型中的细粒度视觉推理
Yuxiang Shen, Hailong Huang, Zhenkun Gao, Xueheng Li, Man Zhou, Chengjun Xie, Haoxuan Che, Xuanhua He, Jie Zhang
机构
*
Institute of Intelligent Machines, Hefei Institutes of Physical Science, Chinese Academy of Sciences(智能机器研究所,合肥物理科学研究院,中国科学院)
;
University of Science and Technology of China(中国科学技术大学)
;
Zhejiang University(浙江大学)
;
East China Normal University(华东师范大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
专题命中
视觉推理
:visual reasoning(title,abstract);multimodal large language model(title,abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI
机构
*
Tandon School of Engineering, New York University(纽约大学工程学院)
;
Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)
;
Brookhaven National Laboratory(布鲁克海文国家实验室)
Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy
超越文本与表格:ComProScanner中视觉-语言模型集成实现从科学图表中高精度提取材料数据
Aritra Roy, Enrico Grisan, Chiara Gattinoni, John Buckeridge
机构
*
Energy, Materials and Environment Research Centre, London South Bank University, London SE1 0AA, UK(能源、材料与环境研究中心,伦敦南银行大学)
;
School of Engineering and Design, London South Bank University, London SE1 0AA, UK(工程与设计学院,伦敦南银行大学)
;
Bioscience and Bioengineering Research Centre, London South Bank University, London SE1 0AA, UK(生物科学与生物工程研究中心,伦敦南银行大学)
;
Department of Physics, Kings College London, London WC2R 2LS, UK(物理系,伦敦国王学院)
机构
*
Kyoto University(京都大学)
;
NII LLMC(日本国立信息与通信技术研究所语言模型中心)
;
RIKEN AIP(日本理化学研究所先进理工研究所)
;
Case Western Reserve University(凯斯西储大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Osaka(大阪大学)
;
University of Tokyo(东京大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV
Comments25 pages. v2: Title updated; added a section on object/spatial imagery and propositional reasoning; added new experimental results for the single-object rotation probe