OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models
OddGridBench: 暴露多模态大语言模型在细粒度视觉差异敏感性方面的不足
Tengjin Weng, Wenhao Jiang, Jingyi Wang, Ming Li, Lin Ma, Zhong Ming
机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机与软件学院)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
;
Shenzhen Technology University(深圳技术大学)
;
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
Meituan(美团)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV
Explaining CLIP Zero-shot Predictions Through Concepts
通过概念解释CLIP零样本预测
Onat Ozdemir, Anders Christensen, Stephan Alaniz, Zeynep Akata, Emre Akbas
机构
*
School of Informatics, University of Edinburgh(爱丁堡大学信息学院)
;
Dept. of Computer Eng., Middle East Technical University (METU)(中东技术大学计算机工程系)
;
Orbital
;
DTU Compute, Technical University of Denmark(丹麦技术大学DTU计算学院)
;
Dept. of Biology, University of Copenhagen(哥本哈根大学生物系)
;
LTCI, Télécom Paris, Institut Polytechnique de Paris(巴黎理工学院巴黎电信学院LTCI实验室)
;
Technical University of Munich (TUM)(慕尼黑工业大学)
;
Helmholtz Munich(亥姆霍兹慕尼黑中心)
;
MCML
;
MDSI
;
Robotics & AI Center (ROMER), METU(中东技术大学机器人与人工智能中心)
机构
*
School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)
;
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Wuhan Vocational College of Software and Engineering(武汉软件工程职业学院)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)
;
Hangzhou Innovation Institute, Beihang University(北京航空航天大学杭州创新研究院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
GUIDED: 通过识别、检测和辨别实现细粒度理解的细粒度开放词汇物体检测
Jiaming Li, Zhijia Liang, Weikai Chen, Lin Ma, Guanbin Li
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Meituan(美团)
;
GuangDong Province Key Laboratory of Information Security Technology(广东省信息安全技术重点实验室)
;
Research Institute, Sun Yat-sen University(中山大学深圳研究院)
From Questions to Queries: An AI-powered Multi-Agent Framework for Spatial Text-to-SQL
从问题到查询:一种基于AI的多智能体框架用于空间文本到SQL
Ali Khosravi Kazazi, Zhenlong Li, M. Naser Lessani, Guido Cervone
机构
*
Geoinformation and Big Data Research Laboratory, Department of Geography, The Pennsylvania State University(宾夕法尼亚州立大学地理系地理信息与大数据研究实验室)
;
Institute for Computational and Data Sciences and Department of Geography, The Pennsylvania State University(宾夕法尼亚州立大学计算与数据科学研究所及地理系)