Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Great Bay University(大湾区大学)
;
Dongguan Key Laboratory for Intelligence and Information Technology(东莞市智能信息技术重点实验室)
专题命中
视觉推理
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
School of Intelligent Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
School of Computer Science, Peking University(北京大学计算机科学学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
超越语义的艺术:用于多关系表示的层状信息对比学习
Ludovica Schaerf, Antonio Purificato, Piera Riccio, Fabrizio Silvestri, Noa Garcia
机构
*
University of Zurich(苏黎世大学)
;
Max Planck Institute Bibliotheca Hertziana(马克斯·普朗克赫兹iana图书馆研究所)
;
Sapienza University of Rome(罗马第一大学)
;
Amazon(亚马逊)
;
University of Amsterdam(阿姆斯特丹大学)
;
The University of Osaka(大阪大学)
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
TSRouter:用于时间序列推理的动态模态-模型选择
Fangxu Yu, Tao Feng, Dehai Min, Lu Cheng, Ge Liu, Tianyi Zhou
机构
*
University of Maryland, College Park(马里兰大学帕克分校)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
MBZUAI(Mohamed Bin Zayed University of Artificial Intelligence)
DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
MAGIS:基于证据的多智能体推理用于可解释的斜视临床决策
Xikai Tang, Yifan Wang, Jiafan Zhuang, Li Luo, Jinming Guo, Xiaoling Xie, Jiacheng Liu, Peiwei Wei, Lihao Zhong, Xiaoli Kang, Jie Cen, Guangqiang Yin, Kunliang Qiu, Ce Zheng, Zhun Fan
机构
*
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
;
Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳高等研究院)
;
Joint Shantou International Eye Center of Shantou University and The Chinese University of Hong Kong(汕头大学·香港中文大学联合汕头国际眼科中心)
;
School of Artificial Intelligence, Guangzhou City Polytechnic(广州城市职业学院人工智能学院)
;
Medical College, Shantou University(汕头大学医学院)
;
College of Engineering, Shantou University(汕头大学工学院)
;
Department of Ophthalmology, Xinhua Hospital Affiliated to Shanghai Jiaotong University School of Medicine(上海交通大学医学院附属新华医院眼科)
;
Shenzhen Loop Area Institute(深圳河套学院)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV