MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
MedLVR: 基于潜在视觉推理的可靠医学视觉问答
Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
机构
*
Department of Radiation Oncology and Winship Cancer Institute, Emory University School of Medicine(埃默里大学医学院放射肿瘤学系与温希普癌症研究所)
;
Department of Biostatistics and Bioinformatics, Emory University(埃默里大学生物统计学与生物信息学系)
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
医学VQA中通过置信度-证据贝叶斯增益实现确定性幻觉检测
Mohammad Asadi, Tahoura Nedaee, Jack W. O'Sullivan, Euan Ashley, Ehsan Adeli
机构
*
Department of Electrical Engineering, Stanford University, CA, USA(电气工程系,斯坦福大学)
;
Department of Biology, Stanford University, CA, USA(生物学系,斯坦福大学)
;
Division of Cardiology, Department of Medicine, Stanford University, CA, USA(心脏病学部,医学系,斯坦福大学)
;
Department of Biomedical Data Science, Stanford University, CA, USA(生物医学数据科学系,斯坦福大学)
;
Department of Computer Science, Stanford University, CA, USA(计算机科学系,斯坦福大学)
;
Department of Psychiatry and Behavioral Sciences, Stanford University, CA, USA(精神病学与行为科学系,斯坦福大学)
专题命中
视觉问答
:MLLM(summary_cn,abstract_cn);visual question answering(abstract);multimodal large language model(abstract);分类 cs.AI
Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model
人工智能生成的以人为中心的视频的多维质量评估:数据集与模型
Sijing Wu, Yunhao Li, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai
机构
*
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所)
;
USC-SJTU Institute of Cultural and Creative Industry, Shanghai Jiao Tong University(上海交通大学南加州大学文化创意产业学院)
;
Polytech Nantes, Université de Nantes(法国南特大学高等理工学院)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
NLPCC 2026共享任务1概述:难度感知多语言多模态医学教学视频理解评估
Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
机构
*
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与工程学院)
;
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算学系)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉
Pengxu Chen, Yao Zhu, Guangming Zhu, Jun Sheng, Jincai Huang, Xiangyang Ji, Liang Zhang
机构
*
Xidian University(西安电子科技大学)
;
Tsinghua University(清华大学)
;
Shanghai Road Transport Development Center(上海市道路运输发展中心)
;
Hunan Institute of Advanced Technology(湖南先进技术研究院)
Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Great Bay University(大湾区大学)
;
Dongguan Key Laboratory for Intelligence and Information Technology(东莞市智能信息技术重点实验室)
专题命中
视觉推理
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
School of Intelligent Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
School of Computer Science, Peking University(北京大学计算机科学学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
超越语义的艺术:用于多关系表示的层状信息对比学习
Ludovica Schaerf, Antonio Purificato, Piera Riccio, Fabrizio Silvestri, Noa Garcia
机构
*
University of Zurich(苏黎世大学)
;
Max Planck Institute Bibliotheca Hertziana(马克斯·普朗克赫兹iana图书馆研究所)
;
Sapienza University of Rome(罗马第一大学)
;
Amazon(亚马逊)
;
University of Amsterdam(阿姆斯特丹大学)
;
The University of Osaka(大阪大学)
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
TSRouter:用于时间序列推理的动态模态-模型选择
Fangxu Yu, Tao Feng, Dehai Min, Lu Cheng, Ge Liu, Tianyi Zhou
机构
*
University of Maryland, College Park(马里兰大学帕克分校)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
MBZUAI(Mohamed Bin Zayed University of Artificial Intelligence)
DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
MAGIS:基于证据的多智能体推理用于可解释的斜视临床决策
Xikai Tang, Yifan Wang, Jiafan Zhuang, Li Luo, Jinming Guo, Xiaoling Xie, Jiacheng Liu, Peiwei Wei, Lihao Zhong, Xiaoli Kang, Jie Cen, Guangqiang Yin, Kunliang Qiu, Ce Zheng, Zhun Fan
机构
*
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
;
Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳高等研究院)
;
Joint Shantou International Eye Center of Shantou University and The Chinese University of Hong Kong(汕头大学·香港中文大学联合汕头国际眼科中心)
;
School of Artificial Intelligence, Guangzhou City Polytechnic(广州城市职业学院人工智能学院)
;
Medical College, Shantou University(汕头大学医学院)
;
College of Engineering, Shantou University(汕头大学工学院)
;
Department of Ophthalmology, Xinhua Hospital Affiliated to Shanghai Jiaotong University School of Medicine(上海交通大学医学院附属新华医院眼科)
;
Shenzhen Loop Area Institute(深圳河套学院)