Fairness-Aware Fine-Tuning of Vision-Language Models for Medical Glaucoma Diagnosis
面向医疗青光眼诊断的公平性感知视觉语言模型微调
Zijian Gu, Yuxi Liu, Zhenhao Zhang, Song Wang
机构
*
Department of Computer Science, University of Rochester, NY, USA(罗切斯特大学计算机科学系)
;
Biostatistics and Health Data Science, School of Medicine, Indiana University, Indianapolis, IN, USA(印第安纳大学医学院生物统计学与健康数据科学系)
;
Department of Computer Science, University of Central Florida, FL, USA(佛罗里达州立大学计算机科学系)
3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection
面向MRI多器官异常检测的3D模态感知预训练
Haowen Zhu, Ning Yin, Xiaogen Zhou
机构
*
School of Electronic, Electrical Engineering and Physics, Fujian University of Technology(福建工程学院电子电气工程学院)
;
School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院)
;
Department of Medical Imaging, Suzhou Traditional Chinese Medicine Hospital, China(苏州中医医院影像科)
MOSAIC: A Unified Platform for Cross-Paradigm Comparison and Evaluation of Homogeneous and Heterogeneous Multi-Agent RL, LLM, VLM, and Human Decision-Makers
MOSAIC:一个用于跨范式比较和评估同质和异质多智能体RL、LLM、VLM和人类决策者的统一平台
Abdulhamid M. Mousa, Yu Fu, Rakhmonberdi Khajiev, Jalaledin M. Azzabi, Abdulkarim M. Mousa, Peng Yang, Yunusa Haruna, Ming Liu
机构
*
School of Optics and Photonics, Beijing Institute of Technology, Beijing 100081, China(北京理工大学光学工程学院)
;
School of Automation Science and Electrical Engineering, Beihang University, Beijing 100191, China(北京航空航天大学自动化科学与电气工程学院)
;
Faculty of Science, Ain Shams University, Cairo, Egypt(爱思唯命大学科学学院)
机构
*
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室)
;
Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院)
;
Xi’an Jiaotong University(西安交通大学)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
为什么强化学习比监督微调在泛化能力上更优?一种以数据为中心的视觉语言模型后训练视角
Aojun Lu, Tao Feng, Hangjie Yuan, Wei Li, Yanan Sun
机构
*
College of Computer Science, Sichuan University(四川大学计算机科学学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
Innovator-VL:一种用于科学发现的多模态大语言模型
Zichen Wen, Boxue Yang, Shuang Chen, Yaojie Zhang, Yuhang Han, Junlong Ke, Cong Wang, Yicheng Fu, Jiawang Zhao, Jiangchao Yao, Xi Fang, Zhen Wang, Henxing Cai, Lin Yao, Zhifeng Gao, Yanhui Hong, Nang Yuan, Yixuan Li, Guojiang Zhao, Haoyi Tao, Nan Wang, Han Lyu, Guolin Ke, Ning Liao, Xiaoxing Wang, Kai Chen, Zhiyu Li, Feiyu Xiong, Sihan Hu, Kun Chen, Yanfeng Wang, Weinan E, Linfeng Zhang, Linfeng Zhang
机构
*
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
;
Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI
Chenxu Dang, Jie Wang, Guang Li, Zhiwen Hou, Zihan You, Hangjun Ye, Jie Ma, Long Chen, Yan Wang
机构
*
Huazhong University of Science and Technology(华中科技大学)
;
Xiaomi EV(小米电动车)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
专题命中
VLM训练与架构
:vision-language model(title);vision language model(abstract);分类 cs.CV、cs.AI
机构
*
Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳))
;
Hong Kong Baptist University, China(香港 Baptist大学)
;
City University of Hong Kong, China(香港城市大学)
机构
*
Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学)
;
Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学)
;
Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科)
;
Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院)
;
School of Medicine Shaoxing University(绍兴大学医学院)
专题命中
VLM训练与架构
:multimodal large language model(title);LLaVA(abstract);分类 cs.AI、cs.LG
AI总结
Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
令牌空间中的扩展能力:对大视觉语言模型的分析
Tenghui Li, Guoxu Zhou, Xuyang Zhao, Qibin Zhao
机构
*
School of Automation, Guangdong University of Technology(广东工业大学自动化学院)
;
Key Laboratory of Intelligent Detection and the Internet of Things in Manufacturing, Ministry of Education(教育部智能制造智能检测与物联网重点实验室)
;
Guangdong Provincial Key Laboratory of Intelligent Systems and Optimization Integration(广东省智能系统与优化集成重点实验室)
;
Medical Science Data-driven Mathematics Team, RIKEN Center for Interdisciplinary Theoretical and Mathematical Sciences(RIKEN跨学科理论与数学科学中心医学科学数据驱动数学团队)
;
Medical Data Mathematical Reasoning Special Team, RIKEN Center for Integrative Medical Sciences(RIKEN整合医学科学中心医学数据数学推理特别团队)
;
Department of Artificial Intelligence Medicine, Chiba University(千叶大学人工智能医学系)
;
Tensor Learning Team, RIKEN Center for Advanced Intelligence Project(RIKEN高级人工智能项目中心张量学习团队)
专题命中
VLM训练与架构
:vision language model(title);vision-language model(abstract);分类 cs.AI、cs.LG