arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-17 至 2026-04-17 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 7 篇

2604.14656 2026-04-17 cs.AI cs.CL cs.CV 85%

Rethinking Patient Education as Multi-turn Multi-modal Interaction

重新思考患者教育作为多轮多模态交互

Zonghai Yao, Zhipeng Tang, Chengtao Lin, Xiong Luo, Benlu Wang, Juncheng Huang, Chin Siang Ong, Hong Yu

机构 * VA Bedford Health Care(VA贝德福德医疗中心) UMass Amherst(马萨诸塞大学阿默斯特分校) UMass Lowell(马萨诸塞大学洛厄尔分校) Yale University(耶鲁大学) National University of Singapore(新加坡国立大学) Yale School of Medicine(耶鲁医学院)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出MedImageEdu基准,通过多轮多模态交互提升患者教育效果,评估咨询过程和最终响应质量,发现多模态模型在视觉 grounding、安全性和情绪互动方面存在不足。

Comments Equal contribution for the first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14198 2026-04-17 cs.LG cs.AI cs.CL 81%

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining

MixAtlas: 多模态大语言模型中训练过程中的数据混合优化方法

Bingbing Wen, Sirajul Salekin, Feiyang Kang, Bill Howe, Lucy Lu Wang, Javier Movellan, Manjot Bilkhu

机构 * Apple(苹果公司) University of Washington(华盛顿大学) Virginia Tech(弗吉尼亚理工学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 MixAtlas通过双轴分解训练语料,结合小代理模型与高斯过程代理,优化多模态大语言模型的训练混合策略,提升性能和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14629 2026-04-17 cs.CV 70%

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models

Switch-KD:视觉切换知识蒸馏用于视觉-语言模型

Haoyi Sun, Xiaoxiao Wang, Ning Mao, Qian Wang, Lifu Mu, Wen Zheng, Tao Wei, Wei Chen

机构 * Li Auto Inc(利-auto公司)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 Switch-KD通过统一视觉-语言知识转移,解决多模态对齐问题,使小模型高效蒸馏大模型的多模态知识。

Comments 11 pages, 3 figures

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11176 2026-04-17 cs.CV 70%

Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment

多示踪PET的高精度合成通过VLM调制的校正流用于区分轻度认知障碍

Tuo Liu, Shuijin Lin, Shaozhen Yan, Haifeng Wang, Jie Lu, Jianhua Ma, Chunfeng Lian

机构 * School of Mathematics and Statistics, Xi'an Jiaotong University(西安交通大学数学与统计学学院) Key Laboratory of Biomedical Information Engineering of Ministry of Education, School of Life Science and Technology, Xi'an Jiaotong University(教育部生物医学信息工程重点实验室,西安交通大学生命科学与技术学院) Department of Radiology and Nuclear Medicine, Xuanwu Hospital, Capital Medical University(首都医科大学宣武医院放射科与核医学科) Research Center for Intelligent Medical Equipment and Devices (IMED), Xi'an Jiaotong University(智能医疗设备与器件研究中心(IMED),西安交通大学)

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出DIReCT$++$模型,结合MRI和临床信息,通过校正流和视觉语言模型生成高保真多示踪PET图像,实现轻度认知障碍的精准分层。

Comments Added supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18129 2026-04-17 cs.CV cs.CL 62%

One RL to See Them All: Visual Triple Unified Reinforcement Learning

一个RL看尽一切:视觉三合一强化学习

Yan Ma, Linge Du, Xuyang Shen, Shaoxiang Chen, Pengfei Li, Qibing Ren, Lizhuang Ma, Yuchao Dai, Pengfei Liu, Junjie Yan

机构 * Qwen Team(通义实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出V-Triune方法,通过三个协调抽象提升多模态RL性能,开发Orsta模型在多个任务中表现优异,展示统一RL在视觉语言模型中的优势。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04585 2026-04-17 cs.CV 57%

SAM3-I: Segment Anything with Instructions

SAM3-I: 基于指令的分割

Jingjing Li, Yue Feng, Yuchen Guo, Jincai Huang, Wei Ji, Qi Bi, Yongri Piao, Miao Zhang, Xiaoqi Zhao, Qiang Chen, Shihao Zou, Huchuan Lu, Li Cheng

机构 * University of Alberta(阿尔伯塔大学) Tencent WeChat(腾讯微信) Northwestern University(西北大学) SUSTech(四川大学) Yale University(耶鲁大学) Utrecht University(乌得勒支大学) Dalian University of Technology(大连理工大学) SIAT, Chinese Academy of Sciences(中科院深圳先进技术研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 SAM3-I通过整合概念级 grounding 和指令级推理,提升分割性能,支持复杂自然语言指令。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14474 2026-04-17 cs.LG 50%

Scouting By Reward: VLM-TO-IRL-Driven Player Selection For Esports

通过奖励进行 scouting:由 VLM-TO-IRL 驱动的电子竞技玩家选拔

Qing Yan, Wenyu Yang, Yufei Wang, Wenhao Ma, Linchong Hu, Yifei Jin, Anton Dahbura

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Pennsylvania(宾夕法尼亚大学) Cornell University(康奈尔大学)

专题命中 图文多模态 :multimodal(abstract)

AI总结 本文提出一种基于逆强化学习的电子竞技玩家选拔框架,通过学习专业选手的奖励函数,利用多模态双分支架构和 GAIL 目标实现玩家风格匹配评估,提升人才发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏