arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-12 至 2026-03-12 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 3 篇

2401.16685 2026-03-12 cs.LG cs.DC 78%

Communication-Efficient Multimodal Federated Learning: Joint Modality and Client Selection

高效多模态联邦学习:联合模态与客户端选择

Liangqi Yuan, Dong-Jun Han, Su Wang, Devesh Upadhyay, Christopher G. Brinton

机构 * School of Electrical and Computer Engineering, Purdue University(普渡大学电气与计算机工程学院) Department of Computer Science and Engineering, Yonsei University(延世大学计算机科学与工程系) School of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程学院) Saab Inc.(Saab公司)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出MFedMC框架,通过解耦架构和联合选择算法,实现高效多模态联邦学习,减少通信开销并提升模型泛化能力。

Comments arXiv admin note: text overlap with arXiv:2310.07048

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03666 2026-03-12 cs.CV cs.AI 62%

MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operations

MonitorVLM:一种用于采矿作业中安全违规检测的视觉语言框架

Jiang Wu, Sichao Wu, Yinsong Ma, Guangyuan Yu, Haoyuan Xu, Lifang Zheng, Jingliang Duan

机构 * School of Mechanical Engineering, University of Science and Technology Beijing(北京科技大学机械工程学院) Laboratory for Computational Sensing and Robotics, Johns Hopkins University(约翰霍普金斯大学计算感知与机器人实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MonitorVLM通过视觉语言框架提升采矿作业中安全违规检测的精度和召回率,实现高效自动化监控。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10089 2026-03-12 stat.ME stat.AP 50%

Trajectory-informed graph-based clustering for longitudinal cancer subtyping

基于轨迹的图基团分析用于纵向癌症亚型分类

Lara Cavinato, Marco Rocchi, Luca Viganò, Francesca Ieva

专题命中 视频多模态 :multi-modal(abstract)

AI总结 本文提出了一种基于轨迹的图聚类方法,用于纵向癌症亚型分类,整合多模态临床数据和患者轨迹,以发现具有不同预后轨迹的亚型。

详情

展开后加载摘要…

URL PDF HTML 收藏