arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-09 至 2026-04-09 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 7 篇

2604.06728 2026-04-09 cs.CV cs.AI cs.MM 87%

URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection

URMF:面向多模态讽刺检测的不确定性感知鲁棒多模态融合

Zhenyu Wang, Weichen Cheng, Weijia Li, Junjie Mou, Zongyou Zhao, Guoying Zhang

机构 * School of Artificial Intelligence, China University of Mining and Technology-Beijing(中国矿业大学(北京)人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出URMF框架,通过建模模态可靠性提升多模态讽刺检测的准确性和鲁棒性,采用多头交叉注意力和自注意力机制,并结合不确定性建模和联合训练目标,在公开基准上优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07141 2026-04-09 cs.CV 79%

USCNet: Transformer-Based Multimodal Fusion with Segmentation Guidance for Urolithiasis Classification

USCNet:基于Transformer的多模态融合与分割引导的尿路结石分类

Changmiao Wang, Songqi Zhang, Yongquan Zhang, Yifei Wang, Liya Liu, Nannan Li, Xingzhi Li, Jiexin Pan, Yi Jiang, Xiang Wan, Hai Wang, Ahmed Elazab

机构 * Shenzhen Research Institute of Big Data(深圳市大数据研究院) Zhejiang University of Finance and Economics(浙江财经大学) Anhui University of Finance and Economics(安徽财经大学) The Second Affiliated Hospital of Chinese University of Hong Kong (Longgang District People’s Hospital of Shenzhen)(香港中文大学第二附属医院(深圳市龙岗区人民医院))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 USCNet通过融合CT图像与电子病历数据,利用Transformer架构实现尿路结石的预手术分类,提出动态损失函数平衡分割与分类任务,实验表明其分类效果优于现有方法。

Comments Accepted by IEEE Journal of Biomedical and Health Informatics. Early Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06912 2026-04-09 cs.CV cs.AI 76%

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

Q-Zoom:基于查询的自适应感知以实现高效多模态大语言模型

Yuheng Shi, Xiaohuan Pei, Linfeng Wen, Minjing Dong, Chang Xu

机构 * University of Sydney(悉尼大学) Sun Yat-sen University(中山大学) City University of Hong Kong(香港城市大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

AI总结 Q-Zoom通过高效粗到细的框架优化多模态大语言模型的高分辨率感知,提升推理速度并保持精度,实验表明其在文档识别和高分辨率场景中表现优异。

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07154 2026-04-09 cs.CV cs.AI 62%

Bridging MRI and PET physiology: Untangling complementarity through orthogonal representations

弥合MRI与PET生理学:通过正交表示解构互补性

Sonja Adomeit, Kartikay Tehlan, Lukas Förner, Katharina Weisser, Helen Scholtiseek, David Kaufmann, Julie Steinestel, Constantin Lapa, Thomas Kröncke, Thomas Wendler

机构 * Dept. of diagnostic and interventional Radiology and Neuroradiology, University Hospital Augsburg, Germany(德国奥格斯堡大学医院诊断与介入放射学及神经放射学系) Digital Medicine, University Hospital Augsburg, Germany(德国奥格斯堡大学医院数字医学) Chair for Computer Aided Medical Procedures and Augmented Reality, Technical University of Munich, Germany(德国慕尼黑工业大学计算机辅助医疗流程与增强现实教席) Bavarian Center for Cancer Research (BZKF) Augsburg, Germany(德国奥格斯堡巴伐利亚癌症研究中心) Dept. of Nuclear Medicine, University Hospital Augsburg, Germany(德国奥格斯堡大学医院核医学系) Dept. of Urology, University Hospital Augsburg, Germany(德国奥格斯堡大学医院泌尿外科) Center for Advanced Analytics and Predictive Sciences, University of Augsburg, Germany(德国奥格斯堡大学高级分析与预测科学中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种子空间分解框架,通过正交子空间分离重构多模态融合,区分共享信息与模态特有信息,揭示PSMA PET信号中不可由MRI生理描述恢复的部分。

Comments The code is available at https://github.com/SonjaA14/inrmri2pet

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06267 2026-04-09 cs.LG cs.AI 57%

MO-RiskVAE: A Multi-Omics Variational Autoencoder for Survival Risk Modeling in Multiple MyelomaMO-RiskVAE

MO-RiskVAE: 一种多组学变分自编码器用于多发性骨髓瘤生存风险建模

Zixuan Chen, Heng Zhang, YuPeng Qin, WenPeng Xing, Qiang Wang, Da Wang, Changting Lin, Meng Han

机构 * Binjiang Institute of Zhejiang University(浙江大学滨江研究院) Zhejiang University(浙江大学) School of Medicine, Zhejiang University(浙江大学医学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文提出MO-RiskVAE,通过统一扩展MyeVAE框架,系统研究了潜变量建模选择对多模态生存预测的影响,发现生存驱动训练对潜变量正则化幅度和结构更敏感,通过混合连续-离散公式提升风险排序,改进了风险分层。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06203 2026-04-09 cs.CY cs.AI 57%

Front-End Ethics for Sensor-Fused Health Conversational Agents: An Ethical Design Space for Biometrics

传感器融合健康对话代理的前端伦理:生物特征的伦理设计空间

Hansoo Lee, Rafael A. Calvo

机构 * Imperial College London(伦敦帝国理工学院) Korea Institute of Science and Technology(韩国科学技术研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨了生物特征翻译的伦理设计,提出五个维度分析前端伦理风险,提出适应性披露作为安全机制,确保健康代理支持用户自主性。

Comments Accepted at the Proceedings of the CHI 2026 Workshop: Ethics at the Front-End

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07071 2026-04-09 cs.HC cs.CR 50%

BioMoTouch: Touch-Based Behavioral Authentication via Biometric-Motion Interaction Modeling

BioMoTouch:基于生物-运动交互建模的触控行为认证

Zijian Ling, Jianbang Chen, Hongwei Li, Hongda Zhai, Man Zhou, Jun Feng, Zhengxiong Li, Qi Li, Qian Wang

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 BioMoTouch通过整合电容触控屏与惯性传感器信号,建立生物特征与行为动态的联合模型,实现高鲁棒性的触控认证,实验显示其在99.71%的平衡准确率和0.27%的误判率下表现优异。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏