arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-06 至 2026-03-06 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 10 篇

2603.04890 2026-03-06 cs.LG cs.AI cs.CV 84%

FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation

FedAFD: 通过对抗融合与蒸馏实现多模态联邦学习

Min Tan, Junchao Ma, Yinfu Feng, Jiajun Ding, Wenwen Pan, Tingting Han, Qian Zheng, Zhenzhong Kuang, Zhou Yu

机构 * Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(浙江空间信息感知与传输重点实验室,杭州电子大学) Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University(复杂系统建模与仿真实验室,计算机科学与技术学院,杭州电子大学) Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 FedAFD通过对抗融合与蒸馏方法,解决多模态联邦学习中的模态差异、任务差异和模型异质性问题,提升客户端与服务器的学习性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22695 2026-03-06 cs.DC 82%

Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference

模态膨胀:MLLM推理的能量特性与优化机会

Mona Moghadampanah, Adib Rezaei Shahmirzadi, Farhana Amin, Dimitrios S. Nikolopoulos

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract)

AI总结 本文研究了多模态大语言模型推理中的能量特性与优化机会,发现模态膨胀导致额外能耗,并提出阶段级DVFS优化方法以提升能效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04887 2026-03-06 cs.CV 79%

Federated Modality-specific Encoders and Partially Personalized Fusion Decoder for Multimodal Brain Tumor Segmentation

联邦模态特定编码器和部分个性化融合解码器用于多模态脑肿瘤分割

Hong Liu, Dong Wei, Qian Dai, Xian Wu, Yefeng Zheng, Liansheng Wang

机构 * National Institute for Data Science in Health and Medicine(国家医学数据科学研究院) Department of Computer Science at School of Informatics(信息学院计算机科学系) Jarvis Research Center(Jarvis研究中心) Medical Artificial Intelligence Lab(医学人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出FedMEPD框架,通过联邦模态特定编码器和部分个性化融合解码器,解决多模态医学图像分析中的模态间异质性和个性化需求问题。

Comments Medical Image Analysis 2025. arXiv admin note: substantial text overlap with arXiv:2403.11803

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04562 2026-03-06 cs.CV cs.LG 79%

Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data

深度学习中多模态遥感数据局部气候区分类的融合与分组策略

Ancymol Thomas, Jaya Sreevalsan-Nair

机构 * Graphics-Visualization-Computing Lab, International Institute of Information Technology Bangalore, Karnataka 560100, India(图形可视化计算实验室,国际信息学院班加罗尔,卡纳塔克邦560100,印度)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出了一种基于深度学习的多模态遥感数据局部气候区分类方法,通过融合与分组策略提升分类准确率,最终达到76.6%的整体准确率。

Comments 25 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23508 2026-03-06 cs.CL cs.AI 62%

Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective

为何强化微调能更好地使大语言模型保留先验知识:从数据角度出发

Zhihao Zhang, Qiaole Dong, Qi Zhang, Jun Zhao, Enyu Zhou, Zhiheng Xi, Senjie Jin, Xiaoran Fan, Yuhao Zhou, Mingqi Wu, Yanwei Fu, Tao Ji, Tao Gui, Xuanjing Huang, Kai Chen

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文通过数据视角揭示强化微调(RFT)相较于监督微调(SFT)在保留大语言模型先验知识方面的优势,发现RFT通过强化正确样本与基模型对齐,有效减少对先验知识的干扰。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05465 2026-03-06 cs.CV 57%

HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token

在不生成单个标记的情况下检测视觉-语言模型中的幻觉

Sai Akhil Kogilathota, Sripadha Vallabha E G, Luzhe Sun, Jiawei Zhou

机构 * Stony Brook University(石英溪大学) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 通过探测模型内部表示在生成前检测视觉-语言模型的幻觉风险,展示不同架构中信息丰富的层和模态差异,并验证轻量级探测器在提升安全性和效率方面的潜力。

Journal ref The 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22091 2026-03-06 cs.CV 57%

Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

学习驾驶是免费礼物:从未经处理的现实视频中进行大规模无标签自主性预训练

Matthew Strong, Wei-Jer Chang, Quentin Herau, Jiezhi Yang, Yihan Hu, Chensheng Peng, Wei Zhan

机构 * Applied Intuition Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种无标签、教师引导的框架,通过未经处理的现实视频学习自动驾驶表示,无需姿态、标签或LiDAR,实现高效的自动驾驶感知和规划。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17488 2026-03-06 cs.CV 57%

Optimizing Multi-Modality Trackers via Significance-Regularized Tuning

通过显著性正则化调优优化多模态跟踪器

Zhiwen Chen, Jinjian Wu, Zhiyu Zhu, Yifan Zhang, Guangming Shi, Junhui Hou

机构 * School of Artificial Intelligence, Xidian University(电子科技大学人工智能学院) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) School of Mechatronic Engineering and Automation, Shanghai University(上海大学机电工程与自动化学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出显著性正则化微调框架,通过引入参数显著性提升多模态跟踪器的跨模态可迁移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06515 2026-03-06 cs.CV 57%

RESAR-BEV: An Explainable Progressive Residual Autoregressive Approach for Camera-Radar Fusion in BEV Segmentation

RESAR-BEV:一种可解释的逐步残差自回归方法用于摄像头-雷达融合的鸟瞰图分割

Zhiwen Zeng, Yunfei Yin, Zheng Yuan, Argho Dey, Xianjian Bao

机构 * College of Computer Science, Chongqing University(重庆大学计算机学院) Department of Computer Science, Maharishi University of Management(Maharishi大学管理学院计算机系)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 RESAR-BEV通过残差自回归学习和双路径特征编码,实现摄像头-雷达融合的鸟瞰图分割,取得7类驾驶场景54.0% mIoU的最优性能并保持实时处理能力。

Comments This work was submitted to IEEE Transactions on Intelligent Transportation Systems (T-ITS) on 09-May-2025; revised 5 October 2025 and 26 January 2026; accepted 1 March 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04606 2026-03-06 cs.LG physics.plasm-ph 50%

PDE foundation model-accelerated inverse estimation of system parameters in inertial confinement fusion

偏微分方程基础模型加速惯性约束聚变系统参数反演

Mahindra Rautela, Alexander Scheinker, Bradley Love, Diane Oyen, Nathan DeBardeleben, Earl Lawrence, Ayan Biswas

机构 * AOT-IC Group, Los Alamos National Laboratory(AOT-IC组,洛斯阿拉莫斯国家实验室) CAI Division, Los Alamos National Laboratory(CAI部门,洛斯阿拉莫斯国家实验室) HPC Division, Los Alamos National Laboratory(HPC部门,洛斯阿拉莫斯国家实验室)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本研究利用偏微分方程基础模型加速惯性约束聚变系统参数反演,通过微调和轻量级头部训练实现高精度超光谱重建与参数回归。

详情

展开后加载摘要…

URL PDF HTML 收藏