arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-19 至 2026-03-19 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 8 篇

2603.17705 2026-03-19 cs.CV 83%

Parameter-Efficient Modality-Balanced Symmetric Fusion for Multimodal Remote Sensing Semantic Segmentation

参数高效模态平衡对称融合用于多模态遥感语义分割

Haocheng Li, Juepeng Zheng, Shuangxi Miao, Ruibo Lu, Guosheng Cai, Haohuan Fu, Jianxi Huang

机构 * College of Land Science and Technology, China Agricultural University(中国农业大学土地科学与技术学院) Key Laboratory of Remote Sensing for Agri-Hazards, Ministry of Agriculture and Rural Affairs(农业农村部农业灾害遥感重点实验室) Faculty of Geosciences and Engineering, Southwest Jiaotong University(西南交通大学地质科学与工程学院) School of Artificial Intelligence, Sun Yat-Sen University(中山大学人工智能学院) Henan Polytechnic University(河南理工大学) Key Laboratory of Spatio-Temporal Information and Ecological Restoration of Mines, Ministry of Natural Resources of the People’s Republic of China(矿产资源时空信息与生态修复重点实验室) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) National Supercomputing Center in Shenzhen, Shenzhen, China(深圳国家超算中心) Ministry of Education Key Laboratory for Earth System Modeling and the Department of Earth System Science, Tsinghua University(地球系统模拟教育部重点实验室和清华大学地球系统科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MoBaNet,通过参数高效和模态平衡的对称融合框架,在减少可训练参数的同时提升多模态遥感语义分割的鲁棒性和平衡性。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16421 2026-03-19 cs.CV 83%

HGP-Mamba: Integrating Histology and Generated Protein Features for Mamba-based Multimodal Survival Risk Prediction

HGP-Mamba:整合组织学与生成的蛋白质特征用于基于Mamba的多模态生存风险预测

Jing Dai, Chen Wu, Ming Wu, Qibin Zhang, Zexi Wu, Jingdong Zhang, Hongming Xu

机构 * Cancer Hospital of Dalian University of Technology, Shenyang, China(大连理工大学沈阳医院) School of Biomedical Engineering, Faculty of Medicine, Dalian University of Technology, Dalian, China(大连理工大学生物医学工程学院) Key Laboratory of Integrated Circuit and Biomedical Electronic System, Dalian University of Technology, Dalian, China(大连理工大学集成电路与生物医学电子系统重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出HGP-Mamba框架,通过整合组织学与生成蛋白质特征,提升多模态生存风险预测的效率与性能。

Comments Accepted at IEEE ICME 2026. This arXiv version includes additional supplementary experiments and extended discussions beyond the conference version

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17753 2026-03-19 cs.CV 79%

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

PC-CrossDiff:点-簇双级跨模态微分注意力用于统一的3D指称与分割

Wenbin Tan, Jiawen Lin, Fangyong Wang, Yuan Xie, Yong Xie, Yachao Zhang, Yanyun Qu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PC-CrossDiff框架,通过双级跨模态微分注意力解决复杂多物体场景中指称理解和分割的挑战,提升3D视觉定位的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17347 2026-03-19 cs.MM 79%

Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning

超越强制模态平衡:多模态学习中的内在信息预算

Zechang Xiong, Da Li, Kexin Tang, Pengyuan Li, Wenkang Kong, Yulan Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出IIBalance框架,通过内在信息预算对齐模态贡献,解决多模态学习中的模态不平衡问题,实验表明其优于现有方法。

Comments 6 pages, 4 figures, paper accepted by ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17228 2026-03-19 cs.CV cs.AI cs.LG 73%

From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs

从丢弃到恢复:对MLLMs分割能力的机理分析

Boyong Wu, Sanghwan Kim, Zeynep Akata

机构 * Technical University of Munich(慕尼黑技术大学) Helmholtz Munich(亥姆霍兹慕尼黑) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML))

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文通过逐层线性探测评估MLLMs整个流程,揭示适配器导致的分割表示下降及LLM层通过注意力机制逐步恢复的机制,为未来分割模型设计提供依据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17809 2026-03-19 cs.CV cs.AI 62%

Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients

细粒度后训练量化用于大型视觉语言模型的量化感知集成梯度

Ziwei Xiang, Fanhu Zeng, Hongjian Fang, Rui-Qi Wang, Renxing Chen, Yanan Zhu, Yi Chen, Peipei Yang, Xu-Yao Zhang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,中国科学院自动化所) School of Artificial Intelligence, UCAS(人工智能学院,中国科学院大学) Beijing National Research Center for Information Science and Technology(北京信息科学研究中心) Institute of Artificial Intelligence, USTB(信息科学技术大学人工智能学院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Zhongguancun Academy(中关村学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出细粒度后训练量化方法,通过量化感知集成梯度评估token敏感性,提升大型视觉语言模型的精度与效率。

Comments Accepted by CVPR 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17647 2026-03-19 cs.CV 57%

Part-Aware Open-Vocabulary 3D Affordance Grounding via Prototypical Semantic and Geometric Alignment

面向部分的开放词汇3D功能接地 via 语义和几何对齐

Dongqiang Gou, Xuming He

机构 * ShanghaiTech University(上海科技大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种双阶段跨模态框架,通过增强语义和几何表示,解决开放词汇3D功能接地中的语义一致性、几何对齐和开放词汇泛化问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17311 2026-03-19 cs.CL 57%

Ruyi2.5 Technical Report

Ruyi2.5 技术报告

Huan Song, Shuyu Tian, Qingfei Zhao, Wenhao Hong, Jiang Liu, Ting Long, Jiawei Shao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 Ruyi2.5基于AI Flow框架构建多模态家族模型,通过共享主干架构实现多尺度模型协同训练,提升部署层级语义一致性,同时提出隐私保护摄像头系统及BPPO算法加速强化学习微调。

详情

展开后加载摘要…

URL PDF HTML 收藏