arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2606.30168 2026-06-30 cs.CV 85%

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

潜在噪声掩码:减少多模态大语言模型中的视觉冗余

Kai Jiang, Ruishu Zhu, Siqi Huang, Hongyuan Zhang, Xuelong Li

机构 * School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院、光学和电子学(iOPEN)、西北工业大学) Institute of Artificial Intelligence, China Telecom (TeleAI)(人工智能研究院、中国电信(TeleAI)) Fudan University(复旦大学) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出Lens框架,通过轻量级LET令牌为视觉令牌评分,并注入自适应噪声抑制低相关令牌,在不改变模型结构的情况下提升多模态推理性能。

Comments 21 pages, 7 figures;

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27161 2026-06-26 cs.AI 新提交 85%

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference

TOPS:通过构建令牌最优保留集实现高效多模态大语言模型推理的第一性原理视觉令牌剪枝

Tinghao Wang, Yichen Guo, Rui Huang, Zheng Lu, Qizhe Zhang, Chenxi Li, Yuan Zhang, Jiajun Cao, Zhirong Shen, Yaosong Du, Guangyan Gan, Wenya Wang, Lin William Cong, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室) University of Electronic Science and Technology of China(电子科技大学) Nanyang Technological University(南洋理工大学) Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 提出TOPS方法,基于信息论分析确立任务相关性、信息覆盖和语义多样性三大原则,无训练且模型无关地剪枝视觉令牌,在LLaVA-NeXT上剪除77.8%令牌仍保持甚至提升性能。

Comments 27 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24165 2026-06-24 cs.CV 新提交 85%

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

多模态大语言模型中基于谱演化引导的令牌剪枝

Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang, Jianwei Yin, Zhi Chen

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字化服务技术重点实验室) The University of Queensland(昆士兰大学) Singapore Management University(新加坡管理大学) The University of Southern Queensland(南昆士兰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV

AI总结 提出跨层谱演化(CLSE)框架,通过频域分析令牌表示在Transformer层间的演化来评估重要性,实现无训练剪枝,在保持性能的同时减少计算开销。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20641 2026-06-23 cs.RO cs.AI cs.LG 新提交 85%

MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning

MAGNIFIED: 多模态大语言模型的强化学习微调用于运动规划

Letian Chen, Yiren Lu, Justin Fu, Yichen Xie, Runsheng Xu, Jyh-Jing Hwang, Ben Sapp, Drago Anguelov

机构 * Waymo LLC(Waymo有限责任公司)

专题命中 多模态训练与对齐 :multimodal(title);MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.AI

AI总结 提出MAGNIFIED方法,通过强化学习微调多模态大语言模型,利用令牌级奖励优化规划目标,在Waymo数据集上显著降低重叠率和偏离道路率。

Journal ref ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14700 2026-06-15 cs.CV 新提交 85%

RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space

RepFusion:利用多模态先验在表示空间中进行去噪

Xichen Pan, Aashu Singh, Satya Narayan Shukla, Xiangjun Fan, Shlok Kumar Mishra, Saining Xie

机构 * Meta AI New York University(纽约大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出RepFusion方法,利用多模态大语言模型作为噪声表示编码器,为扩散变压器提供条件信号,在相似推理预算下优于新初始化解码器基线。

Comments Project Page: https://xichenpan.com/repfusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10504 2026-06-10 cs.AI 新提交 85%

Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm

无配对数据的跨模态知识蒸馏:理论基础与算法

Trong Khiem Tran, Anh Duc Chu, Quang Hung Pham, Phi Le Nguyen, Trong Nghia Hoang

机构 * School of Information and Communications Technology, Hanoi University of Science and Technology, Hanoi, Vietnam(信息与通信技术学院,河内科学技术大学,越南河内) School of Electrical Engineering and Computer Science, Washington State University, Pullman, US(电气工程与计算机科学学院,华盛顿州立大学,华盛顿州普尔曼)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 提出无配对数据下的跨模态知识蒸馏框架,通过特征对齐和标签对齐两种分布对齐机制,实现跨模态知识迁移,理论保证且实验效果显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02941 2026-06-02 cs.CV 85%

MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion

MMTalker: 多分辨率3D说话头合成与多模态特征融合

Bin Liu, Zhixiang Xiong, Zhifen He, Bo Li

机构 * IEEE Publication Technology Group(IEEE出版技术组) Piscataway, NJ(新泽西州皮萨卡威)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出一种基于多分辨率表示和多模态特征融合的3D语音驱动面部动画合成方法MMTalker,通过网格参数化、非均匀可微采样、残差图卷积网络和双交叉注意力机制,实现高唇同步精度和逼真面部表情。

Comments This article presents only the preliminary research results, which are not yet complete and lack necessary supplementary experiments. The author has decided to withdraw it to improve the research work, and will submit a more complete version in the future

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06453 2026-06-01 cs.AI 85%

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

ConSensus:面向多模态感知的多智能体协作

Hyungjun Yoon, Mohammad Malekzadeh, Sung-Ju Lee, Fahim Kawsar, Lorena Qendro

机构 * KAIST(韩国科学技术院) Nokia Bell Labs(诺基亚贝尔实验室) University of Glasgow(格拉斯哥大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 提出ConSensus,一种无需训练的多智能体协作框架,通过将多模态感知任务分解为专用智能体并采用混合融合机制,在五个基准上平均准确率提升7.1%,融合token成本降低12.7倍。

Comments Accepted to ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28822 2026-05-29 cs.CL 85%

Lightweight Multimodal LLM-Enabled Cost-Effective Defect Grading of Power Transmission Equipment

轻量级多模态大语言模型驱动的输电设备经济高效缺陷分级

Tao Wang, Lipeng Zhu, Jiayong Li, Feng Gao, Siwen Liang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CL

AI总结 提出基于多模态大语言模型的缺陷分级框架,通过上下文学习最大化商业模型潜力,并利用链式思考问答对微调轻量级模型,实现低成本高精度分级。

Comments 9pages, 6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00857 2026-05-26 cs.LG cs.AI 85%

MultiPUFFIN: A Multimodal Domain-Constrained Foundation Model for Molecular Property Prediction of Small Molecules

MultiPUFFIN:用于小分子性质预测的多模态领域约束基础模型

Idelfonso B. R. Nogueira, Carine M. Rebello, Mumin Enis Leblebici, Erick Giovani Sperandio Nascimento

机构 * Department of Chemical Engineering, Norwegian University of Science and Technology (NTNU)(挪威科学与技术大学化学工程系) Faculty of Industrial Engineering, KU Leuven(鲁文大学工业工程学院) University of Surrey(萨里大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 提出多模态基础模型MultiPUFFIN,融合SMILES、2D图、3D构象及实验条件,通过条件感知精炼和热力学约束头,在小样本下优于ChemBERTa-2,预测小分子热物理性质。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20035 2026-05-20 cs.CV 85%

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

面向高效多模态大语言模型的阶段自适应令牌选择

Zijie Xin, Jie Yang, Ruixiang Zhao, Tianyi Wang, Fengyun Rao, Jing Lyu, Xirong Li

机构 * Renmin University of China(中国人民大学) WeChat Vision, Tencent Inc.(腾讯微信视觉实验室)

专题命中 多模态训练与对齐 :omni-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 本文提出SEATS方法,通过阶段自适应的令牌选择技术,有效提升多模态大语言模型的推理效率,在保留96.3%原始性能的同时,实现9.3倍的FLOPs减少和4.8倍的prefill加速。

Comments Code Link: https://github.com/xxayt/SEATS

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18915 2026-05-20 cs.CR cs.AI 85%

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

DMN: 一种用于多图像输入多模态大语言模型的组合框架

Wenzhuo Xu, Zhipeng Wei, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou

机构 * AI Security Lab(360人工智能安全实验室) International Computer Science Institute(国际计算机科学研究所) UC Berkeley(加州大学伯克利分校) Beihang University(北航大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.AI

AI总结 本文提出DMN框架,通过分布式指令、多模态证据和数字链任务,提升多图像输入多模态大语言模型的 jailbreak 性能,实验表明其在GPT-4o、Gemini-2.5-pro和Claude Sonnet 4上的攻击成功率超过90%。

Comments ACL 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16179 2026-05-18 cs.CV 85%

MAgSeg: Segmentation of Agricultural Landscapes in High-Resolution Satellite Imagery using Multimodal Large Language Models

MAgSeg:利用多模态大语言模型对高分辨率卫星图像进行农业景观分割

Piyush Tiwary, Utkarsh Ahuja, Depanshu Sani, Aishwarya Jayagopal, Sagar Gubbi, Subhashini Venugopalan, Alok Talekar, Vaibhav Rajan

机构 * Google DeepMind(谷歌DeepMind) Google(谷歌) Indian Institute of Science(印度科学研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出MAgSeg,一种无需视觉解码器的多模态大语言模型分割方法,有效解决南半球农业景观分割中的碎片化地块、高类内方差和标注数据稀缺问题,实现高效农业环境制图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11705 2026-05-13 cs.CV 85%

CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection

CAST:面向多模态聚类选择的坍缩感知多尺度拓扑融合

Boran Zhao, Hetian Liu, Zhenxian Hu, Yuqing Yuan, Yu Yan, Pengju Ren

机构 * School of Software Engineering, the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(软件工程学院、人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院) School of Software Engineering(软件工程学院) XJTU-POLIMI Joint School(西交大-波兰理工联合学院) Faculty of Electronic and Information Engineering(电子与信息工程学院) School of Human Settlements and Civil Engineering(人居与土木工程学院) the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出CAST框架,通过多尺度拓扑融合解决多模态数据集选择中的跨模态信息失衡和分布不匹配问题,提升跨架构泛化能力和能效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29162 2026-05-12 cs.MM 85%

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

从自然对齐到条件可控性在多模态对话中

Zeyu Jin, Songtao Zhou, Haoyu Wang, Minghao Tian, Kaifeng Yun, Zhuo Chen, Xiaoyu Qin, Jia Jia

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.MM

AI总结 本文提出MM-Dia数据集和MM-Dia-Bench测试集,通过多模态条件控制提升对话生成的可控性和表达性,实验表明现有框架难以复现人类交互的细腻表达。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15050 2026-04-28 cs.CV 85%

DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning

DRIFT:用于高效MLLM微调的推理先验转移

Chao Huang, Zeliang Zhang, Jiang Liu, Ximeng Sun, Jialian Wu, Xiaodong Yu, Ze Wang, Chenliang Xu, Emad Barsoum, Zicheng Liu

机构 * University of Rochester(罗切斯特大学) AMD

专题命中 多模态训练与对齐 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV

AI总结 DRIFT通过在梯度空间中转移推理知识,实现高效稳定的多模态推理迁移,优于传统方法并在数据和计算上更高效。

Comments ACL 2026 camera-ready; Project Page: https://wikichao.github.io/DRIFT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17087 2026-04-21 cs.CV cs.LG 85%

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

EvoComp: 通过语义引导的进化标注学习多模态大语言模型的视觉令牌压缩

Jiafei Song, Fengwei Zhou, Jin Qu, Wenjin Jason Li, Tong Wu, Gengjian Xue, Zhikang Zhao, Daomin Wei, Yichao Lu, Bailin Na

机构 * OPPO CTG

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出EvoComp框架,通过语义引导的进化标注策略,有效压缩视觉令牌数量,同时保持任务精度,实验显示在3倍压缩下保持99.3%的准确率,并在移动设备上提升1.6倍速度。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16952 2026-04-21 cs.CV 85%

Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder

更少的协同,更好的表现:通过条件和降质掩码自编码器解决异质多模态图像联合预训练

Bowen Peng, Yongxiang Liu, Jie Zhou, Xiaodong Chen, Tianpeng Liu, Xiaogang Yu, Li Liu

机构 * College of Electronic Science and Technology, National University of Defense Technology (NUDT)(电子科学与技术学院,国防科技大学) Beijing Institute of Remote Sensing Information(遥感信息研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出CoDe-MAE,通过Optical-anchored Knowledge Distillation和Conditioned Contrastive Learning解决高分辨率多模态联合预训练中的异质性-分辨率悖论,有效防止表示退化并在多个下游任务中取得新突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26786 2026-03-31 cs.LG cs.AI 85%

A Step Toward Federated Pretraining of Multimodal Large Language Models

迈向多模态大语言模型联邦预训练的一小步

Baochen Xiong, Yifan Xu, Xiaoshan Yang, Yaguang Song, Yaowei Wang, Changsheng Xu

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室) King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) Pengcheng Laboratory(鹏城实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出Fed-MA任务,通过冻结视觉编码器和LLM,协同训练跨模态投影器,解决参数干扰和梯度震荡问题,提出Fed-CMP框架在联邦预训练中取得显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21077 2026-03-30 cs.CV 85%

CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models

CoVFT:面向多模态大语言模型的上下文感知视觉微调

Nan Zhou, Huiqun Wang, Yaoyan Zheng, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment, Beihang University(北京航空航天大学复杂关键软件环境国家重点实验室) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出CoVFT框架,通过整合上下文向量提取和上下文混合专家模块,解决多模态任务中视觉微调的不稳定性问题,实现更稳定的视觉更新。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20808 2026-03-24 cs.CV cs.LG 85%

Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

预测正则化对抗多模态大语言模型中的视觉表征退化

Enguang Wang, Qiang Wang, Yuanchen Wu, Ke Yan, Xinbin Yuan, Shouhong Ding, Xialei Liu, Ming-Ming Cheng

机构 * NKIARI VCIP, CS, Nankai University(VCIP计算机科学系,南开大学) AAIS, Nankai University(AAIS,南开大学) Tencent Youtu Lab(腾讯优设实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文研究多模态大语言模型中的视觉表征退化问题,提出预测正则化方法以维持视觉表征,提升视觉语言性能。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18462 2026-03-20 cs.AI 85%

AlignMamba-2: Enhancing Multimodal Fusion and Sentiment Analysis with Modality-Aware Mamba

AlignMamba-2:通过模态感知Mamba增强多模态融合与情感分析

Yan Li, Yifei Xing, Xiangyuan Lan, Xin Li, Haifeng Chen, Dongmei Jiang

机构 * Pengcheng Laboratory(鹏城实验室) Shaanxi University of Science & Technology(陕西科技大学) School of Computer Science(计算机科学学院) Northwestern Polytechnical University(西北工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI

AI总结 本文提出AlignMamba-2框架,通过双对齐策略和模态感知Mamba层,提升多模态融合与情感分析的效率与效果,实验表明其在动态时间序列和静态图像任务中均达到新水平。

Comments Accepted by Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05566 2026-03-09 cs.LG cs.CL 85%

Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal Alignment

对齐真实语义:基于约束解耦和分布采样的跨模态对齐

Xiang Ma, Lexin Fang, Litian Xu, Caiming Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CL

AI总结 本文提出CDDS算法,通过约束解耦和分布采样方法,解决跨模态对齐中的语义分离和模态差距问题,实验表明其优于现有方法。

Comments AAAI 2026 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21589 2026-02-26 cs.CV 85%

SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction

SEF-MAP:子空间分解专家融合用于鲁棒多模态高精度地图预测

Haoxiang Fu, Lingfeng Zhang, Hao Li, Ruibing Hu, Zhengrong Li, Guanjing Liu, Zimu Tan, Long Chen, Hangjun Ye, Xiaoshuai Hao

机构 * National University of Singapore(新加坡国立大学) Xiaomi EV(小米电动车) Chinese University of Hong Kong(香港中文大学) The University of Manchester(曼彻斯特大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 SEF-MAP通过子空间分解和专家融合,提升多模态HD地图预测的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13606 2026-02-17 cs.NI cs.AI cs.ET cs.LG 85%

Multi-Modal Sensing and Fusion in mmWave Beamforming for Connected Vehicles: A Transformer Based Framework

毫米波波束成形中多模态感知与融合:基于变换器的框架

Muhammad Baqer Mollah, Honggang Wang, Mohammad Ataul Karim, Hua Fang

机构 * Department of Electrical and Computer Engineering, University of Massachusetts Dartmouth(电子与计算机工程系,马萨诸塞大学达特茅斯分校) Department of Graduate Computer Science and Engineering, Katz School of Science and Health, Yeshiva University(研究生计算机科学与工程系,耶鲁大学科学与健康学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出基于变换器的多模态感知与融合框架,用于毫米波波束成形,以减少波束训练开销并提高连接车辆的通信效率。

Comments 13 Pages. arXiv admin note: text overlap with arXiv:2509.11112

Journal ref IEEE Transactions on Vehicular Technology, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00185 2026-02-11 cs.AI cs.LG 85%

Chunking Strategies for Multimodal AI Systems

多模态AI系统中的分块策略

Shashanka B R, Mohith Charan R, Seema Banu F

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 本文综述了多模态系统中分块策略的分类和技术分析,探讨了不同模态的数据处理方法及挑战。

Comments 50 pages, 5 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09483 2026-02-11 cs.CV 85%

Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions

超越单个词对齐:通过令牌交互蒸馏多模态大语言模型

Lin Chen, Xiaoke Zhao, Kun Ding, Weiwei Feng, Changtao Miao, Zili Wang, Wenxuan Guo, Ying Wang, Kaiyuan Zheng, Bo Zhang, Zhe Li, Shiming Xiang

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所信息与智能系统研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 Align-TI通过令牌交互改进知识蒸馏,实现多模态大语言模型的高效压缩与性能提升

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15831 2026-02-11 cs.CV 85%

UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment

UniFit: 向基于多模态大语言模型引导的语义对齐的通用虚拟试衣迈进

Wei Zhang, Yeying Jin, Xin Li, Yan Zhang, Xiaofeng Cong, Cong Wang, Fengcai Qiao, zhichao Lian

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 UniFit通过多模态大语言模型引导的语义对齐模块,解决虚拟试衣中语义差距和数据稀缺问题,实现通用且高性能的试衣框架。

Comments accepted to AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15885 2025-12-19 cs.CV cs.AI cs.CL cs.MM 85%

Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models

超越文字:面向多模态大语言模型的自监督视觉学习

Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Pier Luigi Dovesi, Shaghayegh Roohi, Mark Granroth-Wilding, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) AMD Silo AI

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 JARVIS通过自监督视觉学习提升多模态大语言模型的视觉推理能力,无需依赖语言监督。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20322 2025-12-19 cs.CV 85%

HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models

HyperET: 在超几何空间中为多模态大语言模型实现高效训练

Zelin Peng, Zhengqin Xu, Qingyang Liu, Xiaokang Yang, Wei Shen

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, School of Computer Science, SJTU(摩埃人工智能重点实验室、人工智能学院、计算机科学学院、上海交通大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 HyperET通过在双曲空间中动态调整半径,实现多模态大语言模型的高效训练,提升跨模态对齐性能。

Comments Accepted by NeurIPS2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏