arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-06-24 至 2026-06-24 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 10 篇

2606.23885 2026-06-24 cs.CV cs.AI cs.CL cs.MM 新提交 87%

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

注意头:多模态大语言模型的拓扑表示对齐

Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi

机构 * University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学) University of Pisa(比萨大学) AMD Silo AI

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出头级表示对齐(HeRA)方法,在注意力头级别强制跨模态对齐,通过对比目标匹配局部拓扑结构,选择对齐最差的头进行训练,有效提升视觉任务性能并减少幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24165 2026-06-24 cs.CV 新提交 85%

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

多模态大语言模型中基于谱演化引导的令牌剪枝

Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang, Jianwei Yin, Zhi Chen

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字化服务技术重点实验室) The University of Queensland(昆士兰大学) Singapore Management University(新加坡管理大学) The University of Southern Queensland(南昆士兰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV

AI总结 提出跨层谱演化(CLSE)框架,通过频域分析令牌表示在Transformer层间的演化来评估重要性,实现无训练剪枝,在保持性能的同时减少计算开销。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14001 2026-06-24 cs.CV cs.AI cs.LG 版本更新 84%

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

MOCHA: 多模态对象感知的跨架构对齐

Elena Camuffo, Francesco Barbato, Mete Ozay, Simone Milani, Umberto Michieli

机构 * University of Padova(帕多瓦大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出MOCHA蒸馏框架,将冻结VLM教师的多模态区域知识迁移至轻量视觉检测器,通过双重损失实现局部对齐与全局关系一致性,在少样本个性化检测中平均提升10.1%。

Comments 18 pages main paper, 10 pages supplementary material

Journal ref Proceedings of ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23888 2026-06-24 eess.IV cs.AI cs.CV 新提交 81%

E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

E-MRL: 跨视图对齐的证据驱动多模态强化学习用于可靠的3D肿瘤分析

Sijing Li, Zhongwei Qiu, Zhuoya Wang, Boxiang Yun, Zhenyu Yi, Jianwei Xu, Wenqiao Zhang, Yingda Xia, Ling Zhang

机构 * Zhejiang University(浙江大学) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院) Hupan Lab(华平实验室) Huazhong University of Science and Technology(华中科技大学) East China Normal University(华东师范大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 提出跨视图对齐的证据驱动多模态强化学习框架E-MRL,通过将生成过程建模为“诊断-定位-验证”的马尔可夫决策过程,并引入跨视图一致性奖励,减少视觉幻觉并提升3D CT肿瘤诊断准确性。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16696 2026-06-24 cs.LG cs.AI cs.MM cs.SD 版本更新 81%

FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation

FISHER:多模态工业信号综合表示的基础模型

Pingyi Fan, Anbai Jiang, Shuwei Zhang, Xinhu Zheng, Zhiqiang Lv, Bing Han, Wenrui Liang, Junjie Li, Wei-Qiang Zhang, Yanmin Qian, Xie Chen, Jia Liu

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Institute for Embodied Intelligence and Robotics, Tsinghua University(清华大学智能感知与机器人研究院) Department of Computer Science and Engineering, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系) Huakong AI Plus Company Limited(华冠AIplus有限公司) Didi International Business Group(滴滴国际商务集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI、cs.MM

AI总结 针对工业信号分析中的数据异质性(M5问题),提出FISHER基础模型,采用子带建模处理多采样率问题,通过教师-学生自蒸馏预训练,在19个数据集上以较小规模超越24个SOTA编码器。

Comments Accepted by IEEE TII. FISHER open-sourced on https://github.com/jianganbai/FISHER . RMIS open-sourced on https://jianganbai.github.io/RMIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24156 2026-06-24 cs.CV 新提交 79%

Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction

利用先验校正的令牌缩减加速多模态大语言模型

Zengjie Chen, Yuxiang Cai, Jingcai Guo, Taotao Cai, Jianwei Yin, Zhi Chen

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字化服务技术重点实验室) Hong Kong Polytechnic University(香港理工大学) The University of Southern Queensland(南昆士兰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 提出PriorTR方法,通过分离任务条件注意力与模型先验注意力,在单次前向传播中估计先验并校正令牌重要性,实现无训练令牌缩减,提升多模态大语言模型效率与准确性。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24092 2026-06-24 cs.CV 新提交 57%

Progressive Pixel-Neighborhood Deformable Cross-Attention for Multispectral Object Detection

渐进式像素邻域可变形交叉注意力用于多光谱目标检测

Tian Qiu, Jifeng Shen, Xin Zuo

机构 * School of Electrical and Information Engineering, Jiangsu University(江苏大学电气与信息工程学院) School of Computer Science and Engineering, Jiangsu University of Science and Technology(江苏科技大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 提出PNAFusion,通过像素邻域交叉注意力和自适应可变形对齐模块,渐进式优化多光谱特征融合,在多个数据集上达到高精度并降低计算开销。

Comments Accepted by Sensors

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23724 2026-06-24 cs.IR cs.CL cs.HC 新提交 57%

EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering

EvidenceLens: 用于审计金融问答的声明-证据矩阵

Fengchen Gu, Xiaotian Ren, Zhengyong Jiang, Zhilu Zhang, Ángel F. García-Fernández, Angelos Stefanidis, Mian Zhou, Huakang Li, Jionglong Su

机构 * 1 School of AI Advanced Computing, XJTLU Entrepreneur College (Taicang), Xi’an Jiaotong-Liverpool University, Suzhou, Jiangsu, China 2 ETSI de Telecomunicaci\' o n, Universidad Polit\' e cnica de Madrid, Madrid, Spain IEEE VIS 2026 conditionally accepted version

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 提出EvidenceLens,一种将金融问答视为声明-证据对齐问题的可视化分析工具,通过多模态声明-证据矩阵揭示覆盖、矛盾与模态不平衡,帮助分析师区分有依据的声明与过度自信的综合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21690 2026-06-24 cs.CR cs.CL cs.LG 新提交 57%

A Hybrid, Multi-Layered Pipeline for Phishing and Threat Classification: Independently Validated URL and NLP Engines with a Calibrated Multi-Channel Fusion Stage

用于钓鱼和威胁分类的混合多层管道:独立验证的URL和NLP引擎与校准的多通道融合阶段

Saifelden M. Ismail, Aser O. Ibrahim, Omar A. Mahmoud

机构 * Department of Communications and Information Engineering(通讯与信息工程系) Zewail City of Science and Technology(泽瓦尔科学与技术城)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

AI总结 提出一种混合管道,分别对URL和文本模态使用独立引擎评分并融合结果,在10,677封邮件基准测试中达到F1=0.914,并将真实垃圾邮件误报率降至3.6%。

Comments Graduation project, Zewail City of Science and Technology. Code and documentation: https://github.com/XHCFS/cybersiren. Whole-system fusion results use proxy URL and header channels; treat integrated metrics as preliminary

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01554 2026-06-24 cs.RO 50%

L2M-Calib: One-key Calibration Method for LiDAR and Multiple Magnetic Sensors

L2M-Calib:一种用于激光雷达和多种磁传感器的单键校准方法

Qiyang Lyu, Wei Wang, Zhenyu Wu, Hongming Shen, Huiqin Zhou, Danwei Wang

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出L2M-Calib方法,通过联合估计激光雷达与磁传感器间的外参和磁传感器的内畸变参数,实现多模态传感器融合的高效校准。

详情

展开后加载摘要…

URL PDF HTML 收藏