arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2603.12845 2026-04-24 cs.CV 83%

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

多模态蛋白质语言模型用于酶动力学参数:从底物识别到构象适应

Fei Wang, Xinye Zheng, Kun Li, Yanyan Wei, Yuxin Liu, Ganpeng Hu, Tong Bao, Jingwen Yang

机构 * School of Computer Science and Information Engineering(计算机科学与信息工程学院) Institute of Artificial Intelligence(人工智能研究院) CVLab, College of Information Technology(CV实验室,信息学院) Intelligent Interconnected Systems Laboratory of Anhui Province(安徽省智能互联系统实验室) School of Food Biological Engineering(食品生物工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出多阶段多模态条件建模方法,通过ERBA模块在蛋白质语言模型中注入跨模态信息,提升酶动力学参数预测的准确性与生物合理性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19379 2026-04-22 cs.CV 83%

PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

PanDA: 无监督领域自适应用于自动驾驶中的多模态3D全景分割

Yining Pan, Shijie Li, Yuchen Wu, Xulei Yang, Na Zhao

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(新加坡资讯通信研究院(I2R),A*STAR,新加坡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出PanDA框架,针对多模态3D全景分割的无监督领域自适应问题,通过不对称多模态增强和双专家伪标签细化模块提升鲁棒性和伪标签完整性,实验表明在多种领域转移场景下超越现有SOTA方法。

Comments Accepted at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19083 2026-04-22 cs.CR cs.AI 83%

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

ProjLens: 揭示项目器在多模态模型安全中的作用

Kun Wang, Cheng Qian, Miao Yu, Lilan Peng, Liang Lin, Jiaming Zhang, Tianyu Zhang, Yu Cheng, Yang Wang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 ProjLens通过分析多模态大语言模型中的后门攻击机制,揭示了项目器在安全漏洞中的关键作用,发现后门注入参数编码于低秩子空间,并通过实验验证了激活机制的差异。

Comments 18 pages ,15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12518 2026-04-21 cs.CL 83%

Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment Analysis

增强后再平衡模态协作用于鲁棒多模态情感分析

Kang He, Yuzhe Ding, Xinrong Wang, Fei Li, Chong Teng, Donghong Ji

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航天信息安全部门与可信计算教育部重点实验室,网络安全科学与工程学院,武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出EBMC框架,通过语义解耦和跨模态增强提升表示质量,结合能量引导模态协调机制和实例感知模态信任蒸馏,解决多模态情感分析中的模态不平衡问题,实验显示其在缺失模态情况下表现优异。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16943 2026-04-21 cs.CL 83%

MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation

MNAFT:多模态大语言模型的模态神经感知微调用于图像翻译

Bo Li, Ningyuan Deng, Tianyu Dong, Shaobo Wang, Shaolin Zhu, Lijie Wen

机构 * School of Computer Science and Technology, Tianjin University, Tianjin, China(天津大学计算机科学与技术学院) School of Software, Tsinghua University, Beijing, China(清华大学软件学院) School of Information Resource Management, Renmin University of China,Beijing, China(中国人民大学信息资源管理学院) School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China(上海交通大学人工智能学院) Baidu Inc., Beijing, China(百度公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出MNAFT,通过识别视觉和语言模块中语言无关和语言特定的神经元,改进多模态大语言模型在图像翻译中的表现,实验表明其优于现有方法。

Comments Accepted by SCIS (SCIENCE CHINA Information Science)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23714 2026-04-20 cs.CL 83%

Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph Completion

融合与独立的协作:超复数驱动的鲁棒多模态知识图谱补全

Zhiqiang Liu, Yichi Zhang, Mengshu Sun, Lei Liang, Wen Zhang

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ZJU-Ant Group Joint Lab of Knowledge Graph(浙大蚂蚁集团知识图谱联合实验室) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出M-Hyper方法,结合融合与独立模态表示,利用四元数代数实现多模态交互,提升知识图谱补全的鲁棒性和效率。

Comments ACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17417 2026-04-17 cs.LG cs.AI 83%

Generative Modeling of Class Probability for Multi-Modal Representation Learning

多模态表示学习中的类别概率生成建模

Jungkyoo Shin, Bumsoo Kim, Eunwoo Kim

机构 * Department of AI(人工智能系) School of CSE(计算机科学与工程学院) Chung-Ang University(Chung-Ang 大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种基于类别概率分布的多模态表示学习方法CALM,通过生成和对齐类别概率分布提升多模态对齐效果,并引入跨模态概率变分自编码器以增强模态间关系的建模能力。

Comments To appear in CVPR 2025 (Highlight)

Journal ref Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025, pp. 20737-20746

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14016 2026-04-16 cs.LG cs.AI 83%

MAny: Merge Anything for Multimodal Continual Instruction Tuning

MAny:为多模态连续指令微调合并任何内容

Zijian Gao, Wangwang Jia, Xingxing Zhang, Pengfei Qian, Tao Sun, Bo Ding, Yong Dou, Huaimin Wang, Kele Xu

机构 * College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院) National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology(国防科技大学并行与分布式计算国家重点实验室) State Key Laboratory of Complex & Critical Software Environment(复杂与关键软件环境国家重点实验室) School of Computer Science, Tsinghua University(清华大学计算机科学与技术系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 针对多模态连续指令微调中感知漂移和推理崩溃问题,提出MAny框架,通过跨模态投影合并和低秩参数合并实现任务知识融合,提升模型鲁棒性和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13403 2026-04-16 cs.CV 83%

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

为何多模态上下文学习滞后?揭示内部机制和瓶颈

Yu Wang, Sharon Li

机构 * Department of Computer Sciences, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文分析了多模态上下文学习在零样本设置中表现与少样本演示下退化的原因,揭示了模型在视觉与文本表示对齐和任务映射转移方面的不足,并提出改进方法。

Comments ACL Main 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12319 2026-04-16 cs.CV 83%

RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation

RSGMamba:面向多模态语义分割的可靠性感知自门控状态空间模型

Guoan Xu, Yang Xiao, Guangwei Gao, Dongchen Zhu, Guo-Jun Qi, Wenjing Jia

机构 * Faculty of Engineering and Information Technology, University of Technology Sydney(工程与信息技术学院,悉尼技术大学) PCA Lab, Key Laboratory of Intelligent Perception and Systems for High-Dimensional Information of Ministry of Education, School of Computer Science and Engineering, Nanjing University of Science and Technology(教育部高维信息智能感知与系统重点实验室,南京理工大学计算机科学与工程学院) Bionic Vision Systems Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences(中国科学院上海微系统与信息技术研究所生物视觉系统实验室) Research Center for Industries of the Future and the School of Engineering, Westlake University(未来产业研究中心和工程学院,西湖大学) OPPO Research, Seattle, WA 98101 USA(OPPO研究,美国华盛顿州西雅图98101)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出RSGMamba框架,通过可靠性感知自门控机制提升多模态语义分割性能,实验表明其在RGB-D和RGB-T基准上取得最优结果,参数量仅48.6M。

Comments 7tables,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12380 2026-04-15 cs.CV 83%

Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection

多模态伪装物体检测的模态无关提示学习

Hao Wang, Jiqing Zhang, Xin Yang, Baocai Yin, Lu Jiang, Zetian Mi, Huibing Wang

机构 * Information Science and Technology College, Dalian Maritime University(大连海事大学信息科学与技术学院) Beijing University of Technology(北京理工大学) Dalian University of Technology(大连理工大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种模态无关的提示学习框架,用于多模态伪装物体检测,通过数据驱动内容域与知识驱动提示域的交互,生成统一提示以提升SAM模型性能,并引入轻量级Mask Refine模块以提高检测精度。

Comments 10

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07812 2026-04-10 cs.CV 83%

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

HAWK:多模态模型中的头部重要性感知视觉标记修剪

Qihui Zhu, Tao Zhang, Yuchen Wang, Zijian Wen, Mengjie Zhang, Shuangwu Chen, Xiaobin Tan, Jian Yang, Yang Liu, Zhenhua Dong, Xianzhi Yu, Yinfei Pan

机构 * University of Science and Technology of China(中国科学技术大学) ChangXin Memory Technologies, Inc(长鑫存储技术有限公司) Huawei Noah’s Ark Lab(华为诺亚方舟实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 HAWK通过识别不同注意力头在视觉任务中的重要性,有效保留关键视觉标记,提升多模态模型的推理效率与性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01831 2026-04-08 cs.LG cs.AI 83%

Routing-Based Continual Learning for Multimodal Large Language Models

基于路由的多模态大语言模型持续学习

Jay Mohta, Kenan Emir Ak, Gwang Lee, Dimitrios Dimitriadis, Yan Xu, Mingwei Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种基于路由的持续学习方法,用于多模态大语言模型,有效缓解灾难性遗忘并提升跨模态迁移性能,同时保持训练效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04297 2026-04-07 cs.AI 83%

PanLUNA: An Efficient and Robust Query-Unified Multimodal Model for Edge Biosignal Intelligence

PanLUNA: 一种高效且稳健的多模态模型,用于边缘生物信号智能

Marija Zelic, Anna Tegon, Yawei Li, Thorir Mar Ingolfsson, Luca Benini

机构 * DEI, University of Bologna(博洛尼亚大学DEI)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 PanLUNA是一种紧凑的多模态基础模型,能够高效处理EEG、ECG和PPG信号,实现跨模态早期融合,同时在推理时对缺失模态具有鲁棒性。其在生物信号检测和睡眠分期任务中表现出色。

Comments 5 pages, 5 tables, 1 figure, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20146 2026-04-02 cs.CV 83%

VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection

VMAD:视觉增强的多模态大语言模型用于零样本异常检测

Huilin Deng, Hongchen Luo, Wei Zhai, Yang Cao, Yu Kang

机构 * University of Science and Technology of China(中国科学技术大学) Northeastern University(东北大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 VMAD通过结合视觉信息与多模态大语言模型,提升零样本异常检测的精度与分析能力,提出缺陷敏感结构学习和局部增强令牌压缩等方法,结合RIAD数据集验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27993 2026-03-31 cs.CV 83%

Progressive Prompt-Guided Cross-Modal Reasoning for Referring Image Segmentation

逐步引导的跨模态推理用于指代图像分割

Jiachen Li, Hongyun Wang, Jinyu Xu, Wenbo Jiang, Yanchun Ma, Yongjian Liu, Qing Xie, Bolong Zheng

机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) University of Electronic Science and Technology of China(电子科技大学) Wuhan Vocational College of Software and Engineering(武汉软件工程职业学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出PPCR框架,通过语义理解-空间定位-实例分割流程,改进指代图像分割中语言描述与视觉表示的连接,提升分割精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27321 2026-03-31 cs.LG cs.AI 83%

Multimodal Forecasting for Commodity Prices Using Spectrogram-Based and Time Series Representations

基于频谱和时间序列表示的商品价格多模态预测

Soyeon Park, Doohee Chung, Charmgil Hong

机构 * Impactive AI

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出SEMF方法,结合频谱和时间序列表示,提升多变量时间序列预测的准确性与鲁棒性,通过多模态融合和频谱编码在多个商品价格预测任务中优于七种基线模型。

Comments AAAI 2026 Summer Symposium Series; 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22674 2026-03-31 cs.CV 83%

VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration

VisionTrim: 一种用于无训练多模态大语言模型加速的统一视觉标记压缩方法

Hanxun Yu, Wentong Li, Xuan Qu, Song Wang, Junbo Chen, Jianke Zhu

机构 * State Key Lab of CAD & CG, Zhejiang University(浙江大学CAD&CG国家重点实验室) NUAA(南京航空航天大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出VisionTrim,通过整合DVTS和TGVC模块,有效减少视觉标记,提升多模态大语言模型在高分辨率和视频场景中的计算效率。

Comments ICLR2026, Code Link: https://github.com/hanxunyu/VisionTrim

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17528 2026-03-30 cs.CV 83%

MM-OVSeg:Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing

MM-OVSeg:多模态光学-合成孔径雷达融合用于遥感中的开放词汇分割

Yimin Wei, Aoran Xiao, Hongruixuan Chen, Junshi Xia, Naoto Yokoya

机构 * The University of Tokyo(东京大学) RIKEN AIP(日本理化学研究所先进智能研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MM-OVSeg通过融合光学和SAR数据,提升恶劣天气下的开放词汇分割鲁棒性与泛化能力,采用跨模态统一过程和双编码器融合模块实现多传感器表征对齐与多模态分割。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25555 2026-03-27 cs.CV 83%

Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion

通过多模态图像融合实现眼科手术中的全面实时场景理解

Nikolo Rohrmoser, Ghazal Ghazaei, Michael Sommersperger, Nassir Navab

机构 * Technical University of Munich(慕尼黑工业大学) Carl Zeiss AG(卡尔蔡司公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种多模态实时网络架构,用于联合仪器检测、关键点定位和工具-组织距离估计,通过多模态图像融合提升手术场景理解的准确性和实时性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25077 2026-03-27 cs.CV 83%

Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs

弥合感知与推理:用于多模态大语言模型中RLVR的标记重加权

Jinda Lu, Junkang Wu, Jinghan Li, Kexin Huang, Shuo Yang, Guoyin Wang, Jiancan Wu, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) Peking University(北京大学) Independent Researcher(独立研究员)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出Token-Reweighting策略,通过动态重加权多模态大语言模型中的感知与推理标记,提升RLVR性能,实现视觉 grounding 和推理的协同优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11380 2026-03-26 cs.CV 83%

DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding

DriveXQA: 多模态视觉问答用于恶劣驾驶场景理解

Mingzhe Tao, Ruiping Liu, Junwei Zheng, Yufan Chen, Kedi Ying, M. Saquib Sarfraz, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Hunan University(湖南大学) Mercedes-Benz Tech Innovation(梅赛德斯-奔驰技术创新)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出DriveXQA多模态数据集,用于自主驾驶场景中的视觉问答,通过融合多传感器信息提升对异常驾驶场景的理解能力。

Comments Accepted to CVPR DriveX Workshop. Dataset and Code: https://github.com/jtjmd/DRIVEXQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23276 2026-03-25 cs.CV 83%

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

CCF:互补协作融合用于领域泛化的多模态3D目标检测

Yuchen Wu, Kun Wang, Yining Pan, Na Zhao

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出CCF方法,通过查询解耦损失、LiDAR引导深度先验和互补跨模态掩码,提升多模态3D目标检测在跨领域场景下的鲁棒性,实验表明优于现有方法且保持源域性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22852 2026-03-25 cs.CV 83%

Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy Prediction

Gau-Occ:用于多模态3D占用预测的几何完备高斯分布

Chengxin Lv, Yihui Li, Hongyu Yang, YunHong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,北京航空航天大学,北京,中国) School of Computer Science and Engineering, Beihang University, Beijing, China(计算机科学与工程学院,北京航空航天大学,北京,中国) School of Artificial Intelligence, Beihang University, Beijing, China(人工智能学院,北京航空航天大学,北京,中国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 Gau-Occ通过几何完备的3D高斯分布实现多模态3D占用预测,利用LiDAR完成扩散器恢复缺失结构并融合多视角图像语义,提升空间一致性和语义区分性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21584 2026-03-24 cs.LG cs.CV 83%

SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models

SSAM:奇异子空间对齐用于融合多模态大语言模型

Md Kaykobad Reza, Ameya Patil, Edward Ayrapetian, M. Salman Asif

机构 * University of California Riverside(加州大学河滨分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SSAM通过参数空间对齐融合多模态大语言模型,无需训练数据实现跨模态统一,提升性能并降低资源消耗。

Comments 25 Pages, 9 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14188 2026-03-24 cs.CV 83%

Joint Segmentation and Grading with Iterative Optimization for Multimodal Glaucoma Diagnosis

多模态青光眼诊断的联合分割与分级迭代优化方法

Zhiwei Wang, Yuxing Li, Meilu Zhu, Defeng He, Edmund Y. Lam

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong, China(香港大学电子与电气工程系) College of Information Engineering, Zhejiang University of Technology, Hangzhou, China(浙江工业大学信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种迭代多模态优化模型,通过中层融合策略整合眼底和OCT特征,并利用跨模态特征对齐模块减少模态差异,实现青光眼的精确分割与分级。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03521 2026-03-24 cs.MM cs.LG 83%

Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation

跨空间协同:一种用于对话中多模态情感识别的统一框架

Xiaosen Lyu, Jiayu Xiong, Yuren Chen, Wanlong Wang, Xiaoqing Dai, Jing Wang

机构 * Xiaosen Lyu 1,2(李绍森 1,2) Jiayu Xiong 1,2(熊佳宇 1,2) Yuren Chen 1,2(陈远人 1,2) Wanlong Wang 1,2(王万龙 1,2) Xiaoqing Dai 1,2(戴晓青 1,2) Jing Wang 1,2(王婧 1,2)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出Cross-Space Synergy框架,通过协同多项式融合和帕累托梯度调节器有效提升多模态情感识别的准确性和训练稳定性。

Comments Accepted to AAAI 2026

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(29), 24226-24234 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22862 2026-03-24 cs.LG cs.CV 83%

Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time Adaptation

通过渐进重对齐桥接模态以实现多模态测试时适应

Jiacheng Li, Songhe Feng

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出BriMPR框架,通过分治策略解决多模态测试时适应中的模态间分布偏移和语义对齐问题,通过提示调优和跨模态对比学习提升多模态特征对齐效果。

Comments Accepted by AAAI 2026 (Oral)

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence. 2026, 40(27): 22931-22939

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19623 2026-03-23 cs.CV 83%

Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglement

解耦后再对齐:通过跨尺度特征解耦实现非迭代混合多模态图像配准

Chunlei Zhang, Jiahao Xia, Yun Xiao, Bo Jiang, Jian Zhang

机构 * Faculty of Engineering and IT, University of Technology Sydney(新南威尔士大学工程与信息技术学院) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出HRNet网络,通过解耦表示与混合参数预测,解决多模态图像配准中共享空间不稳定和单类型变换限制的问题,实现非迭代的粗到细配准。

Comments Accepted by CVPR 2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17885 2026-03-23 cs.CV cs.LG 83%

FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning

FastMMoE:通过动态专家激活和路由感知的标记剪枝加速多模态大语言模型

Guoyang Xia, Yifeng Ding, Fengfa Li, Lei Ren, Wei Chen, Fangxiang Feng, Xiaojie Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Li Auto(利亚自动化)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出FastMMoE,一种无需训练的加速框架,通过动态专家激活和路由感知标记剪枝,显著降低计算量并保持性能,优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏