arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2606.04434 2026-06-04 cs.CV cs.LG 83%

Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context Learning

Hyper-ICL:基于双曲锚点蒸馏的注意力校准用于多模态上下文学习

Niloufar Alipour Talemi, Hossein Kashiani, Fatemeh Afghah

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 提出Hyper-ICL,一种轻量级训练框架,通过低秩logit适配器和双曲锚点蒸馏损失校准注意力分布,无需推理时提供上下文示例即可重建演示效果,提升多模态上下文学习的准确性和稳定性。

Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02576 2026-06-04 cs.CV cs.LG 83%

ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning

ProtoAda: 原型引导的自适应适配器扩展与几何整合用于多模态持续指令微调

Yu-Cheng Shi, Zhen-Hao Xie, Jun-Tao Tang, Da-Wei Zhou

机构 * School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院) State Key Laboratory of Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 提出ProtoAda框架,通过格式感知任务原型和几何感知参数整合,解决多模态持续指令微调中任务路由错误和梯度干扰问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03604 2026-06-03 cs.CL 83%

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

超越字面:多模态模因理解中的语用意图分解

Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Luyao Ye, Huimin Wang, Hanqi Yan, Binyang Li, Kam-Fai Wong, Yulan He

机构 * The Chinese University of Hong Kong(香港中文大学) Huawei(华为) Central China Normal University(中央师范大学) Shenzhen University(深圳大学) King’s College London(伦敦国王学院) University of International Relations(国际关系大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

AI总结 针对大型视觉语言模型(LVLMs)在理解模因时倾向于描述字面内容而非语用意图的问题,提出Intent Projection框架,通过表示、输出和目标三层面的字面-语用分解,在六个基准上超越开源模型并缩小与专有模型的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01558 2026-06-02 cs.CV 83%

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning

注意力引导的多模态大语言模型微调提升思维链推理能力

Sanchit Sinha, Guangzhi Xiong, Bohan Liu, Zhenghao He, Aidong Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 针对多模态大语言模型中思维链推理效果不佳的问题,提出注意力引导的微调目标Attentive-CoT,通过延迟答案承诺和维持视觉令牌访问来提升推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15475 2026-06-01 cs.CV 83%

Multimodal Fusion via Self-Consistent Task-Gradient Fields

通过自洽任务梯度场的多模态融合

Jiayu Xiong, Jing Wang, Jun Xue, Wanlong Wang, Jianlong Kwan, Xiaosen Lyu, Zhouqiang Jiang

机构 * Xiamen Key Laboratory of Computer Vision and Pattern Recognition, Huaqiao University, Xiamen, Fujian, China(厦门计算机视觉与模式识别重点实验室,华侨大学,厦门,福建,中国) Huaqiao University, Xiamen, Fujian, China(华侨大学,厦门,福建,中国) Wuhan University, Wuhan, Hubei, China(武汉大学,武汉,湖北,中国) Nakashima Lab, SANKEN, The University of Osaka, Osaka, Japan(Nakashima实验室,SANKEN,大阪大学,大阪,日本)

专题命中 多模态训练与对齐 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

AI总结 提出自洽场自编码器(SCFAE),利用自洽场原理平衡任务学习与特征组织,通过任务损失和重构损失在互补子空间中分离特征,从而鲁棒处理缺失数据和不均匀输入。

Comments ICML 2026 accepted paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28575 2026-05-28 cs.AI 83%

A Conflict-Aware Penalty and Statistical Loss Framework for Balancing Modalities and Enhancing Stability in Multimodal Sentiment Analysis

一种冲突感知惩罚与统计损失框架,用于平衡模态并增强多模态情感分析的稳定性

Jianheng Dai, Jiazhang Liang, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 针对多模态情感分析中文本模态主导导致梯度冲突的问题,提出冲突感知惩罚和统计损失框架,实现模态平衡与训练稳定,在CMU-MOSI上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00682 2026-05-26 cs.IR cs.AI 83%

RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with Dual Semantic Alignment

RecGOAT: 用于LLM增强多模态推荐的图最优自适应传输与双语义对齐

Yuecheng Li, Hengwei Ju, Zeyu Song, Wei Yang, Chi Lu, Peng Jiang, Kun Gai

机构 * Fudan University(复旦大学) University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 针对生成式语言模型表示与ID协同信号之间的语义异质性,提出基于图神经网络和最优传输理论的双粒度语义对齐框架RecGOAT,通过实例级和分布级对齐实现统一特征空间,理论证明其表示误差更低,实验达到最优性能。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21454 2026-05-21 cs.CV q-bio.QM q-bio.TO 83%

ProtoPathway: Biologically Structured Prototype-Pathway Fusion for Multimodal Cancer Survival Prediction

ProtoPathway: 为多模态癌症生存预测设计的生物结构化原型-路径融合

Amaya Gallagher-Syed, Costantino Pitzalis, Myles J. Lewis, Michael R. Barnes, Gregory Slabaugh

机构 * Queen Mary University of London(伦敦女王学院) Imperial College London(帝国理工学院伦敦分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出ProtoPathway框架,通过统一全切片成像和转录组学,利用编码器生成生物基础的表示,以提升癌症生存预测的生物可解释性和计算效率。

Comments Currently under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21059 2026-05-21 cs.CV cs.LG 83%

Multimodal LLMs under Pairwise Modalities

基于成对模态的多模态大语言模型

Yan Li, Yunlong Deng, Yuewen Sun, Gongxu Luo, Kun Zhang, Guangyi Chen

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种基于成对模态训练多模态大语言模型的方法,通过理论分析和表示学习框架,实现了跨模态对齐和重构,提升了模型的跨模态性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18988 2026-05-20 cs.CR cs.AI 83%

Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks

在不可见中存活:面向新颖多轮多模态攻击的预测防御

Doohee You

机构 * Trust and Safety, Google(谷歌信任与安全部)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种预测性防御方法,用于应对新颖多轮多模态攻击,通过动态生存预测和轨迹动态问题来解决静态防御机制的不足,建立了一个计算高效且可解释的安全保障框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17766 2026-05-19 cs.CV 83%

LatentUMM: Dual Latent Alignment for Unified Multimodal Models

LatentUMM: 双重潜在对齐用于统一多模态模型

Yinyi Luo, Wenwen Wang, Hayes Bai, Marios Savvides, Jindong Wang

机构 * Carnegie Mellon University(卡内基梅隆大学) William & Mary(威廉与玛丽学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出LatentUMM,通过构建增强的共享潜在空间,显式对齐映射到和从潜在空间的转换,提高跨模态一致性。实验表明,该方法在多种架构上一致提升了多模态一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16889 2026-05-19 cs.CV 83%

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

通过缺失模态控制决策漂移的多模态情感分析

Chenglizhao Chen, Yuchen Cao, Xinyu Liu, Mengke Song, Guisheng Zhang, Xiaomin Yu

机构 * Qingdao Institute of Software, College of Computer Science and Technology, China University of Petroleum (East China)(青岛软件研究所,计算机科学与技术学院,中国石油大学(华东)) Shandong Key Laboratory of Intelligent Oil & Gas Industrial Software(山东省智能油气工业软件重点实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种两级参考对齐框架,旨在解决多模态情感分析中因缺失模态和质量不平衡导致的决策漂移问题,通过稳定参考提升鲁棒性,实验表明在不同缺失模态设置下方法有效,且在全模态输入下达到最先进的性能。

Comments Accepted by IJCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14654 2026-05-15 cs.CV 83%

Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

超越实例级自监督在3D多模态医学影像中的应用

Tan Pan, Shuhao Mei, Yixuan Sun, Kaiyu Guo, Chen Jiang, Zhaorui Tan, Mengzhu Li, Limei Han, Xiang Zou, Yuan Cheng, Mahsa Baktashmotlagh

机构 * Fudan University, China(复旦大学) University of Queensland, Australia(昆士兰大学) Shanghai Academy of AI for Science, China(上海人工智能科学研究院) Huashan Hospital, National Center for Neurological Disorders, Fudan University, China(华山医院,国家神经系统疾病中心,复旦大学) Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore(生物信息研究所(BII),科技研究局(A*STAR),新加坡)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出利用跨实例拓扑一致性作为监督信号,通过实例内和实例间对齐机制提升多模态医学影像的表征学习,实验显示在7个下游任务中实现了分割和分类任务的平均提升,并在模态缺失时表现出更强的鲁棒性。

Comments ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14458 2026-05-15 cs.AI 83%

OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance

OmniDrop:通过查询引导的分层令牌剪枝实现多模态大语言模型

Yeo Jeong Park, Hyemi Jang, Minseo Choi, Jongsun Lee, Jooyoung Choi, Yongkweon Jeon

机构 * Samsung Research(三星研究院)

专题命中 多模态训练与对齐 :omni-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 本文提出OmniDrop,通过查询引导的分层令牌剪枝方法,提升多模态大语言模型的效率与性能,实验表明其在多个音频视频基准测试中表现优异,减少预填充延迟和内存使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05008 2026-05-15 cs.CV 83%

Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

多模态因果驱动表示学习用于通用化医学图像分割

Xusheng Liang, Lihua Zhou, Nianxin Li, Miao Xu, Ziyang Song, Dong Yi, Jinlin Wu, Jiawei Ma, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * City University of Hong Kong(香港城市大学) Shenzhen Loop Area Institute(深圳河套学院) CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算智能研究所) UESTC(电子科技大学) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MCDRL框架,结合因果推理与VLM解决医学图像分割的领域泛化问题,通过文本提示识别病变区域并消除领域特定影响,提升分割准确性与泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13798 2026-05-14 cs.CV 83%

VoxCor: Training-Free Volumetric Features for Multimodal Voxel Correspondence

VoxCor:无需训练的体积分量用于多模态体素对应

Guney Tombak, Ertunc Erdil, Ender Konukoglu

机构 * Biomedical Image Computing Group, ETH Zurich(生物医学图像计算组,苏黎世联邦理工学院) The LOOP Zurich – Medical Research Center(苏黎世医疗研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 VoxCor通过冻结的2D Vision Transformer基础模型生成可重用的体积分量表示,无需训练即可实现跨模态和跨主体的体素对应,提升跨模态和跨主体的转换性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04804 2026-05-14 cs.CL 83%

OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

OmniSIFT: 多模态非对称令牌压缩用于高效的多模态大语言模型

Yue Ding, Yiyan Ji, Jungang Li, Xuyang Liu, Xinlong Chen, Junfei Wu, Bozhou Li, Bohan Zeng, Yang Shi, Yushuo Guan, Yuanxing Zhang, Jiaheng Liu, Qiang Liu, Pengfei Wan, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR)、自动化研究所、中国科学院(CASIA)) Nanjing University(南京大学) The Hong Kong University of Science(香港科学大学) Sichuan University(四川大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CL

AI总结 OmniSIFT通过非对称令牌压缩框架提升多模态大语言模型效率,采用时空视频修剪和视觉引导音频选择模块,实现低参数高鲁棒性。

Comments [ICML 2026] Code Link: https://github.com/dingyue772/OmniSIFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13059 2026-05-14 cs.CV 83%

BrainAnytime: Anatomy-Aware Cross-Modal Pretraining for Brain Image Analysis with Arbitrary Modality Availability

BrainAnytime: 基于解剖结构的跨模态预训练用于脑图像分析,支持任意模态可用性

Guangqian Yang, Tong Ding, Wenlong Hou, Yue Xun, Ye Du, Qian Niu, Shujun Wang

机构 * Department of Biomedical Engineering, The Hong Kong Polytechnic University, Hong Kong SAR, China(生物医学工程系,香港理工大学,香港特别行政区,中国) Department of Technology Management for Innovation, The University of Tokyo, Japan(创新技术管理系,东京大学,日本) Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hong Kong SAR, China(数据科学与人工智能系,香港理工大学,香港特别行政区,中国)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 BrainAnytime通过跨模态蒸馏和解剖引导课程掩码,在共享的3D掩码自动编码器中学习MRI与PET的结构-分子对应关系,实现对任意模态可用性的统一预训练,提升多任务性能。

Comments Early accepted by MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11444 2026-05-14 cs.CV 83%

Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts

通过混合频率专家利用多模态大语言模型实现一站式图像修复

Eunho Lee, Rei Kawakami, Youngbae Hwang

机构 * Chungbuk National University(Chungbuk国立大学) Institute of Science Tokyo(东京科学研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出一种利用多模态大语言模型指导的图像修复框架,通过混合频率专家模块提升对复合退化特征的建模能力,在多个数据集上取得优异效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11468 2026-05-13 cs.AI 83%

CAMPA: Efficient and Aligned Multimodal Graph Learning via Decoupled Propagation and Aggregation

CAMPA: 通过解耦传播和聚合实现高效且对齐的多模态图学习

Daohan Su, Hao Liu, Xunkai Li, Yinlin Zhu, Xiong Yongfu, Yi Liu, Hongchao Qin, Rong-Hua Li, Guoren Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 CAMPA通过解耦传播和聚合机制,解决多模态图学习中的模态冲突问题,提升效率和可扩展性,优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08592 2026-05-12 cs.CV 83%

Cross-Modal RGB-D Fusion Transformer for 6D Pose Estimation of Non-Cooperative Spacecraft with Stereo-Derived Depth

用于非合作航天器6D位姿估计的跨模态RGB-D融合Transformer

Yongliang Zhen, Bo LÜ, Hang Yang, Xiaotian WU

机构 * School of Physics, Northeast Normal University(东北师范大学物理学院) Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Science(长春光学精密机械与物理研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出基于被动立体视觉的6D位姿估计方法,通过TSCA-Stereo网络处理空间图像中的弱纹理和强光问题,并结合跨模态Transformer融合RGB和立体深度信息,实现高精度位姿估计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12286 2026-05-12 q-bio.GN cs.CL 83%

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer

不再有间隙:通过一个分词器实现零间隙多模态整合

Yanan Li, Christina Yi Jin, Yuan Jin, Manli Luo, Tie Xu, Shuai Jiao, Wei He, Qing Zhang

机构 * Research Center for Frontier Fundamental Studies, Zhejiang Lab(前沿基础研究研究中心,浙江实验室) Research Center for Scientific Data Hub, Zhejiang Lab(科学数据枢纽研究中心,浙江实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出One Tokenizer,通过统一词汇实现多模态无缝整合,克服传统架构的几何模态间隙问题,提升生物推理性能。

Comments Under review at NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02991 2026-05-11 cs.CV 83%

GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection

GraphFusion3D: 动态图注意力卷积与自适应跨模态Transformer用于3D目标检测

Md Sohag Mia, Md Nahid Hasan, Muhammad Abdullah Adnan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出GraphFusion3D框架,结合多模态融合与先进特征学习,通过自适应跨模态Transformer增强几何与语义信息,并引入图推理模块捕捉局部结构与全局语义,实验在SUN RGB-D和ScanNetV2上均取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07356 2026-05-11 cs.CV 83%

UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition

UniD-Shift:通过可解释的共享-私有多模态分解实现统一的语义分割

Shuai Zhang, Zhecheng Shi, Zhuxiao Li, Jing Ou, Tengxi Wang, Yuan Liu, Wufan Zhao

机构 * HKUST(GZ)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出UniD-Shift框架,通过结合视觉和几何编码器提取互补特征,并通过共享-私有子空间分解实现2D-3D语义分割的统一。实验表明在SemanticKITTI和nuScenes数据集上优于基线方法,且具有良好的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06238 2026-05-08 cs.LG cs.AI 83%

Band Together: Untargeted Adversarial Training with Multimodal Coordination against Evasion-based Promotion Attacks

Band Together: 多模态协调的无目标对抗训练对抗基于逃避的推广攻击

Guanmeng Xian, Ning Yang, Philip S. Yu

机构 * Sichuan University(四川大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出UAT-MC方法,通过多模态协调解决逃避攻击中未知目标项的问题,提升系统鲁棒性并保持推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05694 2026-05-08 cs.CV 83%

Adaptive Physical-Facial Representation Fusion via Subject-Invariant Cross-Modal Prompt Tuning for Video-Based Emotion Recognition

基于主体不变跨模态提示调优的自适应物理-面部表征融合用于基于视频的情感识别

Xiwen Luo, Jia Li, Rencheng Song, Yu Liu, Juan Cheng

机构 * Department of Biomedical Engineering(生物医学工程系) Anhui Province Key Laboratory of Measuring Theory and Precision Instrument(安徽省测量理论与精密仪器重点实验室) School of Computer Science and Information Engineering(计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出一种主体不变的跨模态提示调优框架,通过将rPPG波形转换为噪声鲁棒的时间-频率表示,并引入解耦共享-特定适配器以提升跨主体泛化能力,实验证明在MAHNOB-HCI和DEAP基准上优于现有方法。

Comments The source code will be available at https://github.com/MSA-LMC/SCPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00434 2026-05-04 cs.CV 83%

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

LIMSSR:基于大语言模型的训练时不完整多模态观察下的序列到评分推理

Huangbiao Xu, Huanqi Wu, Xiao Ke, Yuxin Peng

机构 * Fujian Provincial Key Laboratory of Networking Computing(网络计算与智能信息处理福建省重点实验室) Intelligent Information Processing, College of Computer(智能信息处理学院) Data Science, Fuzhou University(数据科学,福州大学) Engineering Research Center of Big Data Intelligence, Ministry of Education(大数据智能工程研究中心,教育部) Wangxuan Institute of Computer Technology, Peking University(王宣计算机技术研究所,北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出LIMSSR框架,通过提示引导的上下文感知模态补全和多维表示融合,解决训练时不完整多模态数据下的序列到评分推理问题,并在三个动作质量评估数据集上验证了其有效性。

Comments ICML 2026 [Spotlight]

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24125 2026-04-28 cs.CV 83%

Open-Vocabulary Semantic Segmentation Network Integrating Object-Level Label and Scene-Level Semantic Features for Multimodal Remote Sensing Images

开放词汇语义分割网络:整合对象级标签与场景级语义特征用于多模态遥感图像

Jinkun Dai, Yuanxin Ye, Peng Tang, Tengfeng Tang, Xianping Ma, Jing Xiao, Mi Wang

机构 * Faculty of Geosciences and Engineering, Southwest Jiaotong University, Chengdu 611756, China(西南交通大学地质工程学院,成都 611756,中国) State-Province Joint Engineering Laboratory of Spatial Information Technology for High-Speed Railway Safety, Southwest Jiaotong University, Chengdu 611756, China(高速铁路安全空间信息技术省部共建工程实验室,西南交通大学,成都 611756,中国) School of Artificial Intelligence, Wuhan University, Wuhan 430072, China(武汉大学人工智能学院,武汉 430072,中国) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430072, China(武汉大学测绘遥感信息工程国家重点实验室,武汉 430072,中国)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出TSMNet,通过整合文本监督与视觉表示,实现开放词汇语义分割,利用双分支文本编码器提取场景级语义和对象级标签信息,提升遥感图像分割的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22899 2026-04-28 cs.CV 83%

Text-Guided Multimodal Unified Industrial Anomaly Detection

基于文本引导的多模态统一工业异常检测

Zewen Li, Shuo Ye, Zitong Yu, Weicheng Xie, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(海湾大学计算机与信息学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出基于文本语义的多模态统一工业异常检测框架,解决跨模态对齐和几何建模不足问题,通过几何感知跨模态映射和对象条件文本特征适配器实现多模态特征对齐,实现跨类别的准确异常检测。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16377 2026-04-27 cs.CL cs.CY 83%

GoCoMA: Hyperbolic Multimodal Representation Fusion for Large Language Model-Generated Code Attribution

GoCoMA:超几何多模态表示融合用于大语言模型生成代码归因

Nitin Choudhury, Bikrant Bikram Pratap Maurya, Bhavinkumar Vinodbhai Kuwar, Arun Balaji Buduru

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 GoCoMA通过超几何空间融合代码风格和二进制特征,提升大语言模型生成代码的归因准确性,在两个基准测试中优于单一模态和欧几里得多模态基线。

Comments Accepted to the International Conference on Multimedia & Expo (ICME) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏