arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2603.04890 2026-03-06 cs.LG cs.AI cs.CV 84%

FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation

FedAFD: 通过对抗融合与蒸馏实现多模态联邦学习

Min Tan, Junchao Ma, Yinfu Feng, Jiajun Ding, Wenwen Pan, Tingting Han, Qian Zheng, Zhenzhong Kuang, Zhou Yu

机构 * Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(浙江空间信息感知与传输重点实验室,杭州电子大学) Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University(复杂系统建模与仿真实验室,计算机科学与技术学院,杭州电子大学) Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 FedAFD通过对抗融合与蒸馏方法,解决多模态联邦学习中的模态差异、任务差异和模型异质性问题,提升客户端与服务器的学习性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16479 2026-03-03 eess.IV cs.AI cs.CV 84%

Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization

解耦的多模态学习:组织学与转录组学用于癌症表征

Yupei Zhang, Xiaofei Wang, Anran Liu, Lequan Yu, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Health Technology & Informatics, The Hong Kong Polytechnic University(香港理工大学健康科技与信息学系) Department of Statistics and Actuarial Science, The University of Hong Kong(香港大学统计与精算科学系) Department of Clinical Neurosciences and Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学临床神经科学系和应用数学与理论物理系;邓迪大学科学与工程学院和医学院) School of Science and Engineering and School of Medicine, University of Dundee, UK

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了解耦的多模态学习框架,通过分解组织学和转录组数据以提高癌症表征的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23300 2026-02-27 cs.CL eess.AS 84%

A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations

一种用于对话中多模态情绪识别的专家混合模型

Soumya Dutta, Smruthi Balaji, Sriram Ganapathy

机构 * LEAP Lab, Department of Electrical Engineering(LEAP实验室,电气工程系) Microsoft(微软)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、eess.AS

AI总结 MiSTER-E通过专家混合框架提升对话中多模态情绪识别的准确率,实现跨模态一致性与融合。

Comments Accepted to Elsevier Computer Speech and Language. 30 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15226 2026-02-27 cs.MM cs.CL 84%

Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models

并非所有注意力都是必需的:面向多模态大语言模型的参数和计算高效迁移学习

Qiong Wu, Weihao Ye, Yiyi Zhou, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China(教育部多媒体可信感知与高效计算重点实验室) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CL、cs.MM

AI总结 本文提出高效注意力跳过方法,通过减少冗余注意力计算提升多模态大语言模型的推理效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24072 2026-02-26 cs.CV cs.AI 84%

Uncovering Grounding IDs: How External Cues Shape Multimodal Binding

揭示地面ID:外部线索如何塑造多模态绑定

Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari, Mobin Bagherian, Sadegh Mohammadian, Mohammad Izadi, Mahdieh Soleymani Baghshah

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出地面ID概念,揭示外部线索通过增强多模态绑定的注意力机制,提升跨模态定位精度并减少幻觉。

Comments Under review as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19585 2026-02-24 cs.MM cs.AI 84%

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

三子空间解耦用于多模态情感分析

Chunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu, Rong Fu, Zhongxue Gan, Chun Ouyang

机构 * Fudan University(复旦大学) Peking University(北京大学) University of Macau(澳门大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

AI总结 本文提出三子空间解耦框架,通过分解多模态特征为公共、子模态共享和私人子空间,提升多模态情感分析的性能和鲁棒性。

Comments This study has been Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25889 2026-02-16 cs.CV cs.CL 84%

Multimodal LLM With Hierarchical Mixture-of-Experts for VQA on 3D Brain MRI

多模态大语言模型与分层混合专家用于3D脑部MRI的视觉问答

Arvind Murari Vepa, Yannan Yu, Jingru Gan, Anthony Cuturrufo, Michael F. Romano, Weikai Li, Fabien Scalzo, Wei Wang, Yizhou Sun

机构 * UCLA Computer Science(UCLA计算机科学系) UCSF Radiology(UCSF放射学) UCLA Medicine(UCLA医学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 本研究提出了一种多模态大语言模型mpLLM,用于3D脑部MRI的视觉问答,通过分层混合专家架构提升肿瘤描述的临床可解释性,并在多个数据集上优于现有基线模型。

Comments 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06351 2026-02-09 cs.AI cs.CV 84%

Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion

Trifuse: 通过多模态融合增强基于注意力的GUI定位

Longhui Ma, Di Zhao, Siwei Wang, Zhao Lv, Miao Wang

机构 * College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学) Intelligent Game and Decision Lab, Academy of Military Sciences(智能游戏与决策实验室,军事科学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 Trifuse通过多模态融合提升GUI定位性能,整合注意力、OCR文本和图标描述语义,无需任务微调即可实现高效定位。

Comments 17 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21843 2026-02-06 cs.CV cs.AI 84%

CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition

CMD-HAR:基于交叉模态解耦的可穿戴人类活动识别

Ying Yu, Siyao Li, Yixuan Jiang, Hang Xiao, Jingxi Long, Haotian Tang, Hanyu Liu, Chao Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 CMD-HAR通过交叉模态解耦和时空注意力机制,提升可穿戴设备中人类活动识别的准确性和部署效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05568 2026-02-03 cs.CL cs.AI cs.LG 84%

Large Multimodal Models for Low-Resource Languages: A Survey

大规模多模态模型在低资源语言中的应用:综述

Marian Lupascu, Ana-Cristina Rogoz, Mihai Sorin Stupariu, Radu Tudor Ionescu

机构 * Department of Computer Science, University of Bucharest(布加勒斯特大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了大规模多模态模型在低资源语言中的应用,分析了适应技术及挑战,提供了方法导向的贡献比较和开放源代码资源。

Comments Accepted in Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00505 2026-02-03 cs.CV cs.AI 84%

Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models

稀疏快捷键:促进多模态大语言模型中的高效融合

Jingrui Zhang, Feng Liang, Yong Zhang, Wei Wang, Runhao Zeng, Xiping Hu

机构 * Guangdong-Hong Kong-Macao Joint Laboratory for Emotion Intelligence and Pervasive Computing, Artificial Intelligence Research Institute, Shenzhen MSU-BIT University, Shenzhen(粤港澳大湾区情感智能与 pervasive 计算联合实验室,人工智能研究院,深圳MSU-BIT大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 SparseCut通过稀疏快捷连接和多粒度特征融合模块,提升多模态大语言模型在跨模态融合中的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00143 2026-02-03 cs.LG cs.AI cs.CV 84%

Invariant Representation Guided Multimodal Sentiment Decoding with Sequential Variation Regularization

不变表示引导的多模态情感解码与序列变化正则化

Guoyang Xu, Zhenxi Song, Junqi Xue, Yuxin Liu, Zirui Wang, Zhiguo Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳校区)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种通过模态不变融合和序列变化正则化相结合的方法,以提升多模态情感解码的稳定性和准确性。

Comments change Title, Authors, Abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19327 2026-01-27 cs.CV cs.AI 84%

DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding

DeepInsert: 早期层绕过以实现高效且高性能的多模态理解

Moulik Choraria, Xinbo Wu, Akhil Bhimaraju, Nitesh Sekhar, Yue Wu, Xu Zhang, Prateek Singhal, Lav R. Varshney

机构 * UIUC(伊利诺伊大学) Amazon(亚马逊) Capital One Apple(苹果) Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 DeepInsert通过将多模态令牌插入模型中间层以绕过早期层,从而在降低训练和推理成本的同时保持或提升多模态语言模型的性能。

Comments To be presented at EACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15316 2026-01-23 cs.AI cs.CV 84%

The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection

范式转变:大型视觉语言模型在多模态虚假新闻检测中的全面调查

Wei Ai, Yilong Tan, Yuntao Shou, Tao Meng, Haowen Chen, Zhixiong He, Keqin Li

机构 * College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业科技学院) College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学) College of Economics and Management, Central South University of Forestry and Technology(经济管理学院,中央南大学林业科技学院) Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文通过大型视觉语言模型全面调查多模态虚假新闻检测,分析其发展历程、技术挑战及未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07178 2026-01-13 cs.CV cs.AI 84%

DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection

DIVER: 动态迭代视觉证据推理用于多模态虚假新闻检测

Weilin Zhou, Zonghao Ying, Chunlei Meng, Jiahui Liu, Hengyang Zhou, Quanchen Zou, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang

机构 * Xinjiang University(新疆大学) AI Security Lab(360人工智能安全实验室) Beihang University(北航) Fudan University(复旦大学) Central South University(中南大学) Nanjing University(南京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 DIVER通过动态迭代视觉证据推理提升多模态虚假新闻检测性能,利用文本和视觉证据融合优化检测效果。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02415 2026-01-07 cs.CV cs.AI 84%

Multimodal Sentiment Analysis based on Multi-channel and Symmetric Mutual Promotion Feature Fusion

基于多通道和对称互促特征融合的多模态情感分析

Wangyuan Zhu, Jun Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于多通道和对称互促特征融合的多模态情感分析方法,通过增强模态内特征表示和促进模态间信息交互,提升情感识别的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07307 2026-01-01 cs.CV cs.AI 84%

MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark

MCITlib: 多模态持续指令微调库与基准

Haiyang Guo, Fei Zhu, Hongbo Zhao, Fanhu Zeng, Wenzhuo Liu, Shijie Ma, Da-Han Wang, Xu-Yao Zhang

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新科学研究院人工智能与机器人中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Fujian Key Laboratory of Pattern Recognition and Image Understanding, School of Computer and Information Engineering, Xiamen University of Technology(福建 pattern recognition and image understanding 工程学院,厦门大学科技学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MCITlib提供多模态持续学习的库和基准,支持8种算法并评估3个基准,旨在解决灾难性遗忘和跨模态协调问题。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22605 2025-12-30 cs.AI cs.CV 84%

Learning Multi-Modal Mobility Dynamics for Generalized Next Location Recommendation

学习多模态移动动态以实现通用的下一站推荐

Junshu Dai, Yu Wang, Tongya Zheng, Wei Ji, Qinghong Guo, Ji Cao, Jie Song, Canghong Jin, Mingli Song

机构 * Zhejiang University(浙江大学) Hangzhou City University(杭州市大学) Nanjing University(南京大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出多模态移动(M^3ob)方法,通过构建统一时空关系图和门控机制,提升位置推荐任务的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12822 2025-12-16 cs.CV cs.AI 84%

Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding

Lemon:一种统一且可扩展的3D多模态模型用于通用空间理解

Yongyuan Liang, Xiyao Wang, Yuanchen Ju, Jianwei Yang, Furong Huang

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 Lemon提出一种统一的Transformer架构,通过联合处理3D点云块和语言标记,实现空间-语言融合,提升3D多模态模型的可扩展性和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13115 2025-12-15 eess.IV cs.AI cs.CV 84%

Multimodal Learning for Scalable Representation of High-Dimensional Medical Data

多模态学习用于高维医学数据的可扩展表示

Areej Alsaafin, Abubakr Shafique, Saghir Alfasly, Krishna R. Kalari, H. R. Tizhoosh

机构 * Kimia Lab, Dept. of Artificial Intelligence & Informatics, Mayo Clinic, Rochester, MN, USA(Kimia实验室,人工智能与信息学系,梅奥诊所,罗切斯特,MN,美国) Division of Computational Biology, Dept. of Quantitative Health Sciences, Mayo Clinic, Rochester, MN, USA(计算生物学部门,定量健康科学系,梅奥诊所,罗切斯特,MN,美国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MarbliX通过多模态学习实现高维医学数据的可扩展表示,提升癌症诊断的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07568 2025-12-09 cs.CV cs.AI eess.IV 84%

Dual-Stream Cross-Modal Representation Learning via Residual Semantic Decorrelation

通过残差语义去相关实现双流跨模态表示学习

Xuecheng Li, Weikuan Jia, Alisher Kurbonaliev, Qurbonaliev Alisher, Khudzhamkulov Rustam, Ismoilov Shuhratjon, Eshmatov Javhariddin, Yuanjie Zheng

机构 * School of Information Science & Engineering, Shandong Normal University(信息科学与工程学院,山东师范大学) Tajikistan State University of Law, Business Sughd(塔吉克斯坦法律、商业大学,苏赫德) Tajik State University of Law, Business and Politics Sughd(塔吉克斯坦法律、商业与政治大学,苏赫德)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 DSRSD-Net通过残差分解和语义去相关解决跨模态学习中的模态主导和冗余问题,提升预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03195 2025-12-01 cs.CV cs.AI cs.LG 84%

Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs

未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能

Yunqi Hong, Sohyun An, Andrew Bai, Neil Y. C. Lin, Cho-Jui Hsieh

机构 * Computer Science Department, University of California, Los Angeles(加州大学洛杉矶分校计算机科学系) Mechanical and Aerospace Engineering Department, University of California, Los Angeles(加州大学洛杉矶分校机械与航空航天工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 AutoSEP通过利用未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17585 2025-11-26 cs.LG cs.AI cs.CV 84%

PaSE: Prototype-aligned Calibration and Shapley-based Equilibrium for Multimodal Sentiment Analysis

PaSE:原型对齐校准与基于Shapley的均衡用于多模态情感分析

Kang He, Boyu Chen, Yuzhe Ding, Fei Li, Chong Teng, Donghong Ji

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 PaSE通过原型对齐校准和基于Shapley的均衡框架,有效缓解多模态情感分析中的模态竞争问题,提升模型性能。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18875 2025-11-25 cs.CV cs.MM 84%

Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference

并行视觉令牌调度用于快速准确的多模态大语言模型推理

Wengyi Zhan, Mingbao Lin, Zhihang Lin, Rongrong Ji

机构 * Xiamen University(厦门大学) Rakuten Asia Pte. Ltd.(拉面亚洲有限公司) Xiamen University, China(厦门大学) Shanghai Innovation Institute, China(上海创新研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.MM

AI总结 ParVTS通过并行调度视觉令牌,高效减少多模态大语言模型推理的计算复杂度,实现高剪枝率与性能损失小的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17927 2025-11-25 cs.CV cs.AI 84%

PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning

PA-FAS: 向可解释和通用的多模态人脸反伪装迈进:通过路径增强强化学习

Yingjie Ma, Xun Lin, Yong Xu, Weicheng Xie, Zitong Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 PA-FAS通过路径增强强化学习提升多模态人脸反伪装的可解释性和泛化能力。

Comments Accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17596 2025-11-25 cs.CV cs.AI 84%

Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding

基于重建的多模态表示学习用于自动化媒体理解

Yassir Benhammou, Suman Kalyan, Sujay Kumar

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出多模态自编码器MMAE,通过统一学习文本、音频和视觉数据的表示,实现广播内容的端到端自动化元数据提取与语义聚类,提升广播档案的可检索性与管理效率。

Comments 8 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10671 2025-11-17 cs.CL cs.CV 84%

Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency

Filippo Morbiato, Luca Romano, Alessandro Persona

机构 * University of Padua(帕多瓦大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07941 2025-11-12 cs.CV cs.AI 84%

Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification

Zhenfeng Zhuang, Fangyu Zhou, Liansheng Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11331 2025-10-31 cs.CL cs.MM 84%

Dependency Structure Augmented Contextual Scoping Framework for Multimodal Aspect-Based Sentiment Analysis

Hao Liu, Lijun He, Jiaxi Liang, Zhihan Ren, Haixia Bi, Fan Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21793 2025-10-28 cs.CV cs.AI eess.IV 84%

2D_3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection

Usman Ali, Ali Zia, Abdul Rehman, Umer Ramzan, Zohaib Hassan, Talha Sattar, Jing Wang, Wei Xiang

机构 * GIFT University(GIFT大学) La Trobe University(拉特罗布大学) Department of Primary Industries(初级产业部门)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at 26th International Conference on Digital Image Computing: Techniques and Applications (DICTA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏