arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2509.04870 2025-12-16 eess.IV cs.CV 83%

Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images

多模态不确定性鲁棒树冠覆盖分割用于高分辨率遥感图像

Yuanyuan Gui, Wei Li, Yinjian Wang, Xiang-Gen Xia, Mauro Marty, Christian Ginzler, Zuyuan Wang

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing(空间智能信息处理国家重点实验室) Department of Electrical and Computer Engineering, University of Delaware(电气与计算机工程系,德雷塞尔大学) Swiss Federal Institute for Forest, Snow, and Landscape Research WSL(瑞士森林、雪和景观研究联邦 institute WSL)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MURTreeFormer通过多模态分割框架降低不确定性,提升高分辨率遥感图像中树冠分割的鲁棒性与准确性。

Journal ref IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13261 2025-12-16 cs.CV 83%

Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning

解锁基于强化学习的多模态推理中难度先验的潜力

Mingrui Chen, Haogeng Liu, Hao Liang, Huaibo Huang, Wentao Zhang, Ran He

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) NLPR&MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) Zhongguancun Academy(中关村学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文通过建模问题难度先验信息,改进基于强化学习的多模态推理性能,通过数据筛选、优势分化和难度提示提升模型推理深度和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11901 2025-12-16 cs.CV cs.LG 83%

CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities

CLARGA:任意模态集合上的多模态图表示学习

Santosh Patapati

机构 * Santosh Patapati(独立研究者)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 CLARGA是一种通用的多模态融合架构,通过构建注意力加权图实现多模态表示学习,适用于任意模态集合,具有高效的融合能力和良好的鲁棒性。

Comments WACV; Supplementary material is available on CVF proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19312 2025-12-15 eess.SP cs.AI cs.IT cs.LG math.IT 83%

E2E Learning Massive MIMO for Multimodal Semantic Non-Orthogonal Transmission and Fusion

端到端学习大规模MIMO用于多模态语义非正交传输与融合

Minghui Wu, Zhen Gao

机构 * School of Information and Electronics, Beijing Institute of Technology (BIT)(信息与电子学院,北京理工大学) State Key Laboratory of Environment Characteristics and Effects for Near-space, Beijing(临近空间环境特征与效应国家重点实验室,北京) State Key Laboratory of CNS/ATM, Beijing(CNS/ATM国家重点实验室,北京) MIIT Key Laboratory of Complex-Field Intelligent Sensing, Beijing(工信部复杂场智能感知重点实验室,北京) BIT, Zhuhai(珠海北京理工大学) Advanced Technology Research Institute, BIT, Jinan(北京理工大学济南先进技术研究院) Yangtze Delta Region Academy, BIT, Jiaxing(长江三角洲地区学院,北京理工大学嘉兴)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种端到端学习的多模态语义非正交传输与融合框架,通过联合优化物理层和应用层任务提升大规模MIMO的频谱效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09801 2025-12-11 cs.CV 83%

Modality-Specific Enhancement and Complementary Fusion for Semi-Supervised Multi-Modal Brain Tumor Segmentation

模态特异性增强与互补融合用于半监督多模态脑肿瘤分割

Tien-Dat Chung, Ba-Thinh Lam, Thanh-Huy Nguyen, Thien Nguyen, Nguyen Lan Vi Vu, Hoang-Loc Cao, Phat Kim Huynh, Min Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种半监督多模态脑肿瘤分割框架,通过模态特异性增强模块和互补信息融合模块提升分割性能。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07430 2025-12-09 cs.LG cs.AI 83%

MIDG: Mixture of Invariant Experts with knowledge injection for Domain Generalization in Multimodal Sentiment Analysis

MIDG:基于知识注入的混合不变专家用于多模态情感分析中的领域泛化

Yangle Li, Danli Luo, Haifeng Hu

机构 * School of Electronics and Information Technology(电子信息学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 MIDG通过混合不变专家和跨模态适配器,提升多模态情感分析中领域泛化的性能与语义表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05515 2025-12-08 cs.CV cs.LG 83%

DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis

DashFusion: 基于分层瓶颈融合的双流对齐多模态情感分析

Yuhua Wen, Qifei Li, Yingying Zhou, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学信息科学与技术国家研究中心) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 DashFusion通过双流对齐与分层瓶颈融合技术,提升多模态情感分析的性能与效率。

Comments Accepted to IEEE Transactions on Neural Networks and Learning Systems (TNNLS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16054 2025-12-08 cs.CV 83%

Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model

基于多模态大语言模型的语言引导推理用于群体活动检测

Jihua Peng, Qianxiong Xu, Yichen Liu, Chenxi Liu, Cheng Long, Rui Zhao, Ziyue Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出LIR-GAD框架,通过多模态大语言模型实现群体活动检测,引入活动标记和群体标记以提升语义理解和分类性能。

Comments This work is being incorporated into a larger study

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02558 2025-12-03 cs.AI 83%

Empathy Level Prediction in Multi-Modal Scenario with Supervisory Documentation Assistance

多模态场景中基于监督文档辅助的共情水平预测

Yufei Xiao, Shangfei Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种结合视频、音频和文本信息的多模态共情预测方法,通过监督文档辅助训练提升文本特征提取,实验证明其优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00363 2025-12-02 cs.CV 83%

MM-DETR: An Efficient Multimodal Detection Transformer with Mamba-Driven Dual-Granularity Fusion and Frequency-Aware Modality Adapters

MM-DETR: 一种高效的多模态检测Transformer,采用Mamba驱动的双粒度融合和频率感知模态适配器

Jianhong Han, Yupei Wang, Yuan Zhang, Liang Chen

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) Beijing Institute of Technology Chongqing Innovation Center(北京理工大学重庆创新中心) National Key Laboratory for Space-Born Intelligent Information Processing(空间智能信息处理国家级重点实验室) School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MM-DETR通过Mamba驱动的双粒度融合和频率感知模态适配器,实现高效的多模态目标检测,提升检测精度与轻量化性能。

Comments Manuscript submitted to IEEE Transactions on Geoscience and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23287 2025-12-01 cs.LG cs.CL 83%

Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla

基于Transformer的三融合框架用于低资源孟加拉语多模态作者意图分类

Ariful Islam, Tanvir Mahmud, Md Rifat Hossen

机构 * Department of Computer Science(计算机科学系) Engineering Chittagong University of Engineering(工程学院恰尔达格工程大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出基于Transformer的三融合框架BangACMM,通过结合文本和视觉数据,在低资源孟加拉语社交媒体中实现作者意图分类,达到84.11%的宏F1得分,提升8.4个百分点。

Comments Accepted at the 28th International Conference on Computer and Information Technology (ICCIT 2025). To be published in IEEE proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22897 2025-12-01 cs.CV 83%

From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts

从点到云:学习多模态提示的稳健语义分布

Weiran Li, Yeqiang Liu, Yijie Wei, Mina Han, Xin Liu, Zhenbo Li

机构 * China Agricultural University(中国农业大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出P2C框架,通过动态去噪机制学习语义云分布,提升多模态提示学习的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22103 2025-12-01 cs.CV 83%

MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding

MoE3D:专家混合方法与多模态3D理解

Yu Li, Yuenan Hou, Yingmei Wei, Xinge Zhu, Yuexin Ma, Wenqi Shao, Yanming Guo

机构 * National University of Defense Technology(国防科技大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) ShanghaiTech University(上海科技大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MoE3D通过整合专家混合方法,提升多模态3D理解的性能,尤其在Multi3DRefer任务中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15711 2025-11-27 cs.AI q-bio.NC 83%

Semi-supervised Multimodal Representation Learning through a Global Workspace

通过全局工作空间实现半监督多模态表示学习

Benjamin Devillers, Léopold Maytié, Rufin VanRullen

机构 * CerCo, CNRS UMR 5549, Université de Toulouse and ANITI, Artificial and Natural Intelligence Toulouse Institute(CerCo、CNRS UMR 5549、图卢兹大学和ANITI人工智能与自然智能图卢兹研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种受全局工作空间概念启发的神经网络架构,通过自监督学习实现多模态表示对齐与转换,显著减少对匹配数据的需求。

Comments Under review

Journal ref IEEE Transactions on Neural Networks and Learning Systems 36 (5), 7843-7857 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20167 2025-11-26 cs.MM 83%

FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation

FINE: 通过互信息估计进行因子化多模态情感分析

Yadong Liu, Shangfei Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出了一种基于互信息估计的因子化多模态情感分析框架,通过分解模态为共享和独特表示,抑制噪声并提升情感表示质量,从而在多个数据集上优于现有方法。

Comments 15 pages, 9 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11895 2025-11-26 cs.CV 83%

Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration

通过高效校准实现统一多模态编码器的对抗鲁棒性

Chih-Ting Liao, Zhangquan Chen, Chunlei Meng, Tzu-Yu Huang, Xin Cao, Xu Zheng

机构 * UNSW Sydney(新南威尔士大学悉尼分校) Tsinghua University(清华大学) Fudan University(复旦大学) UTS(澳大利亚UTS大学) HKUST(GZ)(香港理工大学(广州))

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究提出高效对抗校准框架,提升统一多模态编码器的对抗鲁棒性,同时保持清洁性能,提升47.3%的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17952 2025-11-25 cs.CV 83%

Multi-speaker Attention Alignment for Multimodal Social Interaction

多说话者注意力对齐用于多模态社交互动

Liangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang, Ryosuke Furuta, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种多模态多说话者注意力对齐方法,通过动态头选择和自适应注意力偏差提升多模态社交互动理解能力,实现SOTA效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17681 2025-11-25 cs.CV 83%

Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models

基于多模态大语言模型的视觉-运动-参考对齐的指称多目标跟踪

Weiyi Lv, Ning Zhang, Hanyang Sun, Haoran Jiang, Kai Zhao, Jing Xiao, Dan Zeng

机构 * Shanghai University(上海大学) PAII Inc.(PAII公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 VMRMOT通过多模态大语言模型实现视觉-运动-参考对齐,提升指称多目标跟踪的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15122 2025-11-25 cs.IR cs.AI 83%

Multi-Aspect Cross-modal Quantization for Generative Recommendation

多方面跨模态量化用于生成性推荐

Fuwei Zhang, Xiaoyu Liu, Dongbo Xi, Jishen Yin, Huan Chen, Peng Yan, Fuzhen Zhuang, Zhao Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 MACRec通过多方面跨模态量化方法,提升生成性推荐中多模态信息利用和语义ID学习的性能。

Comments Accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14766 2025-11-20 cs.IR cs.MM 83%

OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction

Yang Li, Yajiao Wang, Wenhao Hu, Zhixiong Zhang, Mengting Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14604 2025-11-19 cs.CV 83%

XAttn-BMD: Multimodal Deep Learning with Cross-Attention for Femoral Neck Bone Mineral Density Estimation

Yilin Zhang, Leo D. Westbury, Elaine M. Dennison, Nicholas C. Harvey, Nicholas R. Fuggle, Rahman Attar

机构 * School of Electronics and Computer Science, University of Southampton, UK(电子与计算机科学学院,索姆塞特大学,英国) MRC Lifecourse Epidemiology Centre, University of Southampton, Southampton General Hospital, UK(生命课程流行病学研究中心,索姆塞特大学,南安普顿总医院,英国) NIHR Southampton Biomedical Research Centre, University of Southampton(南安普顿生物医学研究中心,索姆塞特大学) University Hospital NHS Foundation Trust, Southampton, UK(南安普顿国家健康服务基金会信托,英国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11 figures, 10 tables, 38 pages. Submitted to Artificial Intelligence in Medicine (currently with editor)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13755 2025-11-19 cs.LG cs.AI 83%

Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement

Zhe Yang, Wenrui Li, Hongtao Chen, Penghong Wang, Ruiqin Xiong, Xiaopeng Fan

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Zhengzhou Research Institute(哈尔滨工业大学郑州研究所) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究所) School of Mathematical Sciences, University of Electronic Science and Technology of China(数学学院,电子科学与技术大学) School of Electronic Engineering and Computer Science, Institute of Digital Media, Peking University(电子工程与计算机科学系,数字媒体研究所,北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12982 2025-11-18 cs.CR cs.CV 83%

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou, Bo Du, Mang Ye

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17686 2025-11-18 cs.CV 83%

Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration

Yuhang Han, Xuyang Liu, Zihan Zhang, Pengxiang Ding, Junjie Chen, Donglin Wang, Honggang Chen, Qingsen Yan, Siteng Huang

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02557 2025-11-13 eess.IV cs.CV 83%

RL-U$^2$Net: A Dual-Branch UNet with Reinforcement Learning-Assisted Multimodal Feature Fusion for Accurate 3D Whole-Heart Segmentation

Jierui Qu, Jianchun Zhao

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07722 2025-11-13 cs.AI 83%

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Zhi Yu, Qi Zheng, Ming Yan, Jiajun Bu

机构 * Zhejiang Key Laboratory of Accessible Perception and Intelligent Systems, Zhejiang University(浙江可感知智能系统重点实验室,浙江大学) Alibaba Group(阿里巴巴集团) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and DataSecurity(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21486 2025-11-12 cs.CV 83%

Reasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance

Zixuan Wang, Yu Sun, Hongwei Wang, Baoyu Jing, Xiang Shen, Xin Dong, Zhuolin Hao, Hongyu Xiong, Yang Song

机构 * TikTok Inc.(字节跳动公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Camera Ready for EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04078 2025-11-07 cs.CV 83%

Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment

Zehui Feng, Chenqi Zhang, Mingru Wang, Minuo Wei, Shiwei Cheng, Cuntai Guan, Ting Han

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 30 pages, 16 figures, under review as a conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12795 2025-11-07 cs.CV 83%

EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting

Wei Zhang, Miaoxin Cai, Yaqian Ning, Tong Zhang, Yin Zhuang, Shijian Lu, He Chen, Jun Li, Xuerui Mao

机构 * School of Interdisciplinary Science, Beijing Institute of Technology(交叉科学学院,北京理工大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing, Beijing Institute of Technology(空间智能信息处理国家重点实验室,北京理工大学) School of Optics and Photonics, Beijing Institute of Technology(光学与 photonics 学院,北京理工大学) State Key Laboratory of Explosion Science and Safety Protection, Beijing(爆炸科学与安全防护国家重点实验室,北京)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01444 2025-11-04 cs.AI 83%

Robust Multimodal Sentiment Analysis via Double Information Bottleneck

Huiting Huang, Tieliang Gong, Kai He, Jialun Wu, Erik Cambria, Mengling Feng

机构 * School of Computer Science(计算机科学学院) Technology, Xi’an Jiaotong University(技术,西安交通大学) Shaanxi Provincial Key Laboratory of Big Data Knowledge Engineering, Xi’an Jiaotong University(大数据知识工程省级重点实验室,西安交通大学) Saw Swee Hock School of Public Health, National University of Singapore(Saw Swee Hock 公共卫生学院,新加坡国立大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) School of Computer Science, Northwestern Polytechnical University(计算机科学学院,西北工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏