arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2308.09455 2023-08-21 cs.CV cs.AI cs.CL 75%

Artificial-Spiking Hierarchical Networks for Vision-Language Representation Learning

Yeming Chen, Siyu Zhang, Yaoru Sun, Weijian Liang, Haoran Wang

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06798 2023-04-17 cs.AI cs.CL cs.CV 75%

On the Opportunities and Challenges of Foundation Models for Geospatial Artificial Intelligence

Gengchen Mai, Weiming Huang, Jin Sun, Suhang Song, Deepak Mishra, Ninghao Liu, Song Gao, Tianming Liu, Gao Cong, Yingjie Hu, Chris Cundy, Ziyuan Li, Rui Zhu, Ni Lao

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02141 2023-02-07 cs.CV cs.CL cs.MM 75%

LipFormer: Learning to Lipread Unseen Speakers based on Visual-Landmark Transformers

Feng Xue, Yu Li, Deyin Liu, Yincen Xie, Lin Wu, Richang Hong

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13061 2022-08-16 cs.CV cs.AI cs.CL 75%

NewsStories: Illustrating articles with visual summaries

Reuben Tan, Bryan A. Plummer, Kate Saenko, JP Lewis, Avneesh Sud, Thomas Leung

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14411 2025-10-17 cs.LG cs.MM cs.SD eess.AS 74%

Revisit Modality Imbalance at the Decision Layer

Xiaoyu Ma, Hao Chen

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),教育部,中国)

专题命中 多模态训练与对齐 :multimodal(abstract,comments);audio-visual(abstract);分类 cs.MM、eess.AS

Comments Some Insights in Balanced Multimodal Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11472 2026-08-13 cs.CV cs.LG 新提交 74%

Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification

用于多模态IPMN风险分层的堆叠集成的高斯元空间增强

Max A. Nelson, Eminenur Sen Tasci, Zhixiang Wang, Zongwei Zhou, Halil Ertugrul Aktas, Andrea M. Bejar, Elif Keles, Ziliang Hong, Sıtkı Safa Taflan, Muhammed Enes Tasci, Frank H. Miller, Michael B. Wallace, Rajesh N. Keswani, Gorkem Durak, Ulas Bagci

机构 * Northwestern University(西北大学) Johns Hopkins University(约翰斯·霍普金斯大学) Istanbul University(伊斯坦布尔大学) Mayo Clinic Florida(佛罗里达州梅奥诊所)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 该研究针对IPMN风险分层问题,提出cUPMI方法并结合多模态信息融合,构建RF堆叠模型,在多中心分析中取得优于基线的性能。

Comments Accepted at the International Workshop on Machine Learning in Medical Imaging (MLMI 2026), held in conjunction with MICCAI 2026. This is the authors' accepted manuscript; the final version will appear in Springer Lecture Notes in Computer Science (LNCS). 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10489 2026-08-12 cs.CV 新提交 74%

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

当视觉变为文本:视觉语言模型(VLM)中基于跨模态残差引导的视觉令牌剪枝

Congyang Ou, Ruike Song, Yang Zhou, Libo Sun, Haokui Zhang, Zhenbo Luo

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

AI总结 该研究针对视觉语言模型(VLM)海量视觉令牌导致推理成本高的问题,提出基于跨模态残差(CMR)的无训练压缩方法SIEVE,在LLaVA-NeXT-7B上仅保留11.1%视觉令牌,实现显著加速与缓存缩减且性能损失极小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08896 2026-08-11 cs.AI 新提交 74%

From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings

从手册到维护:在低资源场景中微调MedGemma以提供多模态成像系统支持

Bernes Lorier Atabonfack, Zion Kongbi Nfo, Ahmed Tahiru Issah, Tolulope Olusuyi, Clemence Ingabire, Mohammed Hardi Abdul Baaki, Mawuli Deku, Abdulrazaq Zubair, Alyasaa Anas, Raymond Confidence, Maruf Adewole, Udunna C. Anazodo

机构 * Carnegie Mellon University Africa(非洲卡内基梅隆大学) IngenziAI(因根齐人工智能公司) Federal University of Health Sciences Azare(阿扎雷联邦健康科学大学) University of Maiduguri Teaching Hospital(迈杜古里大学教学医院) Lawson Health Research Institute(劳森健康研究所) University of Pennsylvania(宾夕法尼亚大学) McGill University(麦吉尔大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.AI

AI总结 针对中低收入国家医疗设备维护难题,研究人员构建INGENZI_DatasetV1数据集,用QLoRA微调MedGemma-4b-it模型,使维修问答指标大幅提升,为低资源场景AI辅助医疗维护奠定基础。

Comments Accepted at the AFRICAI 2026 Workshop, a satellite event at MICCAI 2026. To appear in Springer Lecture Notes in Computer Science (LNCS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14243 2026-08-07 cs.CV 版本更新 74%

Multi-Representation Geometric Hierarchy Fusion: An Implicit-Submap Driven Framework for Resilient 3D Place Recognition

多表示几何层次融合:一种用于鲁棒三维地点识别的隐式子图驱动框架

Xiaohui Jiang, Haijiang Zhu, Chade Li, Ning An

机构 * College of Information and Technology, Beijing University of Chemical Technology(信息与技术学院,北京化工大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Research Institute of Mine Artificial Intelligence, China Coal Research Institute(矿山人工智能研究院,中国煤炭科研院)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

AI总结 针对激光雷达地点识别中描述子不稳定与表示脆弱性问题,本文提出隐式神经点驱动的多表示几何层次融合框架,在多数据集上实现鲁棒性能,平衡了精度、效率与内存占用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13386 2026-07-16 cs.CV 新提交 74%

FM$^2$: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging

FM$^2$:用于异构多模态医学成像的统一联邦基础模型

Shengchao Chen, Ting Shu

机构 * School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Australian AI Institute, University of Technology Sydney(悉尼科技大学澳大利亚人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 针对医学成像基础模型构建中隐私与任务统一问题,提出FM$^2$框架,通过从头训练核心主干、结合预训练编码器、配备双混合专家模块及正则化器,并引入字幕增强学习,实现跨模态泛化,优于现有联邦基线。

Comments Accepted by ACM MM 2026 (Main Track): the 34th ACM International Conference on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01772 2026-07-03 cs.CV eess.SP 新提交 74%

LLM-Empowered Multimodal Fusion Framework for Autonomous Driving: Semantic Enhancement and Channel-Adaptive Design

LLM赋能的自动驾驶多模态融合框架:语义增强与信道自适应设计

Wen Wang, Yaping Sun, Yejun He, Hao Chen, Zhiyong Chen, Xiaodong Xu, Nan Ma, Shuguang Cui

机构 * National Natural Science Foundation of China(国家自然科学基金) Shenzhen Natural Science Foundation(深圳市自然科学基金) Shenzhen Key Laboratory Evaluation(深圳市重点实验室评估)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 提出LM-SCIP框架,以LLM为核心推理中枢,通过信道自适应语义模块动态融合视觉与雷达特征,实现低信噪比下的视觉主导回退和高信噪比下的协同融合,在nuScenes和VIRAT数据集上显著提升定位精度。

Comments 6 pages, 4 figures. Accepted by 2026 IEEE 37th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29949 2026-06-30 eess.IV cs.AI q-bio.GN 74%

Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction

基于组织病理学的分子预测的数据高效多模态对齐

Dominik Winter, Dominik Vonficht, Loïc Le Bescond, Christian Gebbe, Marco Rosati, Richard J. Chen, Markus Schick, Ross Stewart, Nicolas Brieu

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

AI总结 提出一种轻量级对齐模块,在冻结的组织病理学和RNA-Seq基础模型上训练,实现无需测序或重新训练即可通过基因集签名预测通路活性的开放词汇分子提示,并在多癌队列中实现25倍检索提升。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12957 2026-06-30 cs.LG cs.AI 74%

Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning

揭示-修订:具有多模态注意力的可解释性偏见感知生成建模

Noor Islam S. Mohammad, Uluğ Bayazıt

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

AI总结 本文提出一个统一了跨模态注意力融合、Grad-CAM++属性分析和揭示-修订反馈循环的可解释性偏见感知生成框架,通过多模态MNIST和Fashion MNIST数据集验证了其在图像生成和子组审计中的性能。

Comments We are recently authors in conflict with this work; I am heartily requesting to withdraw this paper as soon as possible

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11687 2026-06-11 cs.CV cs.LG cs.RO 新提交 74%

DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace

DroneShield-AI:一种用于受争议空域中实时自主无人机威胁检测、行为意图分类和群体智能的多模态传感器融合框架

Marius Bayizere

机构 * Independent Researcher(独立研究者)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 提出DroneShield-AI框架,集成RF信号分类、声学检测、YOLOv8视觉检测等六层处理,通过行为意图分类引擎(BICE)实现六类威胁分类并提前30秒预警,以及图神经网络群体智能模块(GNN-SIM)分析多无人机编队,在低成本硬件上达到96.1%检测精度和142ms延迟。

Comments 23 pages, 6 figures, 11 tables. Code available at https://github.com/bayizeremarius/DroneShield-AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06102 2026-06-05 cs.AI cs.LG 74%

Step-adaptive multimodal fusion network with multi-scale cloud feature learning for ultra-short-term solar irradiance forecasting

步进自适应多模态融合网络与多尺度云特征学习用于超短期太阳辐照度预测

Jingxin Zhang Xiaoqin Wang

机构 * School of Automation, Southeast University(自动化学院,东南大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

AI总结 提出一种步进自适应多模态融合网络,通过InceptionNeXt提取多尺度云特征、步进自适应低频补偿单元动态调整低频信息,并结合气象时间序列特征进行超短期太阳辐照度预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03114 2026-06-03 cs.CV 74%

FAF-CD: Frequency-Aware Fusion for Change Detection under Imperfect Multimodal Remote Sensing

FAF-CD: 面向不完美多模态遥感的频率感知融合变化检测

Yufan Wang, Sokratis Makrogiannis, Chandra Kambhamettu

机构 * University of South Florida(佛罗里达州立大学) Delaware State University(特拉华州立大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 提出频率感知混合框架FAF-CD,通过DINOv3预训练ConvNeXt编码器、VMamba解码器及修正感知三支融合模块(可变形空间对齐+傅里叶/哈尔小波比较+自适应门控),在不完美异质遥感(如EO-SAR)和二元光学变化检测中提升精度并降低计算成本。

Comments Code will be released at https://github.com/VimsLab/FAF-CD

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01334 2026-06-02 cs.CV 74%

HOLA: Holistic Multi-Modal Alignment for Open-Set 3D Recognition

HOLA: 面向开放集3D识别的全息多模态对齐

Koby Aharonov, Oren Shrout, Ayellet Tal

机构 * Technion – Israel Institute of Technology(技术ion-以色列理工学院)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 提出HOLA方法,通过解耦多正例对比损失和对齐点云与多视图图像及文本描述,实现开放集3D识别中的全息多模态对齐,在长尾基准上取得最先进零样本性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21075 2026-05-21 cs.CV cs.LG 74%

SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining

SpectralEarth-FM: 将高光谱图像引入多模态地球观测预训练

Nassim Ait Ali Braham, Aaron Banze, Conrad M. Albrecht, Julien Mairal, Jocelyn Chanussot, Xiao Xiang Zhu

机构 * Chair of Data Science in Earth Observation(地球观测数据科学主任) Technical University of Munich(慕尼黑技术大学) Remote Sensing Technology Institute(遥感技术研究所) German Aerospace Center (DLR)(德国航空航天中心) Department of Aerospace Engineering(航空航天工程系) University of the Bundeswehr Munich(联邦国防军慕尼黑大学) LEAP Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Univ. Grenoble Alpes(格勒诺布尔阿尔卑斯大学) Inria(法国国家信息与自动化技术研究院) CNRS(法国国家科学研究中心) Grenoble INP(格勒诺布尔INP) LJK

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 本文提出SpectralEarth-FM,一种用于多传感器地球观测输入的分层变压器,旨在联合处理高光谱图像与低通道观测。通过构建SpectralEarth-MM数据集,采用JEPA风格的目标进行预训练,实现了在高光谱下游任务和标准EO基准上的最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09180 2026-05-21 cs.CV cs.RO 74%

Multimodal Fusion for Sim2real Transfer in Visual Reinforcement Learning

多模态融合用于视觉强化学习中的仿真到现实迁移

Zichun Xu, Jingdong Zhao, Chenyu Guo, Qianxue Zhang, Liao Zhang, Xiao Zhang, Yiming Ren, Lian Zhang, Zengren Zhao

机构 * Medical Artificial Intelligence Lab, The First Hospital of Hebei Medical University, Hebei Medical University(医学人工智能实验室,河北医科大学第一医院,河北医科大学) State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 本文提出基于视觉变换器的多模态融合框架,通过融合RGB和深度信息提升泛化能力,并设计对比学习方案和课程式域随机化方案以提高样本效率和迁移性能,实验结果表明该方法在现实任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03606 2026-05-12 cs.AI 74%

BioBLP: A Modular Framework for Learning on Multimodal Biomedical Knowledge Graphs

BioBLP: 一个用于多模态生物医学知识图谱学习的模块化框架

Daniel Daza, Dimitrios Alivanistos, Payal Mitra, Thom Pijnenburg, Michael Cochez, Paul Groth

机构 * Vrije Universiteit Amsterdam(荷兰阿姆斯特丹自由大学) University of Amsterdam(阿姆斯特丹大学) Elsevier B.V.(埃森哲公司) Discovery Lab, Elsevier(埃森哲发现实验室)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

AI总结 本文提出BioBLP框架,用于处理多模态生物医学知识图谱中的实体属性,支持不同模态的数据编码及缺失属性处理,并通过预训练策略提升训练效率,在药物-蛋白质相互作用预测任务中表现优于基线方法。

Journal ref J Biomed Semant 14, 20 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08288 2026-05-12 cs.LG cs.AI cs.CR cs.DC 74%

UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator Alignment

UMEDA:统一多模态高效数据融合用于隐私保护图联邦学习的谱门控注意力与扩散算子对齐

Shih-Yu Lai, Hirozumi Yamaguchi, Shang-Tse Chen, Yu-Lun Liu, Bing-Yu Chen

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,Location,Country) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,Location,Country)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.AI

AI总结 UMEDA通过谱门控注意力和扩散算子对齐,实现多模态数据在隐私保护下的高效联邦学习,提升准确性和通信效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26524 2026-05-11 cs.LG cs.AI 74%

TAP: Two-Stage Adaptive Personalization of Multi-Task and Multi-Modal Foundation Models in Federated Learning

TAP: 多任务和多模态基础模型在联邦学习中的两阶段自适应个性化

Seohyun Lee, Wenzhi Fang, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. Brinton

机构 * Purdue University(普渡大学) Yonsei University(延世大学) University at Buffalo-SUNY(布法罗大学-苏尼尔)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.AI

AI总结 本文提出TAP方法,通过两阶段自适应个性化在联邦学习中提升多任务和多模态基础模型的泛化能力,解决数据、任务和模态异质性问题。

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00357 2026-05-04 cs.GR cs.HC cs.MM 74%

Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML

面向人类理解的机器学习函数多模态交互表示

Bokang Wang, Yingxuan Liao, Leah Lee, Jack Wesson, Anlan Yang, Ruizi Wang, Yigang Wen

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.MM

AI总结 本文通过交互式可视化增强对机器学习的理解,旨在通过透明数据集探索提升对ML的积极态度,鼓励更多人探索该领域。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07451 2026-05-01 cs.CV 74%

A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion

动态神经网络综述:从计算机视觉到多模态传感器融合

Fabio Montello, Ronja Güldenring, Simone Scardapane, Lazaros Nalpantidis

机构 * DTU Electrical and Photonics Engineering, Technical University of Denmark(丹麦技术大学电气与光子工程系) DIET Department, Sapienza University of Rome(罗马萨皮恩扎大学DIET系)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 本文综述了动态神经网络在计算机视觉中的应用,探讨了其在多模态传感器融合中的优势,提出了一种基于网络组件适应性的分类方法,并提供了相关研究的资源库。

Comments Under review at Image and Vision Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08533 2026-04-29 cs.RO cs.AI 74%

ACROSS: A Deformation-Based Cross-Modal Representation for Robotic Tactile Perception

ACROSS: 一种基于变形的跨模态表示用于机器人触觉感知

Wadhah Zai El Amri, Malte Kuhlmann, Nicolás Navarro-Guerrero

机构 * L3S Research Center(L3S研究所以)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.AI

AI总结 本文提出ACROSS框架,通过利用传感器变形信息实现触觉数据跨传感器转换,将BioTac信号转换为DIGIT传感器数据,解决现有数据集在新设备上的应用问题。

Comments Accepted to 2025 IEEE Conference on Robotics and Automation (ICRA 2025). arXiv admin note: text overlap with arXiv:2410.14310

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21573 2026-04-24 cs.CV q-bio.QM 74%

CHRep: Cross-modal Histology Representation and Post-hoc Calibration for Spatial Gene Expression Prediction

CHRep: 跨模态组织切片表示与事后校准用于空间基因表达预测

Changfan Wang, Xinran Wang, Donghai Liu, Fei Su, Lulu Sun, Zhicheng Zhao, Zhu Meng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室) Peking University Third Hospital(北京大学第三医院)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

AI总结 CHRep通过两阶段框架提升组织切片到基因表达的预测鲁棒性,结合结构感知表示学习与事后校准,提升滑片级变化下的稳定性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18790 2026-04-22 cs.CV 74%

EfficientPENet: Real-Time Depth Completion from Sparse LiDAR via Lightweight Multi-Modal Fusion

EfficientPENet:通过轻量多模态融合实现稀疏LiDAR的实时深度补全

Johny J. Lopez, Md Meftahul Ferdaus, Mahdi Abdelguerfi, Anton Netchaev, Steven Sloan, Ken Pathak, Kendall N. Niles

机构 * Canizaro Livingston Gulf States Center for Environmental Informatics, the University of New Orleans, New Orleans, USA(Canizaro Livingston Gulf States环境信息中心,新奥尔良大学,美国新奥尔良) US Army Corps of Engineers, Engineer Research and Development Center, Vicksburg, Mississippi, USA(美国陆军工程兵团,工程师研究与发展中心,密西西比州维克斯堡,美国)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 本文提出EfficientPENet,通过轻量级多模态融合网络,在稀疏LiDAR和RGB图像上实现实时深度补全,具有更少的参数和更高的速度,同时保持竞争力的精度。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18231 2026-04-21 cs.LG cs.AI 74%

Rethinking Cross-Modal Fine-Tuning: Optimizing the Interaction Between Feature Alignment and Target Fitting

重新思考跨模态微调:优化特征对齐与目标拟合之间的交互

Trong Khiem Tran, Manh Cuong Dao, Phi Le Nguyen, Thao Nguyen Truong, Trong Nghia Hoang

机构 * Washington State University(华盛顿州立大学) National University of Singapore(新加坡国立大学) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院) Hanoi University of Science and Technology(河内科学技术大学)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.AI

AI总结 本文提出一种原理性框架,通过特征-标签扭曲概念解释特征对齐与目标拟合的交互,建立目标误差的可证明泛化界,提升跨模态微调性能。

Comments Accepted AISTATS 20226

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10347 2026-04-14 cs.CV 74%

Multi-modal, multi-scale representation learning for satellite imagery analysis just needs a good ALiBi

多模态、多尺度表示学习用于卫星影像分析只需一个良好的ALiBi

Patrick Kage, Pavlos Andreadis

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 本文提出Scale-ALiBi机制,通过空间编码偏置提升多尺度多模态卫星影像表示学习效果,并在GEO-Bench基准上取得改进。

Comments Originally appeared at the 4th Space Imaging Workshop at the Georgia Institute of Technology, October 7-9, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08018 2026-04-02 cs.CV 74%

Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared

不再缺失:字典引导的跨模态图像融合在红外缺失情况下的应用

Yafei Zhang, Meng Ma, Huafeng Li, Yu Liu

机构 * Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院) Department of Biomedical Engineering, Hefei University of Technology(合肥工业大学生物医学工程系)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

AI总结 本文提出了一种基于共享卷积字典的字典引导框架,解决红外缺失时的跨模态图像融合问题,通过联合字典学习、视觉引导红外推断和自适应融合方法提升感知质量和下游检测性能。

Comments This paper has been accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏