arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2026-06-25 至 2026-06-25 共收录 9 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 红外-可见光融合 1 篇

2512.24755 2026-06-25 eess.SY cs.SY 版本更新 50%

Asymmetry-Aware Routing for Industrial Multimodal Monitoring: A Diagnostic Framework

面向工业多模态监测的非对称感知路由:一个诊断框架

Sungwoo Kang

专题命中 红外-可见光融合 :multimodal fusion(abstract)

AI总结 提出非对称感知路由框架,通过三步诊断(单模态性能差距、门权重归因、模态损坏测试)决定多模态融合策略,在三个数据集上验证了其有效性。

Comments A subsequent extension of this analysis to a larger corpus found that the heterogeneity predictor is collinear with the industrial-versus-general domain split, so the corresponding effect is not identifiable on the available data

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 遥感融合与全色锐化 1 篇

2606.25174 2026-06-25 cs.LG cs.CV eess.IV 新提交 62%

An iterative energy-based multimodal transformer for joint retrieval of wheat soil moisture, leaf area index, and plant height from Sentinel-1 and Sentinel-2 time series

基于迭代能量多模态Transformer的Sentinel-1与Sentinel-2时序联合反演小麦土壤水分、叶面积指数和株高

Shubham Kumar Singh, Peilei Fan, Suraj A. Yadav, Rajendra Prasad, Prashant K Srivastava

机构 * Department of Urban and Environmental Policy and Program, Tufts University(Tufts大学城市与环境政策与项目系) Electrical and Computer Engineering Department, Mississippi State University(密苏里州立大学电气与计算机工程系) Department of Physics, Indian Institute of Technology (BHU)(印度理工学院(BHU)物理系) Institute of Environment and Sustainable Development, Banaras Hindu University(巴纳尔斯赫尔大学环境与可持续发展研究所)

专题命中 遥感融合与全色锐化 :multimodal fusion(abstract);分类 cs.CV、eess.IV

AI总结 提出迭代能量Transformer(iEBT),通过嵌入多模态预测器并迭代更新目标向量,从Sentinel-1和Sentinel-2时序数据中联合反演土壤水分、叶面积指数和株高,在印度瓦拉纳西实测数据上取得高精度,并利用终端能量作为质量诊断指标。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 医学影像融合 2 篇

2601.10386 2026-06-25 cs.CV cs.AI cs.MM 62%

Handling Missing Modalities in Multimodal Survival Prediction for Non-Small Cell Lung Cancer

多模态生存预测中缺失模态的处理:非小细胞肺癌的应用

Filippo Ruffini, Camillo Maria Caruso, Claudia Tacconi, Lorenzo Nibid, Francesca Miccolis, Marta Lovino, Carlo Greco, Edy Ippolito, Michele Fiore, Alessio Cortellini, Bruno Beomonte Zobel, Giuseppe Perrone, Bruno Vincenzi, Claudio Marrocco, Alessandro Bria, Elisa Ficarra, Sara Ramella, Valerio Guarrasi, Paolo Soda

机构 * Unit of Artificial Intelligence and Computer Systems, Department of Engineering(人工智能与计算机系统单位,工程系) Università Campus Bio-Medico di Roma(罗马生物医学大学) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入部门,辐射物理,生物医学工程) Umeå University(乌梅学院) Research Unit of Radiation Oncology, Department of Medicine and Surgery(放射肿瘤研究单位,医学与外科部门) Fondazione Policlinico Universitario Campus Bio-Medico(生物医学大学附属医院基金会) Anatomical Pathology Operative Research Unit(解剖病理学操作研究单位) Research Unit of Anatomical Pathology, Department of Medicine and Surgery(解剖病理学研究单位,医学与外科部门) Department of Medicine and Surgery(医学与外科部门)

专题命中 医学影像融合 :multimodal fusion(abstract);分类 cs.CV、cs.MM

AI总结 本文提出一种缺失感知的多模态生存框架,结合CT、WSI和临床数据,通过基础模型提取特征并融合,提升NSCLC生存预测的准确性与临床适用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10702 2026-06-25 cs.CV cs.AI 版本更新 57%

Backbone-Conditional Behavior of Modality Gating in Multi-Modal Prostate MRI Segmentation: A 5-Fold Cross-Validation and Gate Mechanism Analysis

模态隔离门控融合:用于稳健多模态前列腺MRI分割的架构无关方法

Yongbo Shu, Wenzhao Xie, Shanhu Yao, Zirui Xin, Luo Lei, Kewen Chen, Aijing Luo

机构 * The Second Xiangya Hospital of Central South University(中南大学湘雅医学院第二医院) The Third Xiangya Hospital of Central South University(中南大学湘雅医学院第三医院) School of Life Sciences, Central South University(中南大学生命科学学院) Hunan Provincial Key Laboratory of Medical Information Research (Central South University)(湖南省医学信息研究重点实验室(中南大学)) Hunan Provincial Clinical Medical Research Center for Cardiovascular Intelligent Medicine(湖南省心血管智能医学临床医学研究中心)

专题命中 医学影像融合 :multi-modal fusion(abstract);分类 cs.CV

AI总结 本文提出模态隔离门控融合(MIGF)方法,通过独立编码流和门控阶段提升多模态前列腺MRI分割的鲁棒性,实验表明其在不同模态缺失和伪影情况下均有效。

Comments Major revision. Single-fold analysis replaced by 5-fold cross-validation (180 trained models) plus a direct gate-mechanism analysis; conclusions updated to show that modality gating is backbone-conditional. Supersedes v1

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 机器人多传感器融合 4 篇

2606.24986 2026-06-25 cs.LG cs.AI 新提交 86%

When Multi-Sensor Fusion Fails to Generalize: Cattle Posture Classification Under Animal-Level and Temporal Distribution Shift

当多传感器融合无法泛化:动物级别和时间分布偏移下的牛姿势分类

Leutrim Uka, Severino Pinto, Gundula Hoffmann, Marina M. -C. Höhne

机构 * Institute of Computer Science, University of Potsdam(波茨坦大学计算机科学研究所) Department of Sensors and Modelling, Leibniz Institute for Agricultural Engineering and Bioeconomy - ATB(莱布尼茨农业工程与生物经济研究所传感器与建模系) Department of Data Science in Bioeconomy, Leibniz Institute for Agricultural Engineering and Bioeconomy - ATB(莱布尼茨农业工程与生物经济研究所生物经济数据科学系)

专题命中 机器人多传感器融合 :sensor fusion(title,abstract);multi-sensor fusion(title)

AI总结 研究评估多传感器融合在牛姿势分类中的鲁棒性,发现跨年评估性能大幅下降(F1从0.94降至0.49),表明常用评估高估实际性能,融合可能降低而非提升鲁棒性。

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25953 2026-06-25 cs.RO cs.CV 新提交 62%

DSP-SLAM++: A Unified Framework for Multi-Class, High-Fidelity Object SLAM in the Wild

DSP-SLAM++:面向野外多类高保真物体SLAM的统一框架

Ahmad Kourani, Ghina Daoud, Daniel Asmar, Imad Elhajj

机构 * American University of Beirut(贝鲁特美国大学)

专题命中 机器人多传感器融合 :sensor fusion(abstract);分类 cs.CV、cs.RO

AI总结 提出DSP-SLAM++,通过异步建图流水线和单目鱼眼-激光雷达传感器融合,在支持多类物体的同时实现实时高保真物体建模,将最大物体处理延迟降低70%。

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25796 2026-06-25 cs.RO 新提交 57%

MIL-LC: A Robust Magnetometer-Inertial-LiDAR Fusion Multimodal Localization Framework

MIL-LC:一种鲁棒的磁力计-惯性-LiDAR融合多模态定位框架

Qiyang Lyu, Zhenyu Wu, Wei Wang, Hongming Shen, Danwei Wang

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电气与电子工程学院) Centre for Advanced Robotics Technology Innovation (CARTIN)(先进机器人技术创新中心)

专题命中 机器人多传感器融合 :multimodal fusion(abstract);分类 cs.RO

AI总结 提出MIL-LC框架,融合磁力计、惯性和LiDAR,解决GNSS缺失、几何重复或无纹理环境下的定位问题,通过磁地图与LiDAR退化检测实现鲁棒定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25360 2026-06-25 cs.RO 新提交 57%

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

解耦语义与几何基础:面向语言条件模仿学习的空间视觉提示

Yanzhe Tang, Xinyu Shao, Yuxuan Hu, Siyu Chen, Bowen Yang, Yajun Gao, Tongtong Cao, Xiu Li, Long Zeng

机构 * Tsinghua University(清华大学) Huawei Technologies Co., Ltd.(华为技术有限公司) National University of Singapore(新加坡国立大学)

专题命中 机器人多传感器融合 :feature-level fusion(abstract);分类 cs.RO

AI总结 提出SVP-IL解耦架构,通过视觉语言基础模型将指令解析为零样本几何掩码,生成空间视觉提示并注入连续动作生成器,在低数据条件下显著提升语言条件模仿学习的成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 融合架构与评测 1 篇

2606.25278 2026-06-25 cs.CV cs.AI 新提交 57%

Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation

异构与专长快照蒸馏用于3D语义分割

Xiaopei Wu, Yuenan Hou, Junkai Xu, Wenxiao Wang, Binbin Lin, Yu Li, Ping Li, Haifeng Liu, Deng Cai, Wanli Ouyang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV

AI总结 提出异构和专长快照知识蒸馏(HAS-KD),通过信息导向异构蒸馏(IHD)和专长快照蒸馏(ASD)将多模态模型和多个模型专家的知识高效迁移到点云网络,在ScanNetV2和S3DIS上取得最先进结果。

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏