arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2026-05-19 至 2026-05-19 共收录 12 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 通用Image Fusion 1 篇

2605.17369 2026-05-19 astro-ph.SR 78%

The Deep Learning-Based Dual-Branch Multimodal Fusion Model for Solar Flare Prediction

基于深度学习的双分支多模态融合模型用于太阳耀斑预测

Limin Zhao, Xingyao Chen, Xiaoshuai Zhu, Dong Zhao, and Yihua Yan

专题命中 通用Image Fusion :multimodal fusion(title,abstract)

AI总结 本文提出一种双分支多模态融合深度学习模型,用于预测24小时内的太阳耀斑,通过交叉注意力机制整合磁图和磁参数,并在特征层面进行跨尺度交互以增强多尺度表示,同时实现了二分类和多分类任务,实验结果表明模型在预测X级耀斑方面表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 红外-可见光融合 2 篇

2605.16887 2026-05-19 cs.CV cs.LG 70%

Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet

Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet

Xin Niu, Enyi Li, Jinchao Liu, Yan Wang, Margarita Osadchy, Yongchun Fang

机构 * Tianjin Key Laboratory of Intelligent Robotics, College of Artificial Intelligence, Nankai University, China(天津智能机器人重点实验室,人工智能学院,南开大学,中国) Engineering Research Center of Trusted Behavior Intelligence, Ministry of Education, Nankai University, China(可信行为智能工程研究中心,教育部,南开大学,中国) Department of Computer Science, Haifa University, Israel(计算机科学系,海法大学,以色列) VisionMetric Ltd, Canterbury, Kent, UK(VisionMetric Ltd,坎特伯雷,肯特,英国)

专题命中 红外-可见光融合 :infrared and visible(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出了一种紧凑的编码器-解码器神经模块(cmUNet),通过跨模态转换和模态内重建,学习模态无关的表示,同时保留身份相关的信息。此外,作者提出了MarrNet,通过将cmUNet连接到标准特征提取网络,实现跨模态匹配,并在多个挑战性任务上验证了其优越性能。

Comments Published in IEEE Transactions on Image Processing. See full abstract in the PDF file

Journal ref n IEEE Transactions on Image Processing, vol. 33, pp. 655-670, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17564 2026-05-19 cs.CV 57%

A Conditional U-Net Pipeline with Pre- and Post-Processing for Aerial RGB-to-Thermal Image Translation

具有预处理和后处理的条件U-Net管道用于航空RGB到热图像转换

Tseten Sherpa, Sikandar Ali, Shubham Parab, Haoyun Feng, Matthew Dennis, Keenan Gibbons, Verrah Otiende, Geoffrey H. Siwo

机构 * Department of Data Science, University of Michigan, Ann Arbor, MI, USA(数据科学系,密歇根大学,安阿伯,MI,美国) Department of Information Science, University of Michigan, Ann Arbor, MI, USA(信息科学系,密歇根大学,安阿伯,MI,美国) Department of Computer Science, University of Michigan, Ann Arbor, MI, USA(计算机科学系,密歇根大学,安阿伯,MI,美国) Arcknow, New York, USA(Arcknow,纽约,美国) School of Environmental Sustainability, University of Michigan, Ann Arbor, MI, USA(可持续环境学院,密歇根大学,安阿伯,MI,美国) SmithGroup, Ann Arbor, MI, USA(SmithGroup,安阿伯,MI,美国) Michigan Institute for Data and AI in Society (MIDAS), University of Michigan, Ann Arbor, MI, USA(密歇根数据与人工智能社会研究院(MIDAS),密歇根大学,安阿伯,MI,美国) United States International University (USIU), Nairobi, Kenya(美国国际大学(USIU),内罗毕,肯尼亚) Department of Learning Health Sciences, University of Michigan Medical School, Ann Arbor, MI, USA(学习健康科学系,密歇根大学医学院,安阿伯,MI,美国) Department of Pharmacology, University of Michigan Medical School, Ann Arbor, MI, USA(药理学系,密歇根大学医学院,安阿伯,MI,美国) Center for Global Health Equity, University of Michigan, Ann Arbor, MI, USA(全球健康公平中心,密歇根大学,安阿伯,MI,美国)

专题命中 红外-可见光融合 :image fusion(abstract);分类 cs.CV

AI总结 本文提出了一种基于条件U-Net的简单架构,结合天气数据和针对性预处理与后处理技术,以提高航空RGB到热图像转换的性能,实验结果显示其在PSNR、SSIM和LPIPS指标上优于现有方法。

Comments 8 pages, 7 figures, NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 遥感融合与全色锐化 1 篇

2505.18991 2026-05-19 cs.CV 79%

Fast Kernel-Space Diffusion for Remote Sensing Pansharpening

快速核空间扩散用于遥感全色锐化

Hancong Jin, Zihan Cao, Liang-jian Deng, Jingjing Li

机构 * University of Electronic Science and Technology of China(电子科技大学)

专题命中 遥感融合与全色锐化 :pansharpening(title,abstract);分类 cs.CV

AI总结 本文提出KSDiff框架,通过整合低秩核心张量生成器和统一因子生成器,利用结构感知的多头注意力机制生成增强全局上下文的卷积核,以提升全色锐化质量并加速推理,实验表明其在性能和效率上均优于现有方法。

Comments CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 机器人多传感器融合 4 篇

2605.16414 2026-05-19 cs.CV 86%

NERVE: A Neuromorphic Vision and Radar Ensemble for Multi-Sensor Fusion Research

NERVE:一种用于多传感器融合研究的神经形态视觉与雷达集成系统

Omar Mansour, Pietro Martinello, Ethan Milon, YingFu Xu, Manolis Sifalakis, Guangzhi Tang, Amirreza Yousefzadeh

机构 * 1University of Twente, Netherlands 2University of Modena 3University of Strasbourg, France 4IMEC the Netherlands, Netherlands 5Innatera, Netherlands 6Maastricht University, Netherlands

专题命中 机器人多传感器融合 :sensor fusion(title);multi-sensor fusion(title);multi-modal fusion(abstract);分类 cs.CV

AI总结 NERVE包含257分钟同步记录的多传感器数据,用于评估多模态融合,通过DVS与雷达结合提升人体检测和距离估计性能。

Comments To be published in ICJNN 2026 Maastricht

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02463 2026-05-19 stat.ML cs.LG 50%

Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation

分布变换器:通过实时先验适应实现快速近似贝叶斯推断

George Whittle, Juliusz Ziomek, Jacob Rawling, Maike A. Osborne

机构 * Mind Foundry Ltd(Mind Foundry有限公司)

专题命中 机器人多传感器融合 :sensor fusion(abstract)

AI总结 本文提出分布变换器,一种能够学习任意分布到分布映射的新型架构,通过实时先验适应实现快速近似贝叶斯推断,显著降低计算时间并达到与现有方法相当或更优的对数似然性能。

Comments Spotlight acceptance at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18926 2026-05-19 eess.SY cs.SY 50%

Inertial-Based LQG Control: A New Look at Inverted Pendulum Stabilization

基于惯性力学的LQG控制:对倒立摆稳定化的新视角

Daniel Engelsman, Itzik Klein

专题命中 机器人多传感器融合 :sensor fusion(abstract)

AI总结 本文提出利用局部微分平坦性将高阶动力学纳入系统模型,改进LQG控制器以更稳健地稳定动态稳定移动平台,尤其在传感器有限的户外环境中。

Comments 11 pages, 10 figures, 5 tables

Journal ref International Journal of Mechanical System Dynamics, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08571 2026-05-19 eess.SY cs.SY 50%

Parametric and State Estimation of Stationary MEMS-IMUs: A Tutorial

参数与状态估计在静态MEMS惯性测量单元中的应用:教程

Daniel Engelsman, Yair Stolero, Itzik Klein

专题命中 机器人多传感器融合 :information fusion(abstract)

AI总结 本文探讨了静态MEMS惯性测量单元的参数与状态估计方法,分析了误差与时间、传感器数量的关系,并通过实验验证了多传感器应用的潜力。

Journal ref IEEE Access, volume 12, pages 148592-148604, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 音视频/视觉语言融合 1 篇

2605.17488 2026-05-19 cs.CV cs.MM cs.SD 62%

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

Omni-Customizer: 用于联合音频-视频生成的端到端多模态定制

Yuheng Chen, Qingdong He, Teng Hu, Yuji Wang, Yabiao Wang, Lizhuang Ma, Jiangning Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract);分类 cs.CV、cs.MM

AI总结 本文提出Omni-Customizer,一种端到端多模态定制框架,旨在实现精确的多模态身份信息绑定和无缝融合,通过引入Omni-Context Fusion模块和Masked TTS Cross-Attention机制,提升多模态定制生成的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 融合架构与评测 3 篇

2605.17336 2026-05-19 cs.RO cs.CV eess.SP 78%

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

基于触觉的多模态融合在具身智能中的应用:视觉、语言和接触驱动范式的综述

Zhixiang Cao, Di Tian, Runwei Guan, Yanzhou Mu, Xiaolou Sun, Shaofeng Liang, Daizong Liu, Tao Huang, Yutao Yue, Henghui Ding, Bin Fang, Alex Zhou, Qing-Long Han, Hui Xiong

机构 * School of Electronic Science and Engineering, Xi’an Jiaotong University, China(西安交通大学电子科学与技术学院) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州)人工智能研究所) State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室) Purple Mountain Laboratory, China(紫金山实验室) Institute for Math & AI, Wuhan University, China(武汉大学数学与人工智能学院) Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University, Australia(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China(北京邮电大学人工智能学院) Institute of Big Data, Fudan University, China(复旦大学大数据研究院) Linkerbot (Beijing) Technology Co., Ltd, China(北京链动科技有限公司) School of Engineering, Swinburne University of Technology, Melbourne(斯威本技术大学工程学院)

专题命中 融合架构与评测 :multimodal fusion(title);分类 cs.CV、eess.SP、cs.RO

AI总结 本文综述了多模态触觉融合在具身智能中的研究,探讨了如何通过整合视觉、语言和触觉信息来提升物理交互与语义推理的结合,提出了一种分层的分类体系,并总结了当前的研究挑战和未来方向。

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17038 2026-05-19 cs.AI 78%

Evidential Information Fusion on Possibilistic Structure

可能性结构上的证据信息融合

Qianli Zhou, Ye Cui, Zhen Li, Witold Pedrycz, Yong Deng

机构 * School of Electronics and Information, Northwestern Polytechnical University(电子信息学院,西北工业大学) Department of Electrical and Computer Engineering, University of Alberta(阿尔伯塔大学电气与计算机工程系) China Mobile Information Technology Center(中国移动信息科技中心) Systems Research Institute, Polish Academy of Sciences(波兰科学院系统研究所) Institute of Fundamental and Frontier Science, University of Electronic Science and Technology of China(中国电子科技大学基础与前沿科学研究院)

专题命中 融合架构与评测 :information fusion(title,abstract)

AI总结 本文提出了一种基于可能性结构的证据信息融合方法,通过引入信任演化网络和三角范数家族,实现了更灵活的信息融合框架,适用于非distinct源融合、冲突管理等复杂场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21941 2026-05-19 cs.LG cs.AI 50%

Robust Multimodal Representation Learning in Healthcare

医疗领域鲁棒多模态表征学习

Xiaoguang Zhu, Linxiao Gong, Lianlong Sun, Yang Liu, Haoyu Wang, Jing Liu

机构 * University of California, Davis(加州大学戴维斯分校) HKUST (GZ)(香港科技大学) University of Rochester(罗切斯特大学) Tongji University(同济大学) Georgia Institute of Technology(佐治亚理工学院) Fudan University(复旦大学) The University of British Columbia(不列颠哥伦比亚大学)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文提出双流特征去相关框架,通过结构因果分析处理医疗多模态数据中的系统性偏差,提升模型泛化能力,实验验证在MIMIC-IV、eICU和ADNI数据集上的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏