arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

共收录 1405 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 融合架构与评测 1405 篇

0907.2949 2009-12-01 cs.DC 71%

Distributed anonymous function computation in information fusion and multiagent systems

Julien M. Hendrickx, Alex Olshevsky, John N. Tsitsiklis

专题命中 融合架构与评测 :information fusion(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0604042 2009-12-01 cs.AI 71%

Adaptative combination rule and proportional conflict redistribution rule for information fusion

M. C. Florea, J. Dezert, P. Valin, F. Smarandache, Anne-Laure Jousselme

专题命中 融合架构与评测 :information fusion(title)

Comments Presented at Cogis '06 Conference, Paris, March 2006

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16154 2026-07-20 cs.CV cs.SY eess.SY 新提交 70%

CLIFE: Camera-LiDAR Fusion Framework for Edge-Deployable Roadside VRU Perception

CLIFE:用于边缘可部署路边VRU感知的相机-激光雷达融合框架

Tam Bang, Hoang H. Nguyen, Lei Cheng, Lihao Guo, Siyang Cao, Hussam Abubakr, Tianya Zhang, Austin Harris, Mina Sartipi

专题命中 融合架构与评测 :sensor fusion(abstract);multi-sensor fusion(abstract);分类 cs.CV

AI总结 针对路边VRU感知难题,提出CLIFE框架,集成无目标在线校准与轻量级后期融合跟踪于单嵌入式设备,无需云卸载。经实验验证,该框架提升了传感器感知能力与鲁棒性,核心运行高效,为下游安全应用提供基础并减少相关开销。

Comments Accepted for publication at the 2026 IEEE 29th International Conference on Intelligent Transportation Systems (ITSC 2026). 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14821 2026-07-17 cs.CV 新提交 70%

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

模糊模态边界:从单模态到多模态行人重识别的统一综述

Xiao Wang, Bing Wang, Bin Yang, Cuiqun Chen, Xin Xu, Mang Ye

机构 * Wuhan University of Science and Technology(武汉科技大学) Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System(湖北省智能信息处理与实时工业系统重点实验室) Wuhan University(武汉大学) Anhui University(安徽大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);visible-infrared(abstract);分类 cs.CV

AI总结 综述行人重识别领域从单模态向多模态发展的转变,系统回顾关键跨模态任务,研究多模态融合ReID,提出基于Transformer的可见光-红外ReID基线框架,并概述未来研究方向。

Comments 61 pages, 6 figures, journal

Journal ref International Journal of Computer Vision.2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06222 2026-07-08 cs.RO 新提交 70%

APVI-SLAM: Real-Time Acoustic-Pressure-Visual-Inertial Localization and Photorealistic Mapping System in Complex Underwater Environment

APVI-SLAM:复杂水下环境中的实时声压视觉惯性定位与真实感映射系统

Hanwen Zhang, Yipeng Zhu, Xiaopeng Guo, Huajian Huang, Sai-Kit Yeung

机构 * Hong Kong University of Science and Technology(香港科技大学) Beijing Institute of Technology(北京理工大学)

专题命中 融合架构与评测 :sensor fusion(abstract);multi-sensor fusion(abstract);分类 cs.RO

AI总结 针对水下视觉惯性SLAM在极端环境中的问题,提出APVI-SLAM系统,通过可靠性感知定位框架和滑动窗口冻结策略增强鲁棒性,用四叉树引导映射模块助力重建,还贡献数据集,实验证明其实现了实时领先的定位与重建质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02894 2026-06-04 cs.CV 70%

Tiny Collaborative Inference for Occlusion-Robust Object Detection

用于遮挡鲁棒目标检测的微型协同推理

Chieh-Tung Cheng, Mustafa Aslanov, Eiman Kanjo

机构 * Imperial College London(帝国理工学院伦敦分校) Nottingham Trent University(诺丁汉特伦特大学)

专题命中 融合架构与评测 :feature-level fusion(abstract);decision-level fusion(abstract);分类 cs.CV

AI总结 针对超低端边缘设备,结合MCUNet骨干网络、YOLOv2检测头和TensorFlow Lite量化,评估决策级融合(WBF)相比特征级融合在遮挡场景下提升mAP达+0.2736,并验证了多视角融合与Wi-Fi对等部署的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17049 2026-03-19 cs.CV cs.AI 70%

From Geometric Mimicry to Comprehensive Generation: A Context-Informed Multimodal Diffusion Model for Urban Morphology Synthesis

从几何模仿到综合生成:一种基于上下文的多模态扩散模型用于城市形态合成

Fangshuo Zhou, Huaxia Li, Liuchang Xu, Rui Hu, Sensen Wu, Liang Xu, Hailin Feng, Zhenhong Du

机构 * Zhejiang Agriculture and Forestry University(浙江农业与林业大学) Alibaba Group(阿里巴巴集团) Zhejiang University(浙江大学) Zhejiang University of Technology(浙江工业大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);information fusion(abstract);分类 cs.CV

AI总结 本文提出ControlCity模型,通过多模态信息融合实现城市形态综合生成,提升了形态保真度和空间重叠度,实现了跨城市风格迁移和未知城市零样本生成。

Comments Accepted

Journal ref International Journal of Geographical Information Science (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01974 2026-02-03 eess.SP 70%

Obstacle Detection at Level Crossings under Adverse Weather Conditions -- A Survey

恶劣天气条件下铁路道口障碍物检测——综述

Chenyang Yan, Mats Bengtsson

专题命中 融合架构与评测 :sensor fusion(abstract);multi-sensor fusion(abstract);分类 eess.SP

AI总结 本文综述了恶劣天气下铁路道口障碍物检测的传感器技术和融合策略,分析了各传感器的优缺点及应对措施,提出未来研究方向以提升检测系统的可靠性和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03724 2025-08-07 cs.CV 70%

From Waveforms to Pixels: A Survey on Audio-Visual Segmentation

Jia Li, Yapeng Tian

机构 * Department of Computer Science, The University of Texas at Dallas(计算机科学系,德克萨斯大学达拉斯分校)

专题命中 融合架构与评测 :multimodal fusion(abstract);audio-visual fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08364 2025-07-14 cs.RO 70%

Towards Robust Sensor-Fusion Ground SLAM: A Comprehensive Benchmark and A Resilient Framework

Deteng Zhang, Junjie Zhang, Yan Sun, Tao Li, Hao Yin, Hongzhao Xie, Jie Yin

机构 * Independent(独立研究者) Chongqing University(重庆大学) Nankai University(南开大学) Zhejiang University of Technology(浙江工业大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 融合架构与评测 :sensor fusion(abstract);multi-sensor fusion(abstract);分类 cs.RO

Comments This paper has already been accepted to IROS2025. 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06370 2025-03-27 cs.CV 70%

Robust Multiview Multimodal Driver Monitoring System Using Masked Multi-Head Self-Attention

Yiming Ma, Victor Sanchez, Soodeh Nikan, Devesh Upadhyay, Bhushan Atote, Tanaya Guha

专题命中 融合架构与评测 :feature-level fusion(abstract);decision-level fusion(abstract);分类 cs.CV

Comments 9 pages (1 for reference); accepted by the 6th Multimodal Learning and Applications Workshop (MULA) at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08984 2024-03-15 cs.RO cs.AI cs.MA 70%

Safe Road-Crossing by Autonomous Wheelchairs: a Novel Dataset and its Experimental Evaluation

Carlo Grigioni, Franca Corradini, Alessandro Antonucci, Jérôme Guzzi, Francesco Flammini

专题命中 融合架构与评测 :sensor fusion(abstract);multi-sensor fusion(abstract);分类 cs.RO

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01358 2023-10-03 cs.CV 70%

NEUCORE: Neural Concept Reasoning for Composed Image Retrieval

Shu Zhao, Huijuan Xu

专题命中 融合架构与评测 :multimodal fusion(abstract);information fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05032 2023-09-12 cs.CV 70%

Unified Contrastive Fusion Transformer for Multimodal Human Action Recognition

Kyoung Ok Yang, Junho Koh, Jun Won Choi

专题命中 融合架构与评测 :multimodal fusion(abstract);information fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.14758 2021-01-01 cs.CV cs.AI cs.IT math.IT 70%

Deep Hashing for Secure Multimodal Biometrics

Veeru Talreja, Matthew Valenti, Nasser Nasrabadi

专题命中 融合架构与评测 :multimodal fusion(abstract);feature-level fusion(abstract);分类 cs.CV

Journal ref IEEE Transactions on Information Forensics and Security,vol.16,pp.1306-1321,2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07059 2020-01-22 cs.CV cs.CC 70%

Accuracy vs. Complexity: A Trade-off in Visual Question Answering Models

Moshiur R. Farazi, Salman H. Khan, Nick Barnes

专题命中 融合架构与评测 :multimodal fusion(abstract);multi-modal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.1902 2015-02-04 cs.CV 70%

Quality-based Multimodal Classification Using Tree-Structured Sparsity

Soheil Bahrampour, Asok Ray, Nasser M. Nasrabadi, Kenneth W. Jenkins

专题命中 融合架构与评测 :information fusion(abstract);feature-level fusion(abstract);分类 cs.CV

Comments To Appear in 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2014)

Journal ref CVPR 2014, pp. 4114 - 4121

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27886 2026-06-29 cs.LG 新提交 67%

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

HARMES数据集上多模态人类活动识别融合技术的比较

Ahmed Mohamady, Robin Burchard, Kristof Van Laerhoven

机构 * University of Siegen(锡根大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);sensor fusion(abstract)

AI总结 系统比较七种传感器融合方法在HARMES多模态数据集上的性能,发现门控多模态融合在留一参与者评估中宏F1分数达0.82,优于拼接式后期融合基线(0.76)。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01495 2026-05-05 cs.CL cs.AI 67%

FT-RAG: A Fine-grained Retrieval-Augmented Generation Framework for Complex Table Reasoning

FT-RAG:一种面向复杂表格推理的细粒度检索增强生成框架

Zebin Guo, Weidong Geng, Ruichen Mao

机构 * Georgia Institute of Technology(佐治亚理工学院) Zhejiang University(浙江大学) Zhejiang Lab(浙江实验室)

专题命中 融合架构与评测 :multi-modal fusion(abstract);information fusion(abstract)

AI总结 FT-RAG通过细粒度检索和多模态融合提升表格推理能力,引入多表RAG库增强训练数据,实现表格和单元格命中率显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01447 2026-02-03 cs.CL cs.AI 67%

SentiFuse: Deep Multi-model Fusion Framework for Robust Sentiment Extraction

SentiFuse:一种用于鲁棒情感提取的深度多模型融合框架

Hieu Minh Duong, Rupa Ghosh, Cong Hoan Nguyen, Eugene Levin, Todd Gary, Long Nguyen

机构 * University of Louisville(路易斯维尔大学) Meharry Medical College(梅哈里医学院)

专题命中 融合架构与评测 :feature-level fusion(abstract);decision-level fusion(abstract)

AI总结 SentiFuse通过多模型融合框架提升情感提取的鲁棒性和准确性,特征级融合在F1分数上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00107 2026-02-03 cs.CV cs.RO eess.IV 67%

Efficient UAV trajectory prediction: A multi-modal deep diffusion framework

高效无人机轨迹预测:一种多模态深度扩散框架

Yuan Gao, Xinyu Guo, Wenjing Xie, Zifan Wang, Hongwen Yu, Gongyang Li, Shugong Xu

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、eess.IV、cs.RO

AI总结 本文提出一种多模态深度融合框架,通过融合激光雷达和毫米波雷达数据提升无人机轨迹预测精度,实验显示其比基线模型提升40%。

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17934 2025-07-25 cs.LG cs.AI 67%

Multimodal Fine-grained Reasoning for Post Quality Evaluation

Xiaoxu Guo, Siyan Liang, Yachao Cui, Juxiang Zhou, Lei Wang, Han Cao

机构 * School of Computer Science(计算机科学学院) School of Artificial Intelligence Institute(人工智能研究所) State Key Laboratory of Information Security, Institute of Information Engineering(信息安全国家重点实验室,信息工程研究所) Key Laboratory of Education Informatization for Nationalities(民族教育信息化重点实验室)

专题命中 融合架构与评测 :multimodal fusion(abstract);information fusion(abstract)

Comments 48 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09030 2025-01-17 physics.ins-det 67%

Determination and evaluation of the critical liquid nitrogen for superconducting levitator based on a novel temperature-weight coupling measurement device

Peng Pang, Jun Zheng, Chenling Xian

专题命中 融合架构与评测 :information fusion(abstract);sensor fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02033 2024-08-06 cs.CV cs.LG cs.MM eess.IV 67%

Enhancing Human Action Recognition and Violence Detection Through Deep Learning Audiovisual Fusion

Pooya Janani, Amirabolfazl Suratgar, Afshin Taghvaeipour

专题命中 融合架构与评测 :hybrid fusion(abstract);分类 cs.CV、eess.IV、cs.MM

Comments This work has been submitted to the IEEE for possible publication, 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16923 2024-04-12 cs.CV cs.RO eess.IV 67%

Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation

Ruiping Liu, Jiaming Zhang, Kunyu Peng, Yufan Chen, Ke Cao, Junwei Zheng, M. Saquib Sarfraz, Kailun Yang, Rainer Stiefelhagen

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、eess.IV、cs.RO

Comments Accepted to IEEE IV 2024. The source code is publicly available at https://github.com/RuipingL/MISS

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11303 2018-10-29 cs.IR 67%

Investigating non-classical correlations between decision fused multi-modal documents

Dimitris Gkoumas, Sagar Uprety, Dawei Song

专题命中 融合架构与评测 :multimodal fusion(abstract);multi-modal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07574 2026-08-11 cs.CV cs.RO 新提交 62%

Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion

基于Swin Transformer与临床元数据融合的多模态皮肤病变分类

Nethmi Pathirana, Isuru Munasinghe, Dileeka Alwis

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.RO

AI总结 该研究针对皮肤病变分类的自动化挑战,提出结合Swin Transformer图像特征与临床元数据的多模态框架,在公开数据集上实现92.55%测试准确率,兼具强少数类性能与预测可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22773 2026-07-28 eess.IV cs.CV physics.med-ph 新提交 62%

Metric Surface Reconstruction of Neurosurgical Scenes from Monocular Operating Microscope Images and Microscope Pose

从单目手术显微镜图像和显微镜位姿重建神经外科场景的度量曲面

Thomas Bucher, Didier Neuenschwander, Thomas Petutschnigg, Michael Murek, David Bervini, Andreas Raabe, Manuela Eugster

专题命中 融合架构与评测 :image fusion(abstract);分类 cs.CV、eess.IV

AI总结 研究能否从单目手术显微镜图像和位姿数据重建神经外科场景的三维几何度量,利用预训练模型估计深度、泊松曲面重建点云成网格,结果显示该方法在模型设置中有技术可行性,支持相关手术技术进一步发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15320 2026-07-07 q-bio.QM cs.CV cs.LG cs.MM q-bio.GN 版本更新 62%

GestaltMML: Enhancing Rare Genetic Disease Diagnosis through Multimodal Machine Learning Combining Facial Images and Clinical Text

GestaltMML:通过结合面部图像和临床文本的多模态机器学习增强罕见遗传病诊断

Da Wu, Zhanliang Wang, Hongzhuo Chen, Jingye Yang, Cong Liu, Tzung-Chien Hsieh, Elaine Marchi, Justin Blair, Peter Krawitz, Chunhua Weng, Wendy Chung, Gholson J. Lyon, Ian D. Krantz, Jennifer M. Kalish, Kai Wang

机构 * Raymond G. Perelman Center for Cellular and Molecular Therapeutics, Children’s Hospital of Philadelphia(雷蒙德·G·佩尔曼细胞与分子治疗中心,费城儿童医院) Department of Mathematics, University of Pennsylvania(数学系,宾夕法尼亚大学) Department of Biomedical Informatics, Columbia University Irving Medical Center(生物医学信息学系,哥伦比亚大学伊万斯医疗中心) Department of Human Genetics, New York State Institute for Basic Research in Developmental Disabilities, Staten Island, NY, USA(人类遗传学系,纽约州发育障碍基础研究机构,纽约州史泰登岛) Division of Human Genetics, Children’s Hospital of Philadelphia(人类遗传学部,费城儿童医院) Department of Pediatrics, Boston Children’s Hospital, Harvard Medical School(儿科系,波士顿儿童医院,哈佛医学院) Biology PhD Program, The Graduate Center, The City University of New York(生物学博士项目,纽约市立大学研究生中心) Department of Genetics, Perelman School of Medicine, University of Pennsylvania(遗传学系,宾夕法尼亚大学佩尔曼医学学院) Department of Pediatrics, Perelman School of Medicine, University of Pennsylvania(儿科系,宾夕法尼亚大学佩尔曼医学学院) Department of Pathology and Laboratory Medicine, Perelman School of Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学佩尔曼医学学院)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

AI总结 研究针对罕见遗传病诊断难题,提出基于Transformer架构的多模态机器学习方法GestaltMML,整合面部图像、人口统计学信息和临床笔记,提升预测准确性,缩小诊断差距。

Comments Preprint updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21258 2026-06-23 cs.RO cs.CV 新提交 62%

Spectral GS-SLAM: Observability-Aware, Degeneracy-Robust Tracking for Real-Time 3D Gaussian Splatting SLAM

Spectral GS-SLAM:面向实时3D高斯泼溅SLAM的可观测性感知与退化鲁棒跟踪

Edward Beng Wai Tan, Siew-Kei Lam, Dongshuo Zhang

机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、cs.RO

AI总结 提出Spectral GS-SLAM,通过自适应补偿退化场景中的欠约束方向,结合ICP与特征约束实现鲁棒跟踪,并利用高斯感知平面性加权机制融合几何信息,在无结构/纹理环境中保持实时性能(40.14 FPS)。

Comments This work has been accepted to IROS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏