arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

共收录 1406 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 融合架构与评测 1406 篇

2606.12294 2026-06-11 cs.CV eess.IV 新提交 62%

Bridging the Modality Gap in Forensic Image Retrieval

弥合法医图像检索中的模态差距

Ricardo González-Gazapo, Annette Morales-González, Yoanna Martínez-Díaz, Heydi Méndez-Vázquez, Milton García-Borroto

机构 * Advanced Technologies Application Center (CENATAV)(先进技术应用中心(CENATAV)) Centro de Sistemas Complejos, Facultad de Física, Universidad de La Habana(哈瓦那大学物理学院复杂系统中心)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、eess.IV

AI总结 提出统一检索框架,利用多模态大语言模型生成文本描述并结合视觉与文本特征融合,提升纹身、人脸素描等法医任务的检索精度与鲁棒性。

Comments 23 pages, 5 figures, paper submitted to Elsevier journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09203 2026-05-20 cs.CV cs.RO 62%

3D Modeling and Automated Measurement of Concrete Cracks via Segment Anything Refinement and Visual Inertial LiDAR Fusion

通过段落任何精修和视觉惯性LiDAR融合进行混凝土裂缝的3D建模与自动测量

Pengru Deng, Jiapeng Yao, Chun Li, Su Wang, Xinrun Li, Varun Ojha, Xuhui He

机构 * School of Civil Engineering(土木工程学院) Central South University(中南大学) Hunan Provincial Key Laboratory for Disaster Prevention and Mitigation of Rail Transit Engineering Structures(湖南省铁路工程结构灾害预防与 mitigation 工程结构重点实验室) Nvidia School of Computing(计算学院) Newcastle University(新castle大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、cs.RO

AI总结 本文提出了一种结合计算机视觉技术和多模态同时定位与建图(SLAM)的创新框架,用于二维裂缝检测、三维重建和三维自动裂缝测量,解决了现有方法在适应性和鲁棒性方面的不足,特别是在处理曲线或复杂几何形状时的挑战。

Comments Title and author list updated

Journal ref Computer-Aided Civil and Infrastructure Engineering, Volume 45, 2026, 100019, ISSN 1093-9687

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00826 2026-05-05 cs.IR cs.CV cs.LG cs.MM 62%

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

理解文本到视频检索中的性能平台:全面的实证和语言分析

Maria-Eirini Pegia, Dimitrios Stefanopoulos, Björn Þór Jónsson, Anastasia Moumtzidou, Ilias Gialampoukidis, Stefanos Vrochidis, Ioannis Kompatsiaris

机构 * Information Technologies Institute(信息科技研究所) CERTH-ITI School of Computer Science(计算机科学学院) Reykjavík University(雷克雅未克大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

AI总结 本文通过实证和语言分析,探讨文本到视频检索中模型性能瓶颈,揭示caption特性与模型表现的关系,指出简单清晰的caption提升召回率,而复杂事件仍具挑战性。

Comments Survey, 50 pages, 15 figures, 13 tables, 154 citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23325 2026-04-28 cs.CV cs.AI eess.IV 62%

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence

EAD-Net:具有空间细化和时间一致性的情感感知说话头生成

Yahui Li, Yinfeng Yu, Liejun Wang, Shengjie Shen

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、eess.IV

AI总结 EAD-Net通过引入同步网络监督和时空方向注意力机制,提升情感感知说话头生成的唇同步精度和时间一致性,同时利用大语言模型增强情感语义控制。

Comments Main paper (10 pages). Accepted for publication by ICMR(International Conference on Multimedia Retrieval) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11967 2026-03-31 cs.CV cs.AI cs.RO 62%

Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions

守护天空:反无人机方法的全面调查、基准测试与未来方向

Yifei Dong, Fengyi Wu, Sanjian Zhang, Guangyu Chen, Yuzhi Hu, Masumi Yano, Jingdong Sun, Siyu Huang, Feng Liu, Qi Dai, Zhi-Qi Cheng

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、cs.RO

AI总结 本文综述了反无人机领域,探讨了分类、检测与跟踪等核心方法,分析了数据合成、多模态融合等新兴技术,并指出实时性能、 stealth 检测等领域的不足,旨在推动下一代防御策略的发展。

Comments Accepted to CVPR 2025 Anti-UAV Workshop (Best Paper Award), 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17623 2026-03-04 cs.MM cs.CV 62%

Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?

合成感知:生成图像能否解锁潜在的视觉先验以用于以文本为中心的推理?

Yuesheng Huang, Peng Zhang, Xiaoxin Wu, Riliang Liu, Jiaqi Liang

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

AI总结 本文探讨生成图像能否解锁潜在视觉先验以提升文本中心推理,通过多模态融合架构和提示工程策略实现性能提升。

Comments Accepted as a poster at the International Conference on Machine Learning (ICML 2025) NewInML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16479 2026-03-03 eess.IV cs.AI cs.CV 62%

Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization

解耦的多模态学习:组织学与转录组学用于癌症表征

Yupei Zhang, Xiaofei Wang, Anran Liu, Lequan Yu, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Health Technology & Informatics, The Hong Kong Polytechnic University(香港理工大学健康科技与信息学系) Department of Statistics and Actuarial Science, The University of Hong Kong(香港大学统计与精算科学系) Department of Clinical Neurosciences and Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学临床神经科学系和应用数学与理论物理系;邓迪大学科学与工程学院和医学院) School of Science and Engineering and School of Medicine, University of Dundee, UK

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、eess.IV

AI总结 本文提出了解耦的多模态学习框架,通过分解组织学和转录组数据以提高癌症表征的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19418 2026-02-24 cs.CV cs.AI eess.IV 62%

DEFNet: Multitasks-based Deep Evidential Fusion Network for Blind Image Quality Assessment

DEFNet:基于多任务的深度证据融合网络用于盲图像质量评估

Yiwei Lou, Yuanpeng He, Rongchao Zhang, Yongzhi Cao, Hanpin Wang, Yu Huang

机构 * Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education(高可信软件技术重点实验室(北京大学)教育部) School of Computer Science, Peking University(计算机学院(北京大学))

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、eess.IV

AI总结 DEFNet通过多任务优化和证据学习改进盲图像质量评估,提升鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05208 2026-02-23 eess.IV cs.CV 62%

Context-Aware Asymmetric Ensembling for Interpretable Retinopathy of Prematurity Screening via Active Query and Vascular Attention

具备上下文感知的非对称集成用于通过主动查询和血管注意力的可解释早产视网膜病变筛查

Md. Mehedi Hassan, Taufiq Hasan

机构 * m-Health Lab, Department of Biomedical Engineering, Bangladesh University of Engineering(m-Health实验室,生物医学工程系,孟加拉国工程大学) Center for Bioengineering Innovation(生物工程创新中心) Design, Department of Biomedical Engineering, Johns Hopkins University, Baltimore, MD, USA(设计,生物医学工程系,约翰霍普金斯大学,巴尔的摩,MD,美国)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、eess.IV

AI总结 本文提出具备上下文感知的非对称集成模型,通过主动查询和血管注意力技术,实现早产视网膜病变筛查的高准确率和可解释性。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06363 2026-02-09 eess.IV cs.CV 62%

Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation

Mamba Goes HoME: 分层软专家混合模型用于3D医学图像分割

Szymon Płotka, Gizem Mert, Maciej Chrabaszcz, Ewa Szczurek, Arkadiusz Sitek

机构 * Faculty of Mathematics, Informatics, and Mechanics, University of Warsaw(华沙大学数学、信息学与力学系) Faculty of Mathematics and Computer Science, Jagiellonian University(雅盖隆大学数学与计算机科学系) Institute of AI for Health, Helmholtz Munich(海德堡医学院人工智能与健康研究所) Faculty of Electronics and Information Technology, Warsaw University of Technology(华沙理工大学电子与信息技术系) NASK - National Research Institute(国家研究 institute) Faculty of Radiology, Massachusetts General Hospital(麻省总医院放射学系) Department of Radiology, Harvard Medical School(哈佛医学院放射学系)

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、eess.IV

AI总结 本文提出分层软专家混合模型HoME,通过两级令牌路由提升3D医学图像分割的长上下文建模能力,实现更高效的局部和全局特征提取,从而提升分割性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17045 2026-01-29 cs.CV cs.AI cs.MM 62%

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

RacketVision: 一种多 racket 运动基准数据集用于统一的球和 racket 分析

Linfeng Dong, Yuchen Yang, Hao Wu, Wei Wang, Yuenan Hou, Zhihang Zhong, Xiao Sun

机构 * Shanghai AI Laboratory(上海人工智能实验室)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、cs.MM

AI总结 RacketVision 是一个用于多 racket 运动分析的基准数据集,通过细粒度注释和 CrossAttention 机制提升球轨迹预测性能。

Comments Accepted to AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14961 2026-01-28 cs.CV cs.SD eess.AS eess.IV 62%

Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities

自适应多模态人物识别:一种处理缺失模态的稳健框架

Aref Farhadipour, Teodora Vukovic, Volker Dellwo, Petr Motlicek, Srikanth Madikeri

机构 * Department of Computational Linguistics, University of Zurich(苏黎世大学计算语言学系) Faculty of Information Technology, Brno University of Technology(布拉格技术大学信息学院) Idiap Research Institute(Idiap研究机构)

专题命中 融合架构与评测 :hybrid fusion(abstract);分类 cs.CV、eess.IV

AI总结 本文提出一种自适应多模态人物识别框架,通过融合上半身运动、面孔和语音信息,提升在缺失模态下的识别准确率,达到99.51%的Top-1准确率。

Comments 9 pages and 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08240 2026-01-14 eess.IV cs.CV 62%

Temporal-Enhanced Interpretable Multi-Modal Prognosis and Risk Stratification Framework for Diabetic Retinopathy (TIMM-ProRS)

时间增强的可解释多模态预后和风险分层框架用于糖尿病视网膜病变(TIMM-ProRS)

Susmita Kar, A S M Ahsanul Sarkar Akib, Abdul Hasib, Samin Yaser, Anas Bin Azim

机构 * Department of Robotics, Robo Tech Valley, Dhaka, Bangladesh(机器人系,罗布科技谷,达卡,孟加拉国)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、eess.IV

AI总结 TIMM-ProRS通过融合视网膜图像和时间生物标志物,实现糖尿病视网膜病变的多模态和时间动态分析,达到97.8%的准确率,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17965 2025-11-25 cs.CV cs.MM 62%

Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification

信号:选择性交互与全局-局部对齐用于多模态对象重识别

Yangyang Liu, Yuhao Wang, Pingping Zhang

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、cs.MM

AI总结 Signal通过选择性交互和全局-局部对齐框架提升多模态对象重识别的性能。

Comments Accepted by AAAI2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22829 2025-10-28 cs.CV cs.AI cs.MM 62%

LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction

Aleksandar Pramov

机构 * Georgia Institute of Technology, USA(佐治亚理工学院)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17305 2025-10-22 cs.CV cs.MM 62%

LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding

ZhaoYang Han, Qihan Lin, Hao Liang, Bowen Chen, Zhou Liu, Wentao Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、cs.MM

Comments Submitted to ARR Rolling Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05839 2025-10-15 cs.MM cs.CV 62%

Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality

Hengyang Zhou, Yiwei Wei, Jian Yang, Zhenyu Zhang

机构 * Nanjing University(南京大学) China University of Petroleum(中国石油大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01428 2025-10-13 cs.CV eess.IV 62%

DiffMark: Diffusion-based Robust Watermark Against Deepfakes

Chen Sun, Haiyang Sun, Zhiqing Guo, Yunfeng Diao, Liejun Wang, Dan Ma, Gaobo Yang, Keqin Li

机构 * College of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Silk Road Multilingual Cognitive Computing International Cooperation Joint Laboratory(丝绸之路多语种认知计算国际合作联合实验室) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) Department of Computer Science, State University of New York(纽约州立大学计算机科学系)

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18461 2025-09-24 cs.GR cs.AI cs.CV cs.MM 62%

Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?

Ayan Sar, Sampurna Roy, Tanupriya Choudhury, Ajith Abraham

机构 * School of Computer Sciences, University of Petroleum and Energy Studies (UPES), Dehradun(计算机科学学院,石油与能源研究大学(UPES),德里加恩)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

Comments Published in Foundations and Trends in Signal Processing (#1 in Signal Processing, #3 in Computer Science)

Journal ref Foundations and Trends in Signal Processing (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09182 2025-08-14 eess.IV cs.CV 62%

MedPatch: Confidence-Guided Multi-Stage Fusion for Multimodal Clinical Data

Baraa Al Jorf, Farah Shamout

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03752 2025-08-07 eess.IV cs.AI cs.CV 62%

M$^3$HL: Mutual Mask Mix with High-Low Level Feature Consistency for Semi-Supervised Medical Image Segmentation

Yajun Liu, Zenghui Zhang, Jiang Yue, Weiwei Guo, Dongying Li

机构 * Shanghai Key Laboratory of Intelligent Sensing and Recognition(上海智能感知与识别重点实验室) Shanghai Jiao Tong University(上海交通大学) Department of Endocrinology and Metabolism, Renji Hospital, School of Medicine, Shanghai Jiao Tong University(内分泌与代谢科,仁济医院,医学院,上海交通大学) Center for Digital Innovation, Tongji University(同济大学数字创新中心)

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、eess.IV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18844 2025-06-24 cs.RO cs.CV 62%

Reproducible Evaluation of Camera Auto-Exposure Methods in the Field: Platform, Benchmark and Lessons Learned

Olivier Gamache, Jean-Michel Fortin, Matěj Boxan, François Pomerleau, Philippe Giguère

专题命中 融合架构与评测 :multi-exposure(abstract);分类 cs.CV、cs.RO

Comments 19 pages, 11 figures, pre-print version of the accepted paper for IEEE Transactions on Field Robotics (T-FR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10452 2025-06-13 cs.CV cs.CL cs.LG cs.MM 62%

Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts

Guowei Zhong, Ruohong Huan, Mingzhen Wu, Ronghua Liang, Peng Chen

机构 * College of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

Comments Submitted to TAC. The code is available at https://github.com/gw-zhong/CIDer

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06118 2025-06-11 eess.IV cs.AI cs.CV 62%

The Application of Deep Learning for Lymph Node Segmentation: A Systematic Review

Jingguo Qu, Xinyang Han, Man-Lik Chui, Yao Pu, Simon Takadiyi Gunda, Ziman Chen, Jing Qin, Ann Dorothy King, Winnie Chiu-Wing Chu, Jing Cai, Michael Tin-Cheung Ying

机构 * Department of Health Technology and Informatics, The Hong Kong Polytechnic University(健康科技与信息学系,香港理工大学) Centre for Smart Health and School of Nursing, The Hong Kong Polytechnic University(智能健康中心及护理学院,香港理工大学) Department of Imaging and Interventional Radiology, The Chinese University of Hong Kong(影像与介入放射科系,中国香港大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13309 2025-05-09 eess.IV cs.AI cs.CV 62%

Integrating AI for Human-Centric Breast Cancer Diagnostics: A Multi-Scale and Multi-View Swin Transformer Framework

Farnoush Bayatmakou, Reza Taleei, Milad Amir Toutounchian, Arash Mohammadi

机构 * Concordia Institute for Information Systems Engineering (CIISE)(康科迪亚信息系统工程研究所) Concordia University(康科迪亚大学) Thomas Jefferson University Hospital(泰勒斯·杰弗里斯大学医院) College of Computing & Informatics(计算与信息学院) Drexel University(德雷塞尔大学)

专题命中 融合架构与评测 :hybrid fusion(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00747 2025-05-05 cs.OH cs.CV cs.MA cs.RO 62%

Wireless Communication as an Information Sensor for Multi-agent Cooperative Perception: A Survey

Zhiying Song, Tenghui Xie, Fuxi Wen, Jun Li

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09491 2025-03-13 cs.CV eess.IV 62%

DAMM-Diffusion: Learning Divergence-Aware Multi-Modal Diffusion Model for Nanoparticles Distribution Prediction

Junjie Zhou, Shouju Wang, Yuxia Tang, Qi Zhu, Daoqiang Zhang, Wei Shao

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV、eess.IV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17939 2025-02-26 cs.MM cs.CV 62%

Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point Clouds

Yun Zhang, Zixi Guo, Linwei Zhu, C. -C. Jay Kuo

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15656 2024-11-28 cs.CV cs.RO 62%

SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map Generation

Hao Dong, Weihao Gu, Xianjing Zhang, Jintao Xu, Rui Ai, Huimin Lu, Juho Kannala, Xieyuanli Chen

专题命中 融合架构与评测 :feature-level fusion(abstract);分类 cs.CV、cs.RO

Comments ICRA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01869 2024-10-30 eess.IV cs.CV cs.LG 62%

Let it shine: Autofluorescence of Papanicolaou-stain improves AI-based cytological oral cancer detection

Wenyi Lian, Joakim Lindblad, Christina Runow Stark, Jan-Michaél Hirsch, Nataša Sladoje

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、eess.IV

Comments 16 pages, 12 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏