arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2026-07-07 至 2026-07-07 共收录 8 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 融合架构与评测 8 篇

2607.05311 2026-07-07 cs.CV 新提交 79%

Deep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical Translation

男性不育症精液分析深度学习技术:计算机视觉、多模态融合与临床转化

Runwei Guan, Shaofeng Liang, Jiacheng Weng, Xiaoyi Gu, Jia Weng, Daizong Liu, Duo Pan, Qingxin Zhang, Xiao Liang, Weiping Ding, Suoyu Zhu, Ming Yuan, Yanhua Fei

机构 * Department of Gynaecology and Obstetrics, The Affiliated Jiangyin Hospital of Nantong University(南通大学附属江阴医院妇产科) Department of Oncology, the Affiliated Jiangyin Hospital of Nantong University(南通大学附属江阴医院肿瘤科) Thrust of AI, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能推力实验室) FertiTech AI(生殖科技人工智能公司) Department of Oncology, Suzhou Xiangcheng People’s Hospital(苏州市相城区人民医院肿瘤科) Department of Biological Sciences and Bioinformatics, School of Science, Xi’an Jiaotong-Liverpool University(西交利物浦大学理学院生物科学与生物信息学系) School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院)

专题命中 融合架构与评测 :multimodal fusion(title,abstract);分类 cs.CV

AI总结 针对传统精液分析的主观性强等缺陷,综述聚焦计算机视觉与深度学习驱动的精子分析技术,梳理方法、数据集与落地障碍,提出分阶段临床转化路线。

Comments 46 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04786 2026-07-07 cs.SE cs.AI cs.MA 新提交 71%

An Exploration of Agentic Information Fusion for Test Maintenance Prediction

用于测试维护预测的智能信息融合探索

Jingxiong Liu, Nasser Mohammadiha, Gregory Gay

机构 * Chalmers University of Technology and University of Gothenburg(楚德斯技术大学和戈特堡大学) Ericsson AB(爱立信有限公司) University of Gothenburg(戈特堡大学)

专题命中 融合架构与评测 :information fusion(title)

AI总结 研究针对测试维护预测问题,提出多智能体框架MAST,通过智能融合多种分析方法及后检查程序,在实际场景中评估,提升了测试维护预测水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15320 2026-07-07 q-bio.QM cs.CV cs.LG cs.MM q-bio.GN 版本更新 62%

GestaltMML: Enhancing Rare Genetic Disease Diagnosis through Multimodal Machine Learning Combining Facial Images and Clinical Text

GestaltMML:通过结合面部图像和临床文本的多模态机器学习增强罕见遗传病诊断

Da Wu, Zhanliang Wang, Hongzhuo Chen, Jingye Yang, Cong Liu, Tzung-Chien Hsieh, Elaine Marchi, Justin Blair, Peter Krawitz, Chunhua Weng, Wendy Chung, Gholson J. Lyon, Ian D. Krantz, Jennifer M. Kalish, Kai Wang

机构 * Raymond G. Perelman Center for Cellular and Molecular Therapeutics, Children’s Hospital of Philadelphia(雷蒙德·G·佩尔曼细胞与分子治疗中心,费城儿童医院) Department of Mathematics, University of Pennsylvania(数学系,宾夕法尼亚大学) Department of Biomedical Informatics, Columbia University Irving Medical Center(生物医学信息学系,哥伦比亚大学伊万斯医疗中心) Department of Human Genetics, New York State Institute for Basic Research in Developmental Disabilities, Staten Island, NY, USA(人类遗传学系,纽约州发育障碍基础研究机构,纽约州史泰登岛) Division of Human Genetics, Children’s Hospital of Philadelphia(人类遗传学部,费城儿童医院) Department of Pediatrics, Boston Children’s Hospital, Harvard Medical School(儿科系,波士顿儿童医院,哈佛医学院) Biology PhD Program, The Graduate Center, The City University of New York(生物学博士项目,纽约市立大学研究生中心) Department of Genetics, Perelman School of Medicine, University of Pennsylvania(遗传学系,宾夕法尼亚大学佩尔曼医学学院) Department of Pediatrics, Perelman School of Medicine, University of Pennsylvania(儿科系,宾夕法尼亚大学佩尔曼医学学院) Department of Pathology and Laboratory Medicine, Perelman School of Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学佩尔曼医学学院)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV、cs.MM

AI总结 研究针对罕见遗传病诊断难题,提出基于Transformer架构的多模态机器学习方法GestaltMML,整合面部图像、人口统计学信息和临床笔记,提升预测准确性,缩小诊断差距。

Comments Preprint updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03111 2026-07-07 eess.SP 新提交 57%

Edge-Assisted Multimodal UAV Localization with Resource-Efficient Compression and Robust Fusion

基于资源高效压缩和鲁棒融合的边缘辅助多模态无人机定位

Zhong Ye, Yinghui He, Guanding Yu, Pavel Loskot

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 eess.SP

AI总结 设计多模态无人机定位框架,利用相机、LiDAR和雷达传感模态。通过信息瓶颈压缩等模块应对传感节点资源有限、数据差异大及模态数据退化缺失等挑战,实验表明该框架能准确可靠定位无人机。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03934 2026-07-07 cs.CV 新提交 57%

EgoInertia-MI: A Multimodal Egocentric Vision and IMU Benchmark for Motor Impairment Assessment

EgoInertia-MI:用于运动障碍评估的多模态自我中心视觉与惯性测量单元基准测试

Fatemah Alhamdoosh, Pietro Pala, Abduallah Mohamed, DK Arvind

机构 * University of Florence(佛罗伦萨大学) Meta Reality Labs(元现实实验室) University of Edinburgh(爱丁堡大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 研究针对运动障碍评估,结合自我中心视觉与惯性测量单元信号创建EgoInertia-MI数据集,设动作识别等任务并评估基线,发现多模态融合性能最佳,凸显两者结合用于运动评估的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11730 2026-07-07 cs.CV cs.HC cs.LG 版本更新 57%

Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

视频中矛盾/犹豫识别用于个性化数字健康干预

Manuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi, Muhammad Haseeb Aslam, Lorenzo Sia, Nicolas Richet, Marco Pedersoli, Alessandro Lameiras Koerich, Simon L Bacon, Eric Granger

机构 * LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(ETS蒙特利尔大学系统工程系LIVIA实验室) LIVIA, Dept. of Software and IT Engineering, ETS Montreal, Canada(ETS蒙特利尔大学软件与信息工程系LIVIA实验室) Dept. of Health, Kinesiology, & Applied Physiology, Concordia University, Montreal, Canada(康科迪亚大学健康、运动科学与应用生理学系) Montreal Behavioural Medicine Centre, CIUSSS Nord-de-l’Ile-de-Montréal, Canada(蒙特利尔行为医学中心,蒙特利尔北岛卫生与社会服务局)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本文研究了通过深度学习模型在视频中进行矛盾/犹豫识别,以提升数字健康干预的个性化和成本效益,实验基于新的BAH视频数据集,发现需改进多模态模型以准确识别矛盾/犹豫。

Comments 11 pages, 4 figures, ACII 2026. arXiv admin note: substantial text overlap with arXiv:2505.19328

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02734 2026-07-07 cs.CL cs.AI 新提交 50%

Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity

动荡的回声:用于假新闻和暴力驱动的暴民活动早期预警的多模态NLP框架

Md. Maruf Bangabashi, Tahmid Hasan, Golam Mahmud, Md. Mostafijur Rahman, Md. Toufiqur Rahman, Jahanur Biswas

机构 * Department of Computer Science and Engineering, Dhaka International University(达卡国际大学计算机科学与工程系) Department of Computer Science and Engineering, Bangladesh University of Engineering and Technology(孟加拉国工程技术大学计算机科学与工程系)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 针对社交媒体上错误信息传播引发危害的问题,提出多语言、多模态NLP框架。通过融合多基准数据集创建数据集,集成多种技术进行多模态融合,实验显示该框架在早期错误信息检测中有效。

Comments Accepted for publication as a book chapter (Taylor & Francis, 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.15255 2026-07-07 eess.SY cs.AI cs.SY econ.GN math.OC q-fin.EC 版本更新 50%

Double Fuzzy Probabilistic Interval Linguistic Term Set and a Dynamic Fuzzy Decision Making Model based on Markov Process with tts Application in Multiple Criteria Group Decision Making

双模糊概率区间语言术语集及基于马尔可夫过程的动态模糊决策模型在多准则群体决策中的应用

Zongmin Liu

机构 * Stanford University(斯坦福大学)

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 针对概率语言术语缺陷及动态属性权重确定难问题,提出双模糊概率区间语言术语集概念,定义相关内容,开发模糊语言马尔可夫矩阵等,给出权重确定方法,并用金融风险投资案例说明其在多准则决策中的应用。

Comments submitted to IEEE Transactions on Fuzzy Systems

详情

展开后加载摘要…

URL PDF HTML 收藏