arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2026-08-20 至 2026-08-20 共收录 5 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 医学影像融合 1 篇

2608.19063 2026-08-20 cs.CV 新提交 79%

When Two Tracers Disagree: An Investigation of Multimodal Fusion for Clinical PET/CT Segmentation

当两种示踪剂意见不合时:临床PET/CT分割的多模态融合研究

Jack A. Johnson, Bartłomiej W. Papież

机构 * University of Oxford(牛津大学) Nuffield Department of Medicine(纳菲尔德医学系) Department of Oncology(肿瘤学系) Big Data Institute(大数据研究所) Nuffield Department of Population Health(纳菲尔德人口健康系)

专题命中 医学影像融合 :multimodal fusion(title,abstract);分类 cs.CV

AI总结 该研究针对前列腺癌PSMA与FDG PET/CT的多模态融合分割,训练示踪剂特异性3D nnU-Net基线并对比多种融合策略,发现融合未始终优于单示踪剂模型,需更优架构保留示踪剂特异性表征

Comments 10 pages (8 pages main content and 2 pages of references), 2 figures, 2 tables, accepted to MICCAI 2026 Cancer Prevention, Detection, and IntervenTion (CaPTion) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 机器人多传感器融合 1 篇

2608.18701 2026-08-20 cs.RO 新提交 57%

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

SoftVTBench:面向可变形物体操作的形变感知视觉-触觉数据集与基准

Bowen Jing, Mingxin Wang, Ruiyang Hao, Chenchen Ge, Hanwen Shen, Junjie He, Yang Cui, Yiming Hou, Weitao Zhou, Jiawei Wang, Minglei Li, Dandan Zhang, Ding Zhao, Houde Liu, Xiaofan Li, Si Liu, Ping Luo, Haibao Yu

机构 * Tuojing Intelligence(拓境智能) Tsinghua University(清华大学) Southeast University(东南大学) Stevens Institute of Technology(斯蒂文斯理工学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Manchester(曼彻斯特大学) Simple AI Imperial College London(帝国理工学院) Carnegie Mellon University(卡内基梅隆大学) Zhejiang University(浙江大学) Beihang University(北京航空航天大学) The University of Hong Kong(香港大学)

专题命中 机器人多传感器融合 :multimodal fusion(abstract);分类 cs.RO

AI总结 本研究推出SoftVTBench视觉-触觉数据集与基准,定义形变感知成功率(DSR),发现触觉信息本身未必提升多模态融合,为可变形物体操作的物理交互研究提供资源。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 音视频/视觉语言融合 1 篇

2608.18080 2026-08-20 cs.AI 新提交 50%

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

心理健康领域的大语言模型:应用、创新与伦理挑战的系统综述

Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 该系统综述梳理了大语言模型在心理健康领域的多类应用、技术创新,同时探讨了其面临的伦理与监管挑战,并倡导建立保障其安全公平部署的框架。

Comments Systematic review. Published in Journal of Industrial Integration and Management (2025). Applications of large language models in mental health, including social media analysis, clinical conversational agents, therapy support tools, multimodal learning, and ethical considerations

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 融合架构与评测 2 篇

2608.19052 2026-08-20 cs.CR 新提交 50%

Malformer: A Multi-Modal Malware Detector Using Transformers

Malformer:一种基于Transformer的多模态恶意软件检测器

Samuel Howard, Kshitiz Aryal, Mahmoud Abdelsalam, Maanak Gupta, Andrew Wheeler, Pradip Kunwar

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本研究提出四模态恶意软件检测模型Malformer,融合文本、图像、图形、音频表示,采用多模态Transformer融合,在201549个样本数据集上准确率98.3%,优于单/双模态检测器,为恶意软件防御提供稳健基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18438 2026-08-20 cs.CL cs.AI cs.LG 新提交 50%

Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage

心理健康领域的教学式AI:用于自动化临床监督与风险分诊的三通路微调大语言模型框架

Shreeya Sharma, Ravish Gupta, Saket Kumar, Abhishek Aggarwal

机构 * Microsoft(微软) BigCommerce University at Buffalo, The State University of New York(纽约州立大学布法罗分校) Amazon(亚马逊)

专题命中 融合架构与评测 :multi-modal fusion(abstract)

AI总结 针对精神医疗的监督缺口,提出三通路微调Mistral-7B-instruct框架,实现自动化临床监督与风险分诊,提升效率并解决冷启动问题。

Comments 14 pages, 1 figure, 2 tables. Accepted for publication in AICTC 2026, Lecture Notes in Networks and Systems, vol. 2165, Springer

详情

展开后加载摘要…

URL PDF HTML 收藏