arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-01 至 2026-05-01 共收录 68 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 5 篇

2601.19423 2026-05-01 cs.IR 78%

UniRec: Unified Multimodal Encoding for LLM-Based Recommendations

UniRec:基于LLM的推荐系统的统一多模态编码

Zijie Lei, Tao Feng, Zhigang Hua, Yan Xie, Guanyu Lin, Shuang Yang, Ge Liu, Jiaxuan You

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 UniRec通过统一多模态编码解决推荐系统中多模态信息的理解挑战,提出三元组表示和分层Q-Former结构,实现在多个基准测试中提升15%的性能。

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07451 2026-05-01 cs.CV 74%

A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion

动态神经网络综述:从计算机视觉到多模态传感器融合

Fabio Montello, Ronja Güldenring, Simone Scardapane, Lazaros Nalpantidis

机构 * DTU Electrical and Photonics Engineering, Technical University of Denmark(丹麦技术大学电气与光子工程系) DIET Department, Sapienza University of Rome(罗马萨皮恩扎大学DIET系)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 本文综述了动态神经网络在计算机视觉中的应用,探讨了其在多模态传感器融合中的优势,提出了一种基于网络组件适应性的分类方法,并提供了相关研究的资源库。

Comments Under review at Image and Vision Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27559 2026-05-01 cs.CV cs.AI 73%

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation

RIHA:报告-图像层次对齐用于放射学报告生成

Yucheng Chen, Yang Yu, Yufei Shi, Conghao Xiong, Xulei Yang, Si Yong Yeo

机构 * Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学李科钦医学院) MedVisAI Lab(医学影像人工智能实验室) Centre of AI in Medicine, Singapore(新加坡人工智能医学中心) Chinese University of Hong Kong, Hong Kong SAR, China(香港中文大学(深圳)) Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出RIHA框架,通过多级对齐提升放射学报告生成的准确性,引入视觉和文本特征金字塔及跨模态对齐模块,实验表明其在自然语言生成和临床效果上优于现有方法。

Comments Accepted by Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22500 2026-05-01 cs.CV cs.AI 62%

OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment

OR-VSKC:通过合成数据引导对齐解决手术室中的视觉-语义知识冲突

Weiyi Zhao, Xiaoyu Tan, Liang Liu, Sijia Li, Youwei Song, Xihe Qiu

机构 * Shanghai University of Engineering Science(上海工程技术大学) Tencent YouTu Lab(腾讯优图实验室) Clinical Research Unit, Zhongshan Hospital of Fudan University(复旦大学中山医院临床研究部)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出OR-VSKC基准,通过合成数据引导对齐解决手术室中的视觉-语义知识冲突,评估多模态大语言模型的可靠性,并展示微调后的模型在未见视角上的鲁棒性。

Comments 13 pages, 5 figures. The dataset and appendix are available at https://github.com/zgg2577/VS-KC

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27105 2026-05-01 cs.CV 57%

Automated Detection of Mutual Gaze and Joint Attention in Dual-Camera Settings via Dual-Stream Transformers

通过双流Transformer实现双摄像头设置中互视与联合注意的自动检测

Jakub Kosmydel, Paweł Gajewski, Arkadiusz Białek

机构 * AGH University(AGH大学) Jagiellonian University(雅盖隆大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出双流Transformer架构,用于从同步双摄像头记录中自动检测互视与联合注意,通过冻结的注视感知骨干网络和自定义令牌融合机制,有效提升了多摄像头实验室环境下的检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 3 篇

2604.27899 2026-05-01 cs.AI 74%

Simulating clinical interventions with a generative multimodal model of human physiology

用生成式多模态模型模拟临床干预

Guy Lutsker, Gal Sapir, Jordi Merino, Smadar Shilo, Anastasia Godneva, Eli Meirom, Shie Mannor, Hagai Rossman, Gal Chechik, Eran Segal

机构 * Department of Computer Science and Applied Mathematics, Weizmann Institute of Science(魏茨曼科学研究所计算机科学与应用数学系) Department of Molecular Cell Biology, Weizmann Institute of Science(魏茨曼科学研究所分子细胞生物学系) NVIDIA Novo Nordisk Foundation Center for Basic Metabolic Research, University of Copenhagen(诺沃维克基金会基础代谢研究中心,哥本哈根大学) Faculty of Medical and Health Sciences, Tel Aviv University(特拉维夫大学医学与健康科学学院) The Jesse Z and Sara Lea Shafer Institute for Endocrinology and Diabetes, National Center for Childhood Diabetes, Schneider Children’s Medical Center of Israel(杰西Z和索菲亚·李·沙弗内分泌学与糖尿病研究所,以色列儿童糖尿病国家中心,施耐德儿童医学中心) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他多模态 :multimodal(title);分类 cs.AI

AI总结 本文提出HealthFormer模型,通过训练人类表型项目数据,生成人类生理轨迹,实现对个体生理变化的预测和干预模拟,提升临床风险评分和疾病预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13486 2026-05-01 cs.CV 57%

Uncertainty Quantification Framework for Aerial and UAV Photogrammetry through Error Propagation

通过误差传播的航空和无人机摄影测量不确定性量化框架

Debao Huang, Rongjun Qin

机构 * Geospatial Data Analytics Laboratory, The Ohio State University(地理空间数据分析实验室,俄亥俄州立大学) Department of Civil, Environmental and Geodetic Engineering, The Ohio State University(土木、环境与大地测量工程系,俄亥俄州立大学) Department of Electrical and Computer Engineering, The Ohio State University(电气与计算机工程系,俄亥俄州立大学) Translational Data Analytics Institute, The Ohio State University(转化数据分析研究院,俄亥俄州立大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出通过误差传播的不确定性量化框架,解决多视立体阶段的不确定性估计问题,利用自校准方法提升摄影测量点云的鲁棒性和可验证性。

Comments 27 pages, 12 figures, this manuscript has been accepted to ISPRS Journal of Photogrammetry and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20052 2026-05-01 stat.CO 50%

Annealed Langevin Monte Carlo for Flow ODE Sampling

退火 Langevin 采样用于流 ODE 采样

Hanwen Huang

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文提出 ALMC-ODE 方法,通过退火 Langevin 链生成未归一化分布样本,尤其适用于多模态密度。该方法基于概率流 ODE,利用 Jarzynski 方案降低方差,理论分析证明了估计误差界。

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏