arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-08 至 2026-04-08 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 8 篇

2511.17652 2026-04-08 q-bio.QM cs.CV 83%

TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots

TeamPath: 构建多模态病理专家的推理AI助手

Tianyu Liu, Weihao Xuan, Hao Wu, Peter Humphrey, Marcello DiStasio, Mohamed Kahila, Alfonso Garcia Tan, Heli Qi, Rui Yang, Simeng Han, Tinglin Huang, Fang Wu, Chen Liu, Qingyu Chen, Nan Liu, Irene Li, Hua Xu, Hongyu Zhao

机构 * Interdepartmental Program in Computational Biology and Biomedical Informatics, Yale University(耶鲁大学计算生物学与生物医学信息学跨学科项目) Department of Biostatistics, Yale University(耶鲁大学生物统计学系) Broad Institute of MIT and Harvard(博德研究所) Department of Complexity Science and Engineering, The University of Tokyo(东京大学复杂科学与工程系) Center for Advanced Intelligence Project, RIKEN(理化学研究所先进智能项目中心) Department of Pathology, Yale University(耶鲁大学病理学系) Department of Anatomical Pathology, Singapore General Hospital(新加坡中央医院解剖病理学系) Center for Biomedical Data Science, Duke–NUS Medical School, Singapore, Singapore(杜克-新加坡国立大学医学院生物医学数据科学中心) Department of Computer Science, Yale University(耶鲁大学计算机科学系) Department of Computer Science, Stanford University(斯坦福大学计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 TeamPath通过强化学习和路由增强方案,整合多模态数据提升病理诊断与跨模态生成能力,为临床提供可靠的信息交流系统。

Comments 45 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05497 2026-04-08 cs.AI cs.CV 81%

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models

Thinking Diffusion: 通过惩罚和引导在扩散多模态语言模型中抑制和引导视觉推理

Keuntae Kim, Mingyu Kang, Yong Suk Choi

机构 * Department of Computer Science, Hanyang University(汉阳大学计算机科学系) Department of Artificial Intelligence, Hanyang University(汉阳大学人工智能系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文针对扩散多模态语言模型在视觉推理中生成过早答案的问题,提出位置与步数惩罚和视觉推理引导方法,提升推理准确率并加快推理速度。

Comments CVPR 2026 - main

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04570 2026-04-08 cs.CV cs.CL 81%

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

通过视频思考:视频生成作为一种有前途的多模态推理范式

Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, Yining Zheng, Xinchi Chen, Jun Zhao, Xuanjing Huang, Xipeng Qiu

机构 * Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身智能研究所) Shanghai Innovation Institute(上海创新研究院) Shanghai Key Laboratory of Multimodal Embodied AI(上海市多模态具身人工智能重点实验室) College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) The Chinese University of Hong Kong(香港中文大学) Central South University(中南大学) Fudan University(复旦大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出'通过视频思考'范式,利用视频生成模型实现多模态推理,通过VideoThinkBench评估显示Sora-2在视觉和文本任务中均表现出色,证明视频生成模型在统一多模态理解与生成中的潜力。

Comments 34 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14949 2026-04-08 cs.CL cs.CV cs.LG 81%

DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation

DialectGen:多模态生成中方言鲁棒性评估与改进

Yu Zhou, Sohyun An, Haikang Deng, Da Yin, Clark Peng, Cho-Jui Hsieh, Kai-Wei Chang, Nanyun Peng

机构 * University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文通过构建涵盖六种英语方言的大型基准,评估多模态生成模型在方言输入下的表现,提出一种通用编码器策略提升方言鲁棒性,同时保持标准美式英语性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03323 2026-04-08 cs.GR cs.CV cs.HC cs.LG cs.SD 79%

Listen to Rhythm, Choose Movements: Autoregressive Multimodal Dance Generation via Diffusion and Mamba with Decoupled Dance Dataset

聆听节奏,选择动作:通过扩散与Mamba实现多模态舞蹈生成的自回归方法

Oran Duan, Yinghua Shen, Yingzhu Lv, Luyang Jie, Yaxin Liu, Qiong Wu

机构 * Communication University of China(中国传媒大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出LRCM框架,通过解耦舞蹈数据集实现多模态指导的扩散模型,结合Mamba模块实现流畅的长序列自回归生成,展示了在多模态输入和长序列生成中的潜力。

Comments 12 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05181 2026-04-08 cs.LG 78%

General Multimodal Protein Design Enables DNA-Encoding of Chemistry

通用多模态蛋白质设计使化学编码成为可能

Jarrid Rector-Brooks, Théophile Lambert, Marta Skreta, Daniel Roth, Yueming Long, Zi-Qi Li, Xi Zhang, Miruna Cretu, Francesca-Zhoufan Li, Tanvi Ganapathy, Emily Jin, Avishek Joey Bose, Jason Yang, Kirill Neklyudov, Yoshua Bengio, Alexander Tong, Frances H. Arnold, Cheng-Hao Liu

机构 * California Institute of Technology(加州理工学院) Mila – Québec AI Institute(Mila – 魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Université Paris-Saclay(巴黎-萨克雷大学) McGill University(麦吉尔大学) University of Cambridge(剑桥大学) University of Oxford(牛津大学) Imperial College London(伦敦帝国理工学院) Institut Courtois(库尔图瓦研究所) LawZero AITHYRA FutureHouse

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 DISCO模型通过多模态设计实现蛋白质序列和三维结构的协同优化,能够设计出新型血红素酶,催化新的化学反应,拓展了遗传编码转化的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04403 2026-04-08 cs.AI 57%

MolDA: Molecular Understanding and Generation via Large Language Diffusion Model

MolDA:通过大规模语言扩散模型实现分子理解与生成

Seohyeon Shin, HanJun Choi, Jun-Hyung Park, Hong Kook Kim, Mansu Kim

机构 * Gwangju Institute of Science and Technology(光州科学技术院) Hankuk University of Foreign Studies(韩国外国语大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 MolDA通过大规模语言扩散模型替代传统自回归架构,解决分子生成中非局部约束和结构误差问题,实现分子生成、描述和性质预测的全局一致性与化学有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18297 2026-04-08 cs.CV 57%

Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module

利用自适应共注意和三LSTM模块的图像到文本生成医学报告

Yishen Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出CA-TriNet模型,结合Transformer和多LSTM网络,通过自适应权重运算和三LSTM模块提升医学图像与文本生成的准确性与多样性。

Comments arXiv admin note: This submission has been withdrawn by arXiv administrators due to incorrect authorship. Author list truncated

详情

展开后加载摘要…

URL PDF HTML 收藏