arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

2004.03212 2021-03-23 cs.CV cs.CL 73%

Text-Guided Neural Image Inpainting

Lisai Zhang, Qingcai Chen, Baotian Hu, Shuoran Jiang

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments ACM MM'2020 (Oral). 9 pages, 4 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10767 2026-08-04 cond-mat.mes-hall cond-mat.mtrl-sci quant-ph 版本更新 71%

Lambert W Function Framework for Graphene Nanoribbon Quantum Sensing: Theory, Verification, and Multi-Modal Applications

基于拉姆伯特W函数框架的石墨烯纳米带量子传感:理论、验证与多模应用

F. A. Chishtie, K. Roberts, N. Jisrawi, S. R. Valluri, A. Soni, P. C. Deshmukh

专题命中 多模态生成 :multi-modal(title)

AI总结 基于拉姆伯特W函数的石墨烯纳米带量子传感框架,通过理论验证和多模应用,实现了灵敏度增强和性能预测。

Comments 21 pages, 10 figures, published version at Results in Engineering journal

Journal ref Results in Engineering, Volume 32, 2026, 112238

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21450 2026-06-24 cs.CE q-bio.BM 版本更新 71%

CMADiff: Cross-Modal Aligned Diffusion for Controllable Protein Generation

CMADiff: 用于可控蛋白质生成的跨模态对齐扩散

Changjian Zhou, Yuexi Qiu, Jiafeng Li, Jia Song, Wensheng Xiang

专题命中 多模态生成 :cross-modal(title)

AI总结 提出CMADiff框架,通过条件变分自编码器整合理化特征,并利用对比学习模块BioAligner对齐文本描述与蛋白质特征,实现基于文本驱动的可控蛋白质序列生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16618 2026-06-16 eess.SP 新提交 71%

Acoustic, VOC, and Multimodal Stress Source Localization in the Internet of Plants

植物物联网中的声学、VOC和多模态胁迫源定位

Ahmet B. Kilic, Ozgur B. Akan

专题命中 多模态生成 :multimodal(title)

AI总结 提出一种两阶段粗到细定位流程,结合声学到达时间差多边定位和VOC弥散格林函数模型,实现植物网络中胁迫源的空间定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00535 2026-06-02 cs.LG 71%

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation

DREAM-S: 基于可搜索草稿与目标感知精炼的推测解码用于多模态生成

Zining Liu, Yunhai Hu, Tianhua Xia, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

机构 * New York University(纽约大学) Cerebras Systems Inc.(Cerebras Systems公司) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态生成 :multimodal(title)

AI总结 提出DREAM-S框架,通过神经架构搜索和目标感知超网训练自动优化草稿模型架构与交互策略,结合注意力熵引导的自适应中间特征蒸馏,实现视觉语言模型的高效推测解码,加速比达3.85倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16473 2026-05-19 stat.ML cs.LG cs.NA math.NA math.PR 71%

Dimension-Uniform Discretization Analysis of Preconditioned Annealed Langevin Dynamics for Multimodal Gaussian Mixtures

预处理退火 Langevin 动力学在多模高斯混合中的维度均匀离散化分析

Lorenzo Baldassari, Josselin Garnier, Knut Solna, Maarten V. de Hoop

机构 * University of Basel(巴塞尔大学) Ecole Polytechnique, IP Paris(巴黎高等理工学院) University of California Irvine(加州大学尔湾分校) Rice University(里德大学)

专题命中 多模态生成 :multimodal(title)

AI总结 本文研究了预处理退火 Langevin 动力学在高斯混合中的稳定性问题,通过 Euler-Maruyama 离散化和指数积分方案,证明了在满足特定谱条件时,KL 散度具有维度均匀的上界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13502 2026-05-14 eess.SP 71%

A Multi-Modal Intelligent U2V Channel Model for 6G Sensing-Communication Integration

一种面向6G感知通信一体化的多模智能U2V信道模型

Shuo Wang, Zengrui Han, Lu Bai, Xiang Cheng

专题命中 多模态生成 :multi-modal(title)

AI总结 本文提出一种基于三维散射体预测的新型U2V信道模型,通过构建宽车道场景下的高保真混合感知通信集成U2V仿真数据集,设计了3D-SPADE算法,利用LiDAR点云准确预测散射体分布,提升了动态U2V场景的建模精度与计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04421 2026-04-07 cond-mat.supr-con cond-mat.str-el 71%

Multimodal Terahertz Spectroscopy of the Pairing Symmetry and Normal-State Pseudogap in (La,Pr)$_3$Ni$_2$O$_7$ Films

多模态太赫兹光谱学研究(La,Pr)3Ni2O7薄膜中的配对对称性和正常态伪间隙

Shuxiang Xu, Guangdi Zhou, Hao Wang, Tianyi Wu, Wei Wang, Liyu Shi, Dong Wu, Haoliang Huang, Xinbo Wang, Jinfeng Jia, Qi-Kun Xue, Zhuoyu Chen, Tao Dong, Nanlin Wang

专题命中 多模态生成 :multimodal(title)

AI总结 通过结合体敏感太赫兹时域光谱与太赫兹三次谐波生成,研究(La,Pr)3Ni2O7薄膜的超导配对对称性和正常态伪间隙特性,发现有序态与伪间隙的共存与竞争。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02763 2026-01-13 math.ST cs.NA math.NA math.PR stat.CO stat.TH 71%

Time-complexity of sampling from a multimodal distribution using sequential Monte Carlo

使用序贯蒙特卡罗方法从多模分布中采样的时间复杂度

Ruiyu Han, Gautam Iyer, Dejan Slepčev

专题命中 多模态生成 :multimodal(title)

AI总结 该研究探讨了在低温下使用序贯蒙特卡罗方法从多模分布采样的时间复杂度,分析了几何退火调度与兰格-丹德扩散的效率,并得出了收敛性的数学结果。

Comments 65 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26670 2025-10-31 cs.RO 71%

Hybrid Consistency Policy: Decoupling Multi-Modal Diversity and Real-Time Efficiency in Robotic Manipulation

Qianyou Zhao, Yuliang Shen, Xuanran Zhai, Ce Hao, Duidi Wu, Jin Qi, Jie Hu, Qiaojun Yu

机构 * Shanghai Jiao Tong University(上海交通大学) National University of Singapore(国立新加坡大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态生成 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01622 2025-10-03 cs.RO cs.LG 71%

VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation

Xuanran Zhai, Qianyou Zhao, Qiaojun Yu, Ce Hao

机构 * National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态生成 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14837 2025-09-18 cs.RO cs.LG 71%

Learning Multimodal Attention for Manipulating Deformable Objects with Changing States

Namiko Saito, Mayu Tatsumi, Ayuna Kubo, Kanata Suzuki, Hiroshi Ito, Shigeki Sugano, Tetsuya Ogata

机构 * Future Robotics Organization, Waseda University(早稻田大学未来机器人组织) Microsoft Research Asia(微软亚洲研究院) Department of Modern Mechanical Engineering, Waseda University(早稻田大学现代机械工程系) Artificial Intelligence Laboratories, Fujitsu Limited(Fujitsu 人工智能实验室) Center for Technology Innovation - Controls and Robotics, Research & Development Group, Hitachi, Ltd.(富士通技术研发集团技术创新中心 - 控制与机器人) Faculty of Science and Engineering, Waseda University(早稻田大学工学部) National Institute of Advanced Science and Technology(国家先进科学研究院)

专题命中 多模态生成 :multimodal(title)

Comments Humanoids2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16320 2025-09-16 astro-ph.IM cs.LG 71%

Learning novel representations of variable sources from multi-modal $\textit{Gaia}$ data via autoencoders

P. Huijse, J. De Ridder, L. Eyer, L. Rimoldini, B. Holl, N. Chornay, J. Roquette, K. Nienartowicz, G. Jevardat de Fombelle, D. J. Fritzewski, A. Kemp, V. Vanlaer, M. Vanrespaille, H. Wang, M. I. Carnerero, C. M. Raiteri, G. Marton, M. Madarász, G. Clementini, P. Gavras, C. Aerts

机构 * Institute of Astronomy, KU Leuven, Celestijnenlaan 200D, B-3001 Leuven, Belgium Millennium Institute of Astrophysics, Nuncio Monse\ nor Sotero Sanz 100, Of. 104, Providencia, Santiago, Chile Department of Astronomy, University of Geneva, Chemin Pegasi 51, 1290 Versoix, Switzerland Department of Astronomy, University of Geneva, Chemin d’Ecogia 16, 1290 Versoix, Switzerland Sednai S\`arl, Geneva, Switzerland INAF - Osservatorio Astrofisico di Torino, Via Osservatorio 20, I-10025 Pino Torinese, Italy Konkoly Observatory, HUN-REN Research Centre for Astronomy Earth Sciences, Konkoly Thege 15-17, 1121 Budapest, Hungary CSFK, MTA Centre of Excellence, Konkoly Thege 15-17, 1121, Budapest, Hungary INAF - Osservatorio di Astrofisica e Scienza dello Spazio di Bologna, Via Piero Gobetti 93/3, Bologna 40129, Italy Starion for European Space Agency, Camino bajo del Castillo, s/n, Urbanizacion Villafranca del Castillo, Villanueva de la Ca \ n ada, 28692 Madrid, Spain Department of Astrophysics, IMAPP, Radboud University Nijmegen, PO Box 9010, 6500 GL Nijmegen, The Netherlands Max Planck Institute for Astronomy, Koenigstuhl 17, 69117 Heidelberg, Germany

专题命中 多模态生成 :multi-modal(title)

Comments Manuscript accepted on Astronomy & Astrophysics, 20 pages, 20 figures, 2 tables

Journal ref A&A 701, A150 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23149 2025-09-03 eess.IV 71%

Towards Interpretable Counterfactual Generation via Multimodal Autoregression

Chenglong Ma, Yuanfeng Ji, Jin Ye, Lu Zhang, Ying Chen, Tianbin Li, Mingjie Li, Junjun He, Hongming Shan

专题命中 多模态生成 :multimodal(title)

Comments MICCAI'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23180 2025-07-01 cs.HC 71%

ImprovMate: Multimodal AI Assistant for Improv Actor Training

Riccardo Drago, Yotam Sechayk, Mustafa Doga Dogan, Andrea Sanna, Takeo Igarashi

专题命中 多模态生成 :multimodal(title)

Comments ACM DIS '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16380 2025-06-16 cs.RO cs.HC cs.LG 71%

Learning Multimodal Latent Dynamics for Human-Robot Interaction

Vignesh Prasad, Lea Heitlinger, Dorothea Koert, Ruth Stock-Homburg, Jan Peters, Georgia Chalvatzaki

专题命中 多模态生成 :multimodal(title)

Comments Preprint version of paper accepted at IEEE T-RO. Project website: https://sites.google.com/view/mild-hri

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17807 2025-03-31 cs.LG cs.NA math.NA stat.ML 71%

Neural Network Approach to Stochastic Dynamics for Smooth Multimodal Density Estimation

Z. Zarezadeh, N. Zarezadeh

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00734 2025-02-18 hep-lat cond-mat.stat-mech cs.LG 71%

Flow-based sampling for multimodal and extended-mode distributions in lattice field theory

Daniel C. Hackett, Chung-Chun Hsieh, Sahil Pontula, Michael S. Albergo, Denis Boyda, Jiunn-Wei Chen, Kai-Feng Chen, Kyle Cranmer, Gurtej Kanwar, Phiala E. Shanahan

专题命中 多模态生成 :multimodal(title)

Comments 38+3 pages, 39 figures. v2: major revisions including new application to extended modes

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06897 2024-09-12 eess.IV 71%

RATNUS: Rapid, Automatic Thalamic Nuclei Segmentation using Multimodal MRI inputs

Anqi Feng, Zhangxing Bian, Blake E. Dewey, Alexa Gail Colinco, Jiachen Zhuo, Jerry L. Prince

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09461 2024-08-23 cs.LG cond-mat.mtrl-sci physics.chem-ph q-bio.BM 71%

Advancements in Molecular Property Prediction: A Survey of Single and Multimodal Approaches

Tanya Liyaqat, Tanvir Ahmad, Chandni Saxena

专题命中 多模态生成 :multimodal(title)

Comments Submitted to the journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04380 2024-08-22 cs.RO cs.LG 71%

Deep Generative Models in Robotics: A Survey on Learning from Multimodal Demonstrations

Julen Urain, Ajay Mandlekar, Yilun Du, Mahi Shafiullah, Danfei Xu, Katerina Fragkiadaki, Georgia Chalvatzaki, Jan Peters

专题命中 多模态生成 :multimodal(title)

Comments 20 pages, 11 figures, submitted to TRO

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20218 2024-08-06 physics.acc-ph math.OC 71%

cDVAE: Multimodal Generative Conditional Diffusion Guided by Variational Autoencoder Latent Embedding for Virtual 6D Phase Space Diagnostics

Alexander Scheinker

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13942 2024-06-21 cs.LG 71%

Synthesizing Multimodal Electronic Health Records via Predictive Diffusion Models

Yuan Zhong, Xiaochen Wang, Jiaqi Wang, Xiaokun Zhang, Yaqing Wang, Mengdi Huai, Cao Xiao, Fenglong Ma

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06227 2024-04-10 cs.HC 71%

Multimodal Road Network Generation Based on Large Language Model

Jiajing Chen, Weihang Xu, Haiming Cao, Zihuan Xu, Yu Zhang, Zhao Zhang, Siyao Zhang

专题命中 多模态生成 :multimodal(title)

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06164 2024-01-15 q-fin.GN cs.LG 71%

Multimodal Gen-AI for Fundamental Investment Research

Lezhi Li, Ting-Yu Chang, Hai Wang

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14146 2023-10-24 stat.AP 71%

Cocaine Use Prediction with Tensor-based Machine Learning on Multimodal MRI Connectome Data

Anru R. Zhang, Ryan P. Bell, Chen An, Runshi Tang, Shana A. Hall, Cliburn Chan, Kareem Al-Khalil, Christina S. Meade

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02616 2023-09-07 eess.IV cs.LG cs.NI 71%

Generative AI-aided Joint Training-free Secure Semantic Communications via Multi-modal Prompts

Hongyang Du, Guangyuan Liu, Dusit Niyato, Jiayi Zhang, Jiawen Kang, Zehui Xiong, Bo Ai, Dong In Kim

专题命中 多模态生成 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03933 2023-04-11 stat.ME cs.LG stat.CO stat.ML 71%

Efficient Multimodal Sampling via Tempered Distribution Flow

Yixuan Qiu, Xiao Wang

专题命中 多模态生成 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06826 2023-01-18 cs.RO 71%

Vis2Hap: Vision-based Haptic Rendering by Cross-modal Generation

Guanqun Cao, Jiaqi Jiang, Ningtao Mao, Danushka Bollegala, Min Li, Shan Luo

专题命中 多模态生成 :cross-modal(title)

Comments This paper is accepted at ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14066 2022-11-30 gr-qc astro-ph.HE 71%

Accelerating multimodal gravitational waveforms from precessing compact binaries with artificial neural networks

Lucy M. Thomas, Geraint Pratten, Patricia Schmidt

专题命中 多模态生成 :multimodal(title)

Comments Matches published version in PRD. 19 pages (including appendix and bibliography), 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏