arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

2604.07422 2026-04-10 cs.LG 78%

Multimodal Large Language Models for Multi-Subject In-Context Image Generation

多模态大语言模型用于多主体上下文图像生成

Yucheng Zhou, Dubing Chen, Huan Zheng, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(澳门大学,科技学院,计算机与信息科学系,智慧城市物联网国家重点实验室)

专题命中 多模态生成 :multimodal(title);MLLM(abstract)

AI总结 本文提出MUSIC模型,通过视觉链式思维机制和空间布局规划方法,解决多主体上下文图像生成中的主体缺失和语义漂移问题,并在多主体场景中取得显著优势。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05181 2026-04-08 cs.LG 78%

General Multimodal Protein Design Enables DNA-Encoding of Chemistry

通用多模态蛋白质设计使化学编码成为可能

Jarrid Rector-Brooks, Théophile Lambert, Marta Skreta, Daniel Roth, Yueming Long, Zi-Qi Li, Xi Zhang, Miruna Cretu, Francesca-Zhoufan Li, Tanvi Ganapathy, Emily Jin, Avishek Joey Bose, Jason Yang, Kirill Neklyudov, Yoshua Bengio, Alexander Tong, Frances H. Arnold, Cheng-Hao Liu

机构 * California Institute of Technology(加州理工学院) Mila – Québec AI Institute(Mila – 魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Université Paris-Saclay(巴黎-萨克雷大学) McGill University(麦吉尔大学) University of Cambridge(剑桥大学) University of Oxford(牛津大学) Imperial College London(伦敦帝国理工学院) Institut Courtois(库尔图瓦研究所) LawZero AITHYRA FutureHouse

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 DISCO模型通过多模态设计实现蛋白质序列和三维结构的协同优化,能够设计出新型血红素酶,催化新的化学反应,拓展了遗传编码转化的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07297 2026-03-31 cs.LG q-bio.QM 78%

MM-DADM: Multimodal Drug-Aware Diffusion Model for Virtual Clinical Trials

MM-DADM:多模态药物感知扩散模型用于虚拟临床试验

Qian Shao, Bang Du, Zepeng Li, Qiyuan Chen, Jiahe Chen, Hongxia Xu, Jimeng Sun, Jian Wu, Jintai Chen

机构 * Zhejiang University(浙江大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出MM-DADM,一种多模态药物感知扩散模型,用于生成个性化药物诱导ECG。通过动态交叉注意模块融合外部物理知识,解决形态真实性和病理灵活性的平衡问题,并利用因果特征编码器提取纯药理表示,提升模拟准确性和召回率。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14139 2026-03-31 cs.RO 78%

FlexiCup: Wireless Multimodal Suction Cup with Dual-Zone Vision-Tactile Sensing

FlexiCup: 无线多模态吸盘 with 双区视觉-触觉感知

Junhao Gong, Shoujie Li, Kit-Wa Sou, Changqing Guo, Hourong Huang, Tong Wu, Yifan Xie, Chenxin Liang, Chuqiao Lyu, Xiaojun Liang, Wenbo Ding

机构 * Shenzhen Ubiquitous Data Enabling Key Lab, Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,深圳泛在数据赋能重点实验室) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 FlexiCup通过无线电子集成双区视觉-触觉传感,实现接触感知操控,验证了传感与驱动解耦的有效性,同时在不同气动原理下均表现出色。

Comments Accepted by IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12015 2026-03-24 cs.GR 78%

M^3ashy: Multi-Modal Material Synthesis via Hyperdiffusion

M^3ashy:基于超扩散的多模态材料合成

Chenliang Zhou, Zheyuan Hu, Alejandro Sztrajman, Yancheng Cai, Yaru Liu, Cengiz Oztireli

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出M^3ashy框架,通过超扩散和神经场实现复杂真实材料的高质量重建,支持多种条件下的材料合成,并贡献了新的材料数据集和BRDF评估指标。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, No. 16, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14615 2026-03-20 cs.RO 78%

Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models

利用上下文条件扩散模型加速多模态运动规划

Edward Sandra, Lander Vanroye, Dries Dirckx, Ruben Cartuyvels, Jan Swevers, Wilm Decré

机构 * Department of Mechanical Engineering, KU Leuven(卢森堡鲁文大学机械工程系) Department of Computer Science, KU Leuven(卢森堡鲁文大学计算机科学系)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出CAMPD模型,通过上下文感知的扩散模型实现高效的多模态运动规划,适用于多样场景且无需重新训练,展示出在未见环境中生成高质量轨迹的能力。

Comments Accepted for publication at the 2026 IEEE International Conference on Robotics & Automation (ICRA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14199 2026-03-17 cond-mat.mtrl-sci 78%

Millimeter-Scale, Atomically Controlled 2D Topological Insulators Revealed by Multimodal Spectroscopy

毫米级、原子级控制的二维拓扑绝缘体通过多模式光谱揭示

Woojoo Lee, Qiang Gao, Yufei Zhao, Hui Li, Albert Tsui, Yichao Zhang, Yunhe Bai, Haoran Lin, Khanh Duy Nguyen, Gabriele Berruto, Gangbin Yan, Jianchen Dang, Tongyao Wu, Hossein Rokni, Thomas S. Marchese, Ying Shirley Meng, Chao-Xing Liu, Xiao-Xiao Zhang, Chong Liu, Pinshane Y. Huang, Mark C. Hersam, Binghai Yan, Shuolong Yang

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 研究通过多模式光谱揭示了毫米级原子级控制的二维拓扑绝缘体,展示了Bi2Te3和MnBi2Te4/Bi2Te3异质结构的电子结构和拓扑边缘态,验证了其作为稳健二维拓扑绝缘体的特性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13440 2026-03-17 cs.LG cs.IT math.IT 78%

Improving Channel Estimation via Multimodal Diffusion Models with Flow Matching

通过多模态扩散模型与流匹配改进信道估计

Xiaotian Fan, Xingyu Zhou, Le Liang, Xiao Li, Shi Jin

机构 * School of Information Science and Engineering, Southeast University(信息科学与工程学院,东南大学) Purple Mountain Laboratories(紫金山实验室)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出基于流匹配和扩散变压器的多模态信道估计框架MultiCE-Flow,通过融合LiDAR、相机和位置数据生成语义条件,利用稀疏试点作为结构条件,实现高保真信道重建,优于传统方法和生成模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09394 2026-03-12 cs.HC 78%

EyeAgent: An Agentic AI System for Multimodal Clinical Decision Support in Ophthalmology

EyeAgent:一种用于眼科多模态临床决策支持的智能AI系统

Danli Shi, Xiaolan Chen, Bingjie Yan, Weiyi Zhang, Pusheng Xu, Jiancheng Yang, Ruoyu Chen, Siyu Huang, Bowen Liu, Xinyuan Wu, Meng Xie, Ziyu Gao, Yue Wu, Senlin Lin, Kai Jin, Xia Gong, Yih Chung Tham, Xiujuan Zhang, Li Dong, Yuzhou Zhang, Jason Yam, Guangming Jin, Xiaohu Ding, Haidong Zou, Yalin Zheng, Zongyuan Ge, Mingguang He

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 EyeAgent通过整合53种眼科工具提升诊断准确率,实现多模态临床决策支持,为眼科AI系统提供新范式。

Comments 28 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05433 2026-03-11 cs.LG cs.NE 78%

Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies

多模态大语言模型辅助进化搜索用于程序化控制策略

Qinglong Hu, Xialiang Tong, Mingxuan Yuan, Fei Liu, Zhichao Lu, Qingfu Zhang

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本研究提出多模态大语言模型辅助进化搜索方法,用于生成透明且可验证的程序化控制策略,通过结合视觉反馈与进化搜索提高策略发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01479 2026-03-11 cs.RO 78%

Multimodal Adversarial Quality Policy for Safe Grasping

多模态对抗质量策略用于安全抓取

Kunlin Xie, Chenghao Li, Haolan Zhang, Nak Young Chong

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出多模态对抗质量策略,通过双补丁优化和梯度平衡策略实现安全抓取,提升多模态环境下机器人抓取的安全性和鲁棒性。

Comments submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04466 2026-03-06 cs.RO cs.LG 78%

Act-Observe-Rewrite: Multimodal Coding Agents as In-Context Policy Learners for Robot Manipulation

动作-观察-重写:多模态编码代理作为机器人操作的上下文政策学习者

Vaishak Kumar

机构 * General Bionix, Inc.(通用生物交叉公司)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 AOR框架通过生成可执行代码使机器人操作策略学习无需演示或奖励工程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22246 2026-02-27 cs.CR cs.LG 78%

Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models

自净化缓解多模态扩散语言模型中的后门

Guangnian Wan, Qi Li, Gongfan Fang, Xinyin Ma, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 DiSP通过在推理过程中选择性屏蔽视觉标记,有效去除多模态扩散语言模型中的后门,无需额外模型或干净数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11627 2026-02-26 cond-mat.str-el cond-mat.mes-hall cond-mat.mtrl-sci 78%

Dynamic competition between phason and amplitudon observed by ultrafast multimodal scanning tunneling microscopy

相位振荡与振幅波之间的动态竞争通过超快多模扫描隧道显微镜观测

Seokjin Bae, Arjun Raghavan, Soyeun Kim, Kejian Qu, Chengxi Zhao, Daniel P. Shoemaker, Ziqiang Wang, Fahad Mahmood, Barry Bradlyn, Vidya Madhavan

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 通过超快多模扫描隧道显微镜观测到相位振荡与振幅波之间的动态竞争,揭示了量子材料中集体激发的生成与灭绝机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15857 2026-02-24 cs.IR 78%

Large-scale Benchmarks for Multimodal Recommendation with Ducho

基于Ducho的多模态推荐系统大规模基准测试

Matteo Attimonelli, Danilo Danese, Angela Di Fazio, Daniele Malitesta, Claudio Pomo, Tommaso Di Noia

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文首次提出基于Ducho的多模态推荐系统大规模基准测试,通过统一实验环境评估多模态特征提取器的性能。

Comments Accepted in Expert Systems with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16871 2026-02-12 cs.RO 78%

HOGraspFlow: Taxonomy-Aware Hand-Object Retargeting for Multi-Modal SE(3) Grasp Generation

HOGraspFlow: 一种基于分类意识的多模态SE(3)抓取生成方法

Yitian Shi, Zicheng Guo, Rosa Wolf, Edgar Welte, Rania Rayyes

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 HOGraspFlow通过多模态SE(3)抓取生成方法,在无需显式几何先验的情况下实现高保真的抓取合成。

Comments Accepted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08830 2026-02-10 cs.HC 78%

Enhancing Generative AI Image Refinement with Scribbles and Annotations: A Comparative Study of Multimodal Prompts

通过草图和注释增强生成式AI图像细化:多模态提示的比较研究

Hyerim Park, Phuong Thao Tran, Andre Luckow, Ceenu George, Michael Sedlmair, Malin Eiband

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本研究通过比较多模态提示,探讨草图和注释如何提升生成式AI图像细化,提出原型并揭示设计师在多模态策略中的偏好。

Comments 22 pages, 14 figures. Preprint of an accepted IUI '26 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04780 2026-02-10 cs.LG cond-mat.dis-nn 78%

Dynamical Regimes of Multimodal Diffusion Models

多模态扩散模型的动力学区域

Emil Albrychiewicz, Andrés Franco Valiente, Li-Ching Chen

机构 * Leinweber Institute for Theoretical Physics(莱因韦伯理论物理研究所) Department of Physics, University of California, Berkeley, CA, 94720-7300, USA(加州大学伯克利分校物理系) Theoretical Physics Group, Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室理论物理组) Department of Radiation Oncology, University of California, San Francisco(加州大学旧金山分校放射肿瘤科) Computational Precision Health, University of California San Francisco(加州大学旧金山分校计算精准健康)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 多模态扩散模型通过耦合动态机制,揭示生成过程中的时间层次结构及同步间隙现象,为生成模型的稳定性提供理论框架。

Comments 40 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06597 2026-02-09 cs.LG 78%

DiTS: Multimodal Diffusion Transformers Are Time Series Forecasters

DiTS:多模态扩散变换器是时间序列预测器

Haoran Zhang, Haixuan Liu, Yong Liu, Yunzhong Qiu, Yuxuan Wang, Jianmin Wang, Mingsheng Long

机构 * School of Software, BNRist, Tsinghua University(软件学院、BNRist、清华大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 DiTS通过双流Transformer块处理时间序列的内生和外生变量依赖,实现高效生成预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13646 2026-02-06 cs.HC 78%

Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing

Vistoria:一种多模态系统,通过仪器文本-图像协同编辑支持虚构故事写作

Kexue Fu, Jingfei Huang, Long Ling, Sumin Hong, Yihang Zuo, Ray LC, Toby Jia-jun Li

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 Vistoria通过多模态协同编辑提升虚构故事创作的表达力和创造力,使作家能更自由地探索故事发展方向。

Comments Change the format of the first page

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21006 2026-02-05 physics.plasm-ph 78%

A joint diffusion approach to multi-modal inference in inertial confinement fusion

一种联合扩散方法用于惯性约束聚变中的多模态推断

Michael S. Jones, Justin Kunimune, Daniel Casey, Bogdan Kustowski, Eugene Kur, Kelli Humbird

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出JointDiff方法,通过联合扩散统一正向建模、逆向推断和输出填补,提升惯性约束聚变实验的多模态推断精度与可转移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00508 2026-02-05 eess.SP 78%

Quadrature Over-the-Air-Computing for Multimodal Dual-Stream Signal Processing

正交空天地计算用于多模双流信号处理

Hyeon Seok Rou, Kengo Ando, Giuseppe Thadeu Freitas de Abreu, David González G

专题命中 多模态生成 :multimodal(title);multi-modal(abstract)

AI总结 Q-OTAC通过利用复信号的同相和正交分量,实现双流同时计算,提升计算效率,适用于多模B5G应用。

Comments Accepted at the IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02927 2026-02-04 stat.ML cs.LG 78%

Training-Free Self-Correction for Multimodal Masked Diffusion Models

无需训练的多模态掩码扩散模型自校正

Yidong Ouyang, Panwen Hu, Zhengyan Wan, Zhe Wang, Liyan Xie, Dmitriy Bespalov, Ying Nian Wu, Guang Cheng, Hongyuan Zha, Qiang Sun

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Mohamed bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) East China Normal University(华东师范大学) University of Virginia(弗吉尼亚大学) University of Minnesota(明尼苏达大学) Drexel university(德雷塞尔大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Toronto(多伦多大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出无需训练的多模态掩码扩散模型自校正方法,通过减少采样步骤提升生成质量,适用于文本到图像和多模态理解任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00641 2026-02-03 stat.ML cs.LG stat.CO 78%

Sampling from multi-modal distributions on Riemannian manifolds with training-free stochastic interpolants

在黎曼流形上采样多模分布的无训练随机插值法

Alain Durmus, Maxence Noble, Thibaut Pellerin

机构 * CMAP, CNRS, Ecole polytechnique(CMAP、法国国家科学研究中心、巴黎高等师范学院)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 本文提出了一种无训练的随机插值方法,用于在黎曼流形上高效采样多模分布,通过非平衡动力学和随机插值实现无需训练的高效采样。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22407 2026-02-02 cond-mat.mtrl-sci 78%

Nanoscale mapping of phase-transformation pathways in medium-Mn TRIP steel by multimodal STEM

中等锰TRIP钢中相变路径的纳米级映射:多模态STEM

Marc Raventós-Tato, S. Leila Panahi, Núria Bagués, David Frómeta, Oleg Usoltsev, Núria Cuadrado, Joaquín Otón

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本研究通过多模态STEM技术,实现了中等锰TRIP钢中相变路径的纳米级映射,结合电子衍射与能谱分析,实现了铁、奥氏体和马氏体的相分离与晶格参数细化。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21950 2026-01-30 cs.LG 78%

Embracing Aleatoric Uncertainty in Medical Multimodal Learning with Missing Modalities

在医疗多模态学习中拥抱概率不确定性与缺失模态

Linxiao Gong, Yang Liu, Lianlong Sun, Yulai Bi, Jing Liu, Xiaoguang Zhu

机构 * HKUST (GZ)(香港科技大学) Tongji University(同济大学) University of Rochester(罗切斯特大学) Meta Fudan University(复旦大学) The University of British Columbia(不列颠哥伦比亚大学) University of California, Davis(加州大学戴维斯分校)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出AUM框架,通过建模单模态概率不确定性来应对医疗多模态学习中的缺失模态问题,在死亡预测任务中取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21129 2026-01-30 cs.RO 78%

WheelArm-Sim: A Manipulation and Navigation Combined Multimodal Synthetic Data Generation Simulator for Unified Control in Assistive Robotics

WheelArm-Sim: 一种结合 manipulation 和 navigation 的多模态合成数据生成模拟器,用于辅助机器人中的统一控制

Guangping Liu, Tipu Sultan, Vittorio Di Giorgio, Nick Hawkins, Flavio Esposito, Madi Babaiasl

机构 * Aerospace and Mechanical Engineering Department, Saint Louis University(圣路易斯大学航空航天与机械工程系) Computer Science Department, Saint Louis University(圣路易斯大学计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 WheelArm-Sim 是一种用于辅助机器人统一控制的多模态合成数据生成模拟器,通过集成轮椅和机械臂控制,为数据驱动的机器学习模型提供支持。

Comments Accepted to IEEE International Symposium on Medical Robotics (ISMR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20406 2026-01-27 cs.RO cs.LG 78%

PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

PointMapPolicy: 结构化点云处理用于多模态模仿学习

Xiaogang Jia, Qian Wang, Anrui Wang, Han A. Wang, Balázs Gyenes, Emiliyan Gospodinov, Xinkai Jiang, Ge Li, Hongyi Zhou, Weiran Liao, Xi Huang, Maximilian Beck, Moritz Reuss, Rudolf Lioutikov, Gerhard Neumann

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Reality Labs, Meta(Meta现实实验室) Johannes Kepler University Linz(林茨约翰尼斯·开普勒大学)

专题命中 多模态生成 :multi-modal(title,abstract)

AI总结 PointMapPolicy通过结构化点云处理提升多模态模仿学习的精度与泛化能力,利用xLSTM融合点云与RGB数据,在RoboCasa和CALVIN基准中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02900 2026-01-22 cs.CV cs.AI cs.GR cs.HC cs.MM 78%

Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions

推进说话头生成:多模态方法、数据集、评估指标和损失函数的全面综述

Vineet Kumar Rakesh, Soumya Mazumdar, Research Pratim Maity, Sarbajit Pal, Amitabha Das, Tapas Samanta

机构 * Engineering Sciences, Homi Bhabha National Institute Training School Complex(工程科学,霍米·巴赫瓦全国研究所培训学校) Computer and Informatics Group, VECC 1/AF(计算机与信息组,VECC 1/AF)

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.AI、cs.MM

AI总结 本文全面综述了说话头生成的多模态方法、数据集、评估指标和损失函数,探讨了技术挑战与未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12381 2026-01-21 q-bio.QM q-bio.BM q-bio.GN 78%

Multimodal Spatial Omics: From Data Acquisition to Computational Integration

多模态空间组学:从数据获取到计算整合

Esra Busra Isik, Yusuf Hakan Usta, Haozhe Liu, Maryam Riazi, William Roach, Hongpeng Zhou, Magnus Rattray, Sokratia Georgaka

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文综述了多模态空间组学数据整合的计算方法,涵盖从概率模型到深度学习的多种算法原理,以整合不同分子层和图像数据。

详情

展开后加载摘要…

URL PDF HTML 收藏