arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

2410.23402 2025-02-11 cs.SE 78%

VisualCoder: Guiding Large Language Models in Code Execution with Fine-grained Multimodal Chain-of-Thought Reasoning

Cuong Chi Le, Hoang-Chau Truong-Vinh, Huy Nhat Phan, Dung Duy Le, Tien N. Nguyen, Nghi D. Q. Bui

专题命中 多模态生成 :multimodal(title,abstract)

Comments NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12849 2025-01-31 cs.RO 78%

Whole-Body Trajectory Optimization for Robot Multimodal Locomotion

Giuseppe L'Erario, Gabriele Nava, Giulio Romualdi, Fabio Bergonti, Valentino Razza, Stefano Dafarra, Daniele Pucci

专题命中 多模态生成 :multimodal(title,abstract)

Comments Paper accepted in Humanoids 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08825 2025-01-16 eess.SP 78%

A Multi-modal Intelligent Channel Model for 6G Multi-UAV-to-Multi-Vehicle Communications

Lu Bai, Mengyuan Lu, Ziwei Huang, Xiang Cheng

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07333 2025-01-14 eess.SP 78%

Synesthesia of Machines Based Multi-Modal Intelligent V2V Channel Model

Zengrui Han, Lu Bai, Ziwei Huang, Xiang Cheng

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12933 2025-01-14 cs.SE 78%

ACTesting: Automated Cross-modal Testing Method of Text-to-Image Software

Siqi Gu, Chunrong Fang, Quanjun Zhang, Zhenyu Chen

专题命中 多模态生成 :cross-modal(title,abstract)

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17420 2024-12-03 cs.CE eess.IV 78%

Cross-modal Medical Image Generation Based on Pyramid Convolutional Attention Network

Fuyou Mao, Lixin Lin, Ming Jiang, Dong Dai, Chao Yang, Hao Zhang, Yan Tang

专题命中 多模态生成 :cross-modal(title);multimodal(abstract)

Comments 18 pages, 6 figures, Machine Vision and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15220 2024-11-26 cs.LG cs.NA math.NA stat.CO stat.ML 78%

Sampling with Adaptive Variance for Multimodal Distributions

Björn Engquist, Kui Ren, Yunan Yang

专题命中 多模态生成 :multimodal(title,abstract)

Comments 26 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11394 2024-11-19 cs.RO 78%

InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models

Yu Yan, Rongtao Xu, Jiazhao Zhang, Peiyang Li, Xiaodan Liang, Jianqin Yin

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09117 2024-11-15 cs.LG cs.DS math.PR stat.ML 78%

Efficiently learning and sampling multimodal distributions with data-based initialization

Frederic Koehler, Holden Lee, Thuy-Duong Vuong

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03711 2024-11-07 eess.SP 78%

Multi-Modal Intelligent Channel Modeling: A New Modeling Paradigm via Synesthesia of Machines

Lu Bai, Ziwei Huang, Mingran Sun, Xiang Cheng, Lizhen Cui

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.04003 2024-10-30 eess.SP 78%

Modeling Time-dependent CO$_2$ Intensities in Multi-modal Energy Systems with Storage

Christopher Ripp, Florian Steinke

专题命中 多模态生成 :multi-modal(title,abstract)

Comments This work has been submitted to the Elsevier Applied Energy for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13782 2024-10-18 cs.LG q-bio.QM 78%

DPLM-2: A Multimodal Diffusion Protein Language Model

Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, Quanquan Gu

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10347 2024-10-10 cs.IT math.IT 78%

Editable-DeepSC: Cross-Modal Editable Semantic Communication Systems

Wenbo Yu, Bin Chen, Qinshan Zhang, Shu-Tao Xia

专题命中 多模态生成 :cross-modal(title,abstract)

Comments published at VTC2024-Spring

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00712 2024-10-04 q-bio.NC cs.LG 78%

NECOMIMI: Neural-Cognitive Multimodal EEG-informed Image Generation with Diffusion Models

Chi-Sheng Chen

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08790 2024-10-04 eess.IV 78%

A Multimodal Approach for Fluid Overload Prediction: Integrating Lung Ultrasound and Clinical Data

Tianqi Yang, Nantheera Anantrasirichai, Oktay Karakuş, Marco Allinovi, Alin Achim

专题命中 多模态生成 :multimodal(title,abstract)

Comments In the experiment, for the classification tasks, the network was informed with ground truth during training, significantly improving the performance. This makes the results invalid. Therefore, corrections and more validations are needed to evaluate the performance of the method

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17506 2024-09-27 cs.NI 78%

Optimizing Resource Allocation for Multi-modal Semantic Communication in Mobile AIGC Networks: A Diffusion-based Game Approach

Jian Liu, Ming Xiao, Jinbo Wen, Jiawen Kang, Ruichen Zhang, Tao Zhang, Dusit Niyato, Weiting Zhang, Ying Liu

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08269 2024-09-13 cs.RO 78%

Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation

Samanta Rodriguez, Yiming Dou, Miquel Oller, Andrew Owens, Nima Fazeli

专题命中 多模态生成 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08078 2024-08-20 cs.CV cs.AI cs.CL 78%

Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation

Wenting Chen, Linlin Shen, Jingyang Lin, Jiebo Luo, Xiang Li, Yixuan Yuan

专题命中 多模态生成 :image-text(title);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2024

Journal ref https://aclanthology.org/2024.acl-long.514/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04423 2024-08-09 cs.RO 78%

UNMuTe: Unifying Navigation and Multimodal Dialogue-like Text Generation

Niyati Rawal, Roberto Bigazzi, Lorenzo Baraldi, Rita Cucchiara

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15132 2024-07-23 q-bio.NC cs.LG 78%

Deep multimodal saliency parcellation of cerebellar pathways: linking microstructure and individual function through explainable multitask learning

Ari Tchetchenian, Leo Zekelman, Yuqian Chen, Jarrett Rushmore, Fan Zhang, Edward H. Yeterian, Nikos Makris, Yogesh Rathi, Erik Meijering, Yang Song, Lauren J. O'Donnell

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06947 2024-07-18 eess.SP cs.LG cs.NI 78%

A Multi-Modal Simulation Framework to Enable Digital Twin-based V2X Communications in Dynamic Environments

Lorenzo Cazzella, Francesco Linsalata, Maurizio Magarini, Matteo Matteucci, Umberto Spagnolini

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05996 2024-07-09 cs.RO 78%

Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Moritz Reuss, Ömer Erdinç Yağmurlu, Fabian Wenzel, Rudolf Lioutikov

专题命中 多模态生成 :multimodal(title,abstract)

Comments RSS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04997 2024-06-07 stat.ML cs.LG q-bio.QM 78%

Generative Flows on Discrete State-Spaces: Enabling Multimodal Flows with Applications to Protein Co-Design

Andrew Campbell, Jason Yim, Regina Barzilay, Tom Rainforth, Tommi Jaakkola

专题命中 多模态生成 :multimodal(title,abstract)

Comments 60 pages, 11 figures, 6 tables; ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00681 2024-06-04 cs.LG 78%

Learning Multimodal Behaviors from Scratch with Diffusion Policy Gradient

Zechu Li, Rickmer Krohn, Tao Chen, Anurag Ajay, Pulkit Agrawal, Georgia Chalvatzaki

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06458 2024-06-04 cs.CL cs.AI cs.CV 78%

ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation

Bang Yang, Fenglin Liu, Yuexian Zou, Xian Wu, Yaowei Wang, David A. Clifton

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by TPAMI (Our code and data are available at https://github.com/yangbang18/ZeroNLG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06927 2024-05-14 cs.IR 78%

Multimodal Pretraining and Generation for Recommendation: A Tutorial

Jieming Zhu, Chuhan Wu, Rui Zhang, Zhenhua Dong

专题命中 多模态生成 :multimodal(title,abstract)

Comments Published in WWW 2024 Tutorial. Find the tutorial materials at https://mmrec.github.io/tutorial/www2024/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05273 2024-04-30 cs.IR 78%

Formalizing Multimedia Recommendation through Multimodal Deep Learning

Daniele Malitesta, Giandomenico Cornacchia, Claudio Pomo, Felice Antonio Merra, Tommaso Di Noia, Eugenio Di Sciascio

专题命中 多模态生成 :multimodal(title,abstract)

Comments Accepted in the Special Issue on Knowledge Transferring for Recommender Systems (KT4Rec) in ACM Transactions on Recommender Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13898 2024-04-23 cs.NI 78%

Cross-Modal Generative Semantic Communications for Mobile AIGC: Joint Semantic Encoding and Prompt Engineering

Yinqiu Liu, Hongyang Du, Dusit Niyato, Jiawen Kang, Zehui Xiong, Shiwen Mao, Ping Zhang, Xuemin Shen

专题命中 多模态生成 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16075 2024-03-28 eess.SY cs.SY math.OC 78%

Redesigning Large-Scale Multimodal Transit Networks with Shared Autonomous Mobility Services

Max T. M. Ng, Hani S. Mahmassani, Ömer Verbas, Taner Cokyasar, Roman Engelhardt

专题命中 多模态生成 :multimodal(title,abstract)

Comments 48 pages, 18 figures, accepted for publication in Transportation Research Part C: Emerging Technologies, and presentation in the 25th International Symposium on Transportation and Traffic Theory (ISTTT25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16988 2024-03-26 physics.chem-ph physics.optics 78%

Multimodal operando microscopy reveals that interfacial chemistry and nanoscale performance disorder dictate perovskite solar cell stability

Kyle Frohna, Cullen Chosy, Amran Al-Ashouri, Florian Scheler, Yu-Hsien Chiang, Milos Dubajic, Julia E. Parker, Jessica M. Walker, Lea Zimmermann, Thomas A. Selby, Yang Lu, Bart Roose, Steve Albrecht, Miguel Anaya, Samuel D. Stranks

专题命中 多模态生成 :multimodal(title,abstract)

Comments Main text and supplementary information. Main text 26 pages, 4 figures. Supplementary information 79 pages, 76 figures. Kyle Frohna and Cullen Chosy contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏