arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2409.19684 2024-10-01 cs.CV 79%

MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation

Lijian Xu, Hao Sun, Ziyu Ni, Hongsheng Li, Shaoting Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09877 2024-10-01 cs.CV 79%

LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer

Ning Yu, Chia-Chih Chen, Zeyuan Chen, Rui Meng, Gang Wu, Paul Josel, Juan Carlos Niebles, Caiming Xiong, Ran Xu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ECCV'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18127 2024-09-27 cs.CV 79%

EgoLM: Multi-Modal Language Model of Egocentric Motions

Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim, Yuting Ye, Richard Newcombe, Ziwei Liu, Lingni Ma

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Project Page: https://hongfz16.github.io/projects/EgoLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16818 2024-09-26 eess.IV cs.CV 79%

Towards General Text-guided Image Synthesis for Customized Multimodal Brain MRI Generation

Yulin Wang, Honglin Xiong, Kaicong Sun, Shuwei Bai, Ling Dai, Zhongxiang Ding, Jiameng Liu, Qian Wang, Qian Liu, Dinggang Shen

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16297 2024-09-26 cs.MM 79%

Analyzing Recursiveness in Multimodal Generative Artificial Intelligence: Stability or Divergence?

Javier Conde, Tobias Cheung, Gonzalo Martínez, Pedro Reviriego, Rik Sarkar

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14720 2024-09-24 cs.CV 79%

ControlEdit: A MultiModal Local Clothing Image Editing Method

Di Cheng, YingJie Shi, ShiXin Sun, JiaFu Zhang, WeiJing Wang, Yu Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12099 2024-09-19 cs.CV 79%

Brain-Streams: fMRI-to-Image Reconstruction with Multi-modal Guidance

Jaehoon Joo, Taejin Jeong, Seongjae Hwang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11010 2024-09-18 cs.CV 79%

MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance

Debin Meng, Christos Tzelepis, Ioannis Patras, Georgios Tzimiropoulos

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ECCV 2024 AIM workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09149 2024-09-17 cs.CV 79%

Adaptive Multi-Modal Control of Digital Human Hand Synthesis Using a Region-Aware Cycle Loss

Qifan Fu, Xiaohang Yang, Muhammad Asad, Changjae Oh, Shanxin Yuan, Gregory Slabaugh

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments This paper has been accepted by the ECCV 2024 HANDS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03961 2024-09-09 cs.CV 79%

Generating Faithful and Salient Text from Multimodal Data

Tahsina Hashem, Weiqing Wang, Derry Tanti Wijaya, Mohammed Eunus Ali, Yuan-Fang Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16889 2024-09-02 cs.CL cs.LG 79%

LLaVA-Chef: A Multi-modal Generative Model for Food Recipes

Fnu Mohbat, Mohammed J. Zaki

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06702 2024-08-30 cs.CV 79%

Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization

Jinlu Zhang, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du, Gen Luo, Jun Peng, Xiaoshuai Sun, Rongrong Ji

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02905 2024-08-28 cs.MM 79%

MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model

Sen Wang, Jiangning Zhang, Xin Tan, Zhifeng Xie, Chengjie Wang, Lizhuang Ma

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12117 2024-08-23 cs.CL 79%

The Curious Case of Nonverbal Abstract Reasoning with Multi-Modal Large Language Models

Kian Ahrabian, Zhivar Sourati, Kexuan Sun, Jiarui Zhang, Yifan Jiang, Fred Morstatter, Jay Pujara

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05455 2024-08-13 cs.CV cs.NI 79%

Multimodal generative semantic communication based on latent diffusion model

Weiqi Fu, Lianming Xu, Xin Wu, Haoyang Wei, Li Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11280 2024-08-06 cs.NI cs.AI 79%

Image Generative Semantic Communication with Multi-Modal Similarity Estimation for Resource-Limited Networks

Eri Hosonuma, Taku Yamazaki, Takumi Miyoshi, Akihito Taya, Yuuki Nishiyama, Kaoru Sezaki

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments 14 pages, 15 figures, this paper has been submitted to IEICE Transactions on Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21333 2024-08-01 cs.CV 79%

Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM

Can Wang, Hongliang Zhong, Menglei Chai, Mingming He, Dongdong Chen, Jing Liao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Main paper with supplemental materials

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19180 2024-07-30 cs.CV 79%

Data Processing Techniques for Modern Multimodal Models

Yinheng Li, Han Ding, Hang Chen

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16204 2024-07-24 cs.CV 79%

CLII: Visual-Text Inpainting via Cross-Modal Predictive Interaction

Liang Zhao, Qing Guo, Xiaoguang Li, Song Wang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14936 2024-07-23 cs.MM 79%

EidetiCom: A Cross-modal Brain-Computer Semantic Communication Paradigm for Decoding Visual Perception

Linfeng Zheng, Peilin Chen, Shiqi Wang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11213 2024-07-17 cs.CV 79%

OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models

Zijian Zhou, Zheng Zhu, Holger Caesar, Miaojing Shi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08473 2024-07-12 cs.AR cs.AI 79%

Natural language is not enough: Benchmarking multi-modal generative AI for Verilog generation

Kaiyan Chang, Zhirong Chen, Yunhao Zhou, Wenlong Zhu, kun wang, Haobo Xu, Cangyuan Li, Mengdi Wang, Shengwen Liang, Huawei Li, Yinhe Han, Ying Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by ICCAD 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05808 2024-07-10 cs.CV eess.IV 79%

Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution

Junxiong Lin, Yan Wang, Zeng Tao, Boyang Wang, Qing Zhao, Haorang Wang, Xuan Tong, Xinji Mai, Yuxuan Lin, Wei Song, Jiawen Yu, Shaoqi Yan, Wenqiang Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05340 2024-07-10 cs.CV eess.IV 79%

Unified Multi-Modal Image Synthesis for Missing Modality Imputation

Yue Zhang, Chengtao Peng, Qiuli Wang, Dan Song, Kaiyan Li, S. Kevin Zhou

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE TMI accepted final version

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04736 2024-07-09 eess.SP cs.AI cs.LG 79%

SCDM: Unified Representation Learning for EEG-to-fNIRS Cross-Modal Generation in MI-BCIs

Yisheng Li, Shuqiang Wang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.AI

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00621 2024-07-04 cs.IR cs.MM 79%

Multimodal Pretraining, Adaptation, and Generation for Recommendation: A Survey

Qijiong Liu, Jieming Zhu, Yanting Yang, Quanyu Dai, Zhaocheng Du, Xiao-Ming Wu, Zhou Zhao, Rui Zhang, Zhenhua Dong

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by KDD 2024. See our tutorial materials at https://mmrec.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14109 2024-07-04 cs.AI 79%

Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training

Cheng Tan, Jingxuan Wei, Zhangyang Gao, Linzhuang Sun, Siyuan Li, Ruifeng Guo, Bihui Yu, Stan Z. Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14676 2024-07-02 cs.CV cs.GR 79%

DreamPBR: Text-driven Generation of High-resolution SVBRDF with Multi-modal Guidance

Linxuan Xin, Zheng Zhang, Jinfu Wei, Wei Gao, Duan Gao

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00118 2024-07-02 cs.LG cs.AI 79%

From Efficient Multimodal Models to World Models: A Survey

Xinji Mai, Zeng Tao, Junxiong Lin, Haoran Wang, Yang Chang, Yanlan Kang, Yan Wang, Wenqiang Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18542 2024-06-28 cs.CV eess.SP 79%

Generative AI Empowered LiDAR Point Cloud Generation with Multimodal Transformer

Mohammad Farzanullah, Han Zhang, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 4 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏