arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2403.14141 2024-03-22 cs.CV 79%

Empowering Segmentation Ability to Multi-modal Large Language Models

Yuqi Yang, Peng-Tao Jiang, Jing Wang, Hao Zhang, Kai Zhao, Jinwei Chen, Bo Li

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16274 2024-03-22 cs.CV 79%

Towards Flexible, Scalable, and Adaptive Multi-Modal Conditioned Face Synthesis

Jingjing Ren, Cheng Xu, Haoyu Chen, Xinran Qin, Lei Zhu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08460 2024-03-20 cs.CV cs.RO 79%

Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model

Ruibin Zhang, Donglai Xue, Yuhan Wang, Ruixu Geng, Fei Gao

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures, submitted to RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17336 2024-03-13 cs.CV cs.RO 79%

Robust 3D Object Detection from LiDAR-Radar Point Clouds via Cross-Modal Feature Augmentation

Jianning Deng, Gabriel Chan, Hantao Zhong, Chris Xiaoxuan Lu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to ICRA 2024. 8 pages, 4 figures. Equal contribution for Gabriel Chan and Hantao Zhong, listed randomly

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06470 2024-03-12 cs.CV 79%

3D-aware Image Generation and Editing with Multi-modal Conditions

Bo Li, Yi-ke Li, Zhi-fen He, Bin Liu, Yun-Kun Lai

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04290 2024-03-08 eess.IV cs.CV cs.LG 79%

MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant

Chenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang, Jian Wu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04014 2024-03-08 cs.HC cs.AI 79%

PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement

Zhijie Wang, Yuheng Huang, Da Song, Lei Ma, Tianyi Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments To appear in the 2024 CHI Conference on Human Factors in Computing Systems (CHI '24), May 11--16, 2024, Honolulu, HI, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10855 2024-02-19 cs.CV 79%

Control Color: Multimodal Diffusion-based Interactive Image Colorization

Zhexin Liang, Zhaochen Li, Shangchen Zhou, Chongyi Li, Chen Change Loy

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Project Page: https://zhexinliang.github.io/Control_Color/; Demo Video: https://youtu.be/tSCwA-srl8Q

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09036 2024-02-15 cs.CV 79%

Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?

Tiantian Feng, Daniel Yang, Digbalay Bose, Shrikanth Narayanan

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05803 2024-02-09 cs.CV cs.GR 79%

AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning

Wamiq Reyaz Para, Abdelrahman Eldesokey, Zhenyu Li, Pradyumna Reddy, Jiankang Deng, Peter Wonka

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06911 2024-02-07 cs.LG cs.CL q-bio.BM 79%

GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text

Pengfei Liu, Yiming Ren, Jun Tao, Zhixiang Ren

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments The article has been accepted by Computers in Biology and Medicine, with 14 pages and 4 figures

Journal ref Computers in Biology and Medicine, 108073, 2024, ISSN 0010-4825

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00375 2024-02-02 eess.IV cs.CV 79%

Disentangled Multimodal Brain MR Image Translation via Transformer-based Modality Infuser

Jihoon Cho, Xiaofeng Liu, Fangxu Xing, Jinsong Ouyang, Georges El Fakhri, Jinah Park, Jonghye Woo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17664 2024-02-01 cs.CV cs.GR 79%

Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation

Yuanhuiyi Lyu, Xu Zheng, Lin Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01827 2024-01-04 cs.CV 79%

Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions

David Junhao Zhang, Dongxu Li, Hung Le, Mike Zheng Shou, Caiming Xiong, Doyen Sahoo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments project page: https://showlab.github.io/Moonshot/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17611 2024-01-01 cs.CV 79%

P2M2-Net: Part-Aware Prompt-Guided Multimodal Point Cloud Completion

Linlian Jiang, Pan Chen, Ye Wang, Tieru Wu, Rui Ma

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Best Poster Award of CAD/Graphics 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15900 2023-12-27 cs.CV 79%

Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

Zunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li, Xiu Li

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments AAAI-2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04445 2023-12-19 cs.LG cs.CV 79%

Multi-modal Latent Diffusion

Mustapha Bounoua, Giulio Franzese, Pietro Michiardi

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10512 2023-12-19 cs.CL cs.HC 79%

IMAD: IMage-Augmented multi-modal Dialogue

Viktor Moskvoretskii, Anton Frolov, Denis Kuznetsov

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Main part contains 6 pages, 4 figures. It was accepted on AINL. We wait the publication and DOI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08762 2023-12-15 cs.AI 79%

Multi-modal Latent Space Learning for Chain-of-Thought Reasoning in Language Models

Liqi He, Zuchao Li, Xiantao Cai, Ping Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07944 2023-12-05 cs.AI 79%

AutoRepo: A general framework for multi-modal LLM-based automated construction reporting

Hongxu Pu, Xincong Yang, Jing Li, Runhao Guo, Heng Li

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments We believe that keeping this version of the paper publicly available may lead to confusion or misinterpretation regarding our current research direction and findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10868 2023-12-04 cs.CL 79%

Retrieving Multimodal Information for Augmented Generation: A Survey

Ruochen Zhao, Hailin Chen, Weishi Wang, Fangkai Jiao, Xuan Long Do, Chengwei Qin, Bosheng Ding, Xiaobao Guo, Minzhi Li, Xingxuan Li, Shafiq Joty

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05199 2023-11-10 cs.CV 79%

BrainNetDiff: Generative AI Empowers Brain Network Generation via Multimodal Diffusion Model

Yongcheng Zong, Shuqiang Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20251 2023-11-01 cs.MM 79%

An Implementation of Multimodal Fusion System for Intelligent Digital Human Generation

Yingjie Zhou, Yaodong Chen, Kaiyue Bi, Lian Xiong, Hui Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16131 2023-10-26 cs.CL 79%

GenKIE: Robust Generative Multimodal Document Key Information Extraction

Panfeng Cao, Ye Wang, Qiang Zhang, Zaiqiao Meng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2023, Findings paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13787 2023-10-24 cs.LG cs.AI 79%

Enhancing Illicit Activity Detection using XAI: A Multimodal Graph-LLM Framework

Jack Nicholls, Aditya Kuppa, Nhien-An Le-Khac

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05355 2023-10-10 cs.CV 79%

C^2M-DoT: Cross-modal consistent multi-view medical report generation with domain transfer network

Ruizhi Wang, Xiangtao Wang, Jie Zhou, Thomas Lukasiewicz, Zhenghua Xu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03364 2023-09-08 cs.SD eess.AS 79%

Highly Controllable Diffusion-based Any-to-Any Voice Conversion Model with Frame-level Prosody Feature

Kyungguen Byun, Sunkuk Moon, Erik Visser

专题命中 多模态生成 :any-to-any(title,abstract);分类 eess.AS

Comments 5 pages, 3 figures, submitted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01981 2023-09-06 cs.RO cs.AI cs.GR 79%

Graph-Based Interaction-Aware Multimodal 2D Vehicle Trajectory Prediction using Diffusion Graph Convolutional Networks

Keshu Wu, Yang Zhou, Haotian Shi, Xiaopeng Li, Bin Ran

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14748 2023-08-29 cs.GR cs.CV 79%

MagicAvatar: Multimodal Avatar Generation and Animation

Jianfeng Zhang, Hanshu Yan, Zhongcong Xu, Jiashi Feng, Jun Hao Liew

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://magic-avatar.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.13592 2023-08-25 cs.CV 79%

Multimodal Image Synthesis and Editing: The Generative AI Era

Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, Shijian Lu, Lingjie Liu, Adam Kortylewski, Christian Theobalt, Eric Xing

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments TPAMI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏