arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2312.03011 2024-02-16 cs.CV cs.AI 62%

InstructBooth: Instruction-following Personalized Text-to-Image Generation

Daewon Chae, Nokyung Park, Jinkyu Kim, Kimin Lee

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01241 2024-02-05 cs.CV cs.AI 62%

Can Shape-Infused Joint Embeddings Improve Image-Conditioned 3D Diffusion?

Cristian Sbrolli, Paolo Cudrano, Matteo Matteucci

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12456 2024-01-24 cs.CV cs.AI cs.GR 62%

Exploration and Improvement of Nerf-based 3D Scene Editing Techniques

Shun Fang, Ming Cui, Xing Feng, Yanan Zhang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07709 2024-01-24 cs.CV cs.AI 62%

Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks

Siyu Zou, Jiji Tang, Yiyi Zhou, Jing He, Chaoyi Zhao, Rongsheng Zhang, Zhipeng Hu, Xiaoshuai Sun

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06780 2024-01-17 eess.IV cs.AI cs.CV 62%

HA-HI: Synergising fMRI and DTI through Hierarchical Alignments and Hierarchical Interactions for Mild Cognitive Impairment Diagnosis

Xiongri Shen, Zhenxi Song, Linling Li, Min Zhang, Lingyan Liang Honghai Liu, Demao Deng, Zhiguo Zhang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01811 2024-01-15 cs.CV cs.AI 62%

DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder

Tao Liu, Chenpeng Du, Shuai Fan, Feilong Chen, Kai Yu

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.AI

Comments 5 pages, Accepted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12604 2024-01-15 cs.CV cs.CL 62%

PromptMRG: Diagnosis-Driven Prompts for Medical Report Generation

Haibo Jin, Haoxuan Che, Yi Lin, Hao Chen

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted to AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02173 2024-01-05 cs.CV cs.AI 62%

Prompt Decoupling for Text-to-Image Person Re-identification

Weihao Li, Lei Tan, Pingyang Dai, Yan Zhang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16012 2023-12-27 cs.CV cs.AI 62%

Detection-based Intermediate Supervision for Visual Question Answering

Yuhang Liu, Daowan Peng, Wei Wei, Yuanyuan Fu, Wenfeng Xie, Dangyang Chen

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI24

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15770 2023-12-27 cs.CV cs.AI 62%

A Recipe for Scaling up Text-to-Video Generation with Text-free Videos

Xiang Wang, Shiwei Zhang, Hangjie Yuan, Zhiwu Qing, Biao Gong, Yingya Zhang, Yujun Shen, Changxin Gao, Nong Sang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Comments Project page: https://tf-t2v.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15247 2023-12-27 cs.CV cs.AI 62%

Prompt-Propose-Verify: A Reliable Hand-Object-Interaction Data Generation Framework using Foundational Models

Gurusha Juneja, Sukrit Kumar

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted at International Workshop on AI for Digital Human in AAAI Conference on Articial Intelligence (AAAI, 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08870 2023-12-19 cs.CV cs.AI 62%

UPGPT: Universal Diffusion Model for Person Image Generation, Editing and Pose Transfer

Soon Yau Cheong, Armin Mustafa, Andrew Gilbert

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07550 2023-12-14 cs.CV cs.CL cs.CR cs.LG 62%

Understanding (Un)Intended Memorization in Text-to-Image Generative Models

Ali Naseh, Jaechul Roh, Amir Houmansadr

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07066 2023-12-13 cs.CL cs.CV 62%

DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising Models

Shengguang Wu, Mei Yuan, Qi Su

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06225 2023-12-12 cs.CV cs.AI 62%

DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Fa-Ting Hong, Li Shen, Dan Xu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at TPAMI; CVPR 2022 extension

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04549 2023-12-08 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 62%

PlayFusion: Skill Acquisition via Diffusion from Language-Annotated Play

Lili Chen, Shikhar Bahl, Deepak Pathak

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments In CoRL 2023. Website at https://play-fusion.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14702 2023-12-08 cs.CV cs.AI cs.RO 62%

BM2CP: Efficient Collaborative Perception with LiDAR-Camera Modalities

Binyu Zhao, Wei Zhang, Zhaonian Zou

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 8 figures. Accepted by CoRL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05189 2023-11-30 cs.CL cs.CV 62%

SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language Models

Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin, Liang Lin

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09520 2023-11-21 cs.CV cs.AI 62%

MDFL: Multi-domain Diffusion-driven Feature Learning

Daixun Li, Weiying Xie, Jiaqing Zhang, Yunsong Li

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02329 2023-11-13 cs.CV cs.AI 62%

Complex Organ Mask Guided Radiology Report Generation

Tiancheng Gu, Dongnan Liu, Zhiyuan Li, Weidong Cai

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 7 images. Accepted by WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05464 2023-11-10 cs.CV cs.MM 62%

3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion Models

Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao, Zhineng Chen, Tao Mei

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05463 2023-11-10 cs.CV cs.MM 62%

ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors

Jingwen Chen, Yingwei Pan, Ting Yao, Tao Mei

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04260 2023-11-09 cs.RO cs.CL cs.CV 62%

Fully Automated Task Management for Generation, Execution, and Evaluation: A Framework for Fetch-and-Carry Tasks with Natural Language Instructions in Continuous Space

Motonari Kambara, Komei Sugiura

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted at presentation for CVPR 2023 Embodied AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00399 2023-11-02 cs.CV cs.CL 62%

Enhanced Knowledge Injection for Radiology Report Generation

Qingqiu Li, Jilan Xu, Runtian Yuan, Mohan Chen, Yuejie Zhang, Rui Feng, Xiaobo Zhang, Shang Gao

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by BIBM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17372 2023-10-27 cs.AI cs.CL cs.RO 62%

Dialogue-based generation of self-driving simulation scenarios using Large Language Models

Antonio Valerio Miceli-Barone, Alex Lascarides, Craig Innes

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 6 figures, SpLU-RoboNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15355 2023-10-25 cs.CL cs.AI 62%

Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language Generation

Adam Bouyamourn

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15319 2023-10-25 cs.CL cs.AI cs.LG 62%

Hallucination Detection for Grounded Instruction Generation

Lingjun Zhao, Khanh Nguyen, Hal Daumé

专题命中 多模态生成 :image-text(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11210 2023-10-18 cs.CV cs.MM 62%

Learning Comprehensive Representations with Richer Self for Text-to-Image Person Re-Identification

Shuanglin Yan, Neng Dong, Jun Liu, Liyan Zhang, Jinhui Tang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05881 2023-10-10 cs.CV cs.CL 62%

Controllable Chest X-Ray Report Generation from Longitudinal Representations

Francesco Dalla Serra, Chaoyang Wang, Fani Deligianni, Jeffrey Dalton, Alison Q O'Neil

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to the Findings of EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.07495 2023-09-15 cs.CV cs.AI 62%

HDTR-Net: A Real-Time High-Definition Teeth Restoration Network for Arbitrary Talking Face Generation Methods

Yongyuan Li, Xiuyuan Qin, Chao Liang, Mingqiang Wei

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 15pages, 6 figures, PRCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏