arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2312.07553 2024-05-14 cs.AI cs.CL 81%

Hijacking Context in Large Multi-modal Models

Joonhyun Jeong

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Technical Report. Preprint

Journal ref ICLR 2024 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18591 2024-04-30 cs.CV cs.AI 81%

FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion

Abhishek Kumar Singh, Ioannis Patras

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15100 2024-04-24 cs.CV cs.MM 81%

Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation

Xun Wu, Shaohan Huang, Furu Wei

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13791 2024-04-23 cs.CV cs.AI 81%

Universal Fingerprint Generation: Controllable Diffusion Model with Multimodal Conditions

Steven A. Grosz, Anil K. Jain

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16749 2024-04-18 cs.CV cs.AI eess.IV 81%

MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model

Chunyi Li, Guo Lu, Donghui Feng, Haoning Wu, Zicheng Zhang, Xiaohong Liu, Guangtao Zhai, Weisi Lin, Wenjun Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 page, 11 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17546 2024-04-10 cs.CV cs.AI cs.LG 81%

PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor

Vidit Goel, Elia Peruzzo, Yifan Jiang, Dejia Xu, Xingqian Xu, Nicu Sebe, Trevor Darrell, Zhangyang Wang, Humphrey Shi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in CVPR 2024, Project page https://vidit98.github.io/publication/conference-paper/pair_diff.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07362 2024-04-03 cs.CL cs.CV 81%

Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Seongyun Lee, Sue Hyun Park, Yongrae Jo, Minjoon Seo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00588 2024-04-02 cs.CV cs.AI 81%

Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation

Yitian Tao, Liyan Ma, Jing Yu, Han Zhang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15059 2024-03-25 cs.CV cs.AI 81%

MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration

Zhichao Wei, Qingkun Su, Long Qin, Weizhi Wang

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04302 2024-03-22 cs.CV cs.CL 81%

Prompt Highlighter: Interactive Control for Multi-Modal LLMs

Yuechen Zhang, Shengju Qian, Bohao Peng, Shu Liu, Jiaya Jia

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments CVPR 2024; Project Page: https://julianjuaner.github.io/projects/PromptHighlighter

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04789 2024-03-12 cs.CL cs.AI cs.LG 81%

TopicDiff: A Topic-enriched Diffusion Approach for Multimodal Conversational Emotion Detection

Jiamin Luo, Jingjing Wang, Guodong Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13587 2024-03-08 cs.CL cs.CV 81%

A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation

Yunxin Li, Baotian Hu, Wenhan Luo, Lin Ma, Yuxin Ding, Min Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16117 2024-02-27 cs.RO cs.AI cs.CV 81%

RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

Yao Mu, Junting Chen, Qinglong Zhang, Shoufa Chen, Qiaojun Yu, Chongjian Ge, Runjian Chen, Zhixuan Liang, Mengkang Hu, Chaofan Tao, Peize Sun, Haibao Yu, Chao Yang, Wenqi Shao, Wenhai Wang, Jifeng Dai, Yu Qiao, Mingyu Ding, Ping Luo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07276 2024-01-30 cs.CL cs.AI cs.LG q-bio.BM 81%

BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations

Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, Rui Yan

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by Empirical Methods in Natural Language Processing 2023 (EMNLP 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13298 2024-01-25 cs.CL cs.AI 81%

Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models

Hongzhan Lin, Ziyang Luo, Wei Gao, Jing Ma, Bo Wang, Ruichao Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments The first work towards explainable harmful meme detection by harnessing advanced LLMs

Journal ref The ACM Web Conference 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11631 2024-01-23 cs.CV cs.CL cs.LG 81%

Text-to-Image Cross-Modal Generation: A Systematic Review

Maciej Żelaszczyk, Jacek Mańdziuk

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05134 2024-01-11 cs.AI cs.CL 81%

Yes, this is what I was looking for! Towards Multi-modal Medical Consultation Concern Summary Generation

Abhisek Tiwari, Shreyangshu Bera, Sriparna Saha, Pushpak Bhattacharyya, Samrat Ghosh

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02433 2024-01-08 cs.CV cs.AI cs.LG 81%

FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-Clients

DaiXun Li, Weiying Xie, ZiXuan Wang, YiBing Lu, Yunsong Li, Leyuan Fang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15296 2023-12-21 cs.CV cs.AI cs.LG 81%

MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image Generation

Marco Bellagente, Manuel Brack, Hannah Teufel, Felix Friedrich, Björn Deiseroth, Constantin Eichenberg, Andrew Dai, Robert Baldock, Souradeep Nanda, Koen Oostermeijer, Andres Felipe Cruz-Salinas, Patrick Schramowski, Kristian Kersting, Samuel Weinbach

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Proceedings of Advances in Neural Information Processing Systems: Annual Conference on Neural Information Processing Systems (NeurIPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09767 2023-12-20 cs.CL cs.AI 81%

VLIS: Unimodal Language Models Guide Multimodal Language Generation

Jiwan Chung, Youngjae Yu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted as main paper in EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06647 2023-12-12 cs.CV cs.AI cs.LG 81%

4M: Massively Multimodal Masked Modeling

David Mizrahi, Roman Bachmann, Oğuzhan Fatih Kar, Teresa Yeo, Mingfei Gao, Afshin Dehghan, Amir Zamir

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2023 Spotlight. Project page at https://4m.epfl.ch/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09909 2023-12-05 cs.CV cs.CL 81%

Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

Chaoyi Wu, Jiayu Lei, Qiaoyu Zheng, Weike Zhao, Weixiong Lin, Xiaoman Zhang, Xiao Zhou, Ziheng Zhao, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16483 2023-11-29 cs.CV cs.CL 81%

ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Yucheng Han, Chi Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, Hanwang Zhang

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Code and model on https://tingxueronghua.github.io/ChartLlama/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10056 2023-11-03 cs.CV cs.MM 81%

GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation

Can Qin, Ning Yu, Chen Xing, Shu Zhang, Zeyuan Chen, Stefano Ermon, Yun Fu, Caiming Xiong, Ran Xu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15585 2023-10-25 cs.CL cs.CV cs.LG 81%

Multimodal Representations for Teacher-Guided Compositional Visual Reasoning

Wafa Aissa, Marin Ferecatu, Michel Crucianu

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.CL

Journal ref Advanced Concepts for Intelligent Vision Systems, 21st International Conference (ACIVS 2023), Aug 2023, Kumamoto, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14296 2023-10-24 cs.CV cs.MM 81%

Research on Key Technologies of Infrastructure Digitalization based on Multimodal Spatial Data

Zhanyuan Tian, Tianrui Zhu, Zerui Tian, Zhen Dong

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 20 pages, in Chinese language, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04175 2023-10-24 cs.CR cs.CV cs.MM 81%

Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data Poisoning

Shengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu, Yuejian Fang, Hang Su

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Carmera-ready version. To appear in ACM MM 2023. Code will be released at: https://github.com/sf-zhai/BadT2I

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09755 2023-10-19 cs.CV cs.AI 81%

Beyond Segmentation: Road Network Generation with Multi-Modal LLMs

Sumedh Rasal, Sanjay Kumar Boddhu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08027 2023-10-13 cs.CL cs.CV 81%

Exploring Large Language Models for Multi-Modal Out-of-Distribution Detection

Yi Dai, Hao Lang, Kaisheng Zeng, Fei Huang, Yongbin Li

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments EMNLP2023 Findings Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15564 2023-09-29 cs.LG cs.CL cs.CV 81%

Jointly Training Large Autoregressive Multimodal Models

Emanuele Aiello, Lili Yu, Yixin Nie, Armen Aghajanyan, Barlas Oguz

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏