arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2503.08156 2025-03-12 cs.CV cs.LG 79%

Towards Large-scale Chemical Reaction Image Parsing via a Multimodal Large Language Model

Yufan Chen, Ching Ting Leung, Jianwei Sun, Yong Huang, Linyan Li, Hao Chen, Hanyu Gao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06296 2025-03-11 cs.CL cs.LG 79%

MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering

Vinay Kumar Verma, Shreyas Sunil Kulkarni, Happy Mittal, Deepak Gupta

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments To appear at NAACL Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01115 2025-03-11 cs.CV 79%

WeGen: A Unified Model for Interactive Multimodal Generation as We Chat

Zhipeng Huang, Shaobin Zhuang, Canmiao Fu, Binxin Yang, Ying Zhang, Chong Sun, Zhizheng Zhang, Yali Wang, Chen Li, Zheng-Jun Zha

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10171 2025-03-11 cs.RO cs.AI 79%

Imagine-2-Drive: Leveraging High-Fidelity World Models via Multi-Modal Diffusion Policies

Anant Garg, K Madhava Krishna

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Submitted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16795 2025-03-11 cs.AI 79%

Scene-Aware Explainable Multimodal Trajectory Prediction

Pei Liu, Haipeng Liu, Xingyu Liu, Yiqun Li, Junlan Chen, Yangfan He, Jun Ma

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06142 2025-03-11 cs.CV 79%

VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models

Xinan He, Yue Zhou, Bing Fan, Bin Li, Guopu Zhu, Feng Ding

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14922 2025-03-04 cs.CV 79%

GDTS: Goal-Guided Diffusion Model with Tree Sampling for Multi-Modal Pedestrian Trajectory Prediction

Ge Sun, Sheng Wang, Lei Zhu, Ming Liu, Jun Ma

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00413 2025-03-04 cs.CV cs.LG 79%

CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering

Tianyu Huai, Jie Zhou, Xingjiao Wu, Qin Chen, Qingchun Bai, Ze Zhou, Liang He

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages,4 figures,accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19293 2025-02-28 cs.CV 79%

Pathology Report Generation and Multimodal Representation Learning for Cutaneous Melanocytic Lesions

Ruben T. Lucassen, Sander P. J. Moonemans, Tijn van de Luijtgaarden, Gerben E. Breimer, Willeke A. M. Blokx, Mitko Veta

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CV

Comments 11 pages, 2 figures. arXiv admin note: text overlap with arXiv:2502.19285

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13452 2025-02-28 cs.CV 79%

EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion

Jiangchuan Wei, Shiyue Yan, Wenfeng Lin, Boyuan Liu, Renjie Chen, Mingyu Guo

专题命中 多模态生成 :multimodal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13656 2025-02-25 cs.CV 79%

GLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation Detection

Xiaocan Chen, Qilin Yin, Jiarui Liu, Wei Lu, Xiangyang Luo, Jiantao Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11453 2025-02-18 cs.LG cs.AI 79%

Connector-S: A Survey of Connectors in Multi-modal Large Language Models

Xun Zhu, Zheng Zhang, Xi Chen, Yiming Shi, Miao Li, Ji Wu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10458 2025-02-18 cs.LG cs.AI 79%

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models

Zhenxing Mi, Kuan-Chieh Wang, Guocheng Qian, Hanrong Ye, Runtao Liu, Sergey Tulyakov, Kfir Aberman, Dan Xu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Project page: https://mizhenxing.github.io/ThinkDiff, 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06823 2025-02-13 cs.LG cs.CV cs.GR cs.IR 79%

CTR-Driven Advertising Image Generation with Multimodal Large Language Models

Xingye Chen, Wei Feng, Zhenbang Du, Weizhen Wang, Yanyin Chen, Haohan Wang, Linkai Liu, Yaoyu Li, Jinyuan Zhao, Yu Li, Zheng Zhang, Jingjing Lv, Junjie Shen, Zhangang Lin, Jingping Shao, Yuanjie Shao, Xinge You, Changxin Gao, Nong Sang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06474 2025-02-11 cs.CV 79%

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Weijia Mao, Zhenheng Yang, Mike Zheng Shou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18102 2025-02-11 cs.NE cs.AI 79%

Multiple Global Peaks Big Bang-Big Crunch Algorithm for Multimodal Optimization

Fabio Stroppa, Ahmet Astar

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 16 pages

Journal ref Evolutionary Intelligence, Springer, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05178 2025-02-10 cs.CV 79%

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Yue Zhao, Fuzhao Xue, Scott Reed, Linxi Fan, Yuke Zhu, Jan Kautz, Zhiding Yu, Philipp Krähenbühl, De-An Huang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Tech report. Project page: https://nvlabs.github.io/QLIP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05150 2025-02-10 cs.CL 79%

CodeSCM: Causal Analysis for Multi-Modal Code Generation

Mukur Gupta, Noopur Bhatt, Suman Jana

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02555 2025-02-06 eess.IV cs.CV 79%

AAD-DCE: An Aggregated Multimodal Attention Mechanism for Early and Late Dynamic Contrast Enhanced Prostate MRI Synthesis

Divya Bharti, Sriprabha Ramanarayanan, Sadhana S, Kishore Kumar M, Keerthi Ram, Harsh Agarwal, Ramesh Venkatesan, Mohanasankar Sivaprakasam

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13925 2025-01-24 cs.CV 79%

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad S. Khan, Salman Khan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12173 2025-01-22 cs.CV 79%

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

Shiyue Zhang, Zheng Chong, Xi Lu, Wenqing Zhang, Haoxiang Li, Xujie Zhang, Jiehui Huang, Xiao Dong, Xiaodan Liang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11233 2025-01-22 cs.IR cs.CL cs.MA 79%

PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents

Kanika Goswami, Puneet Mathur, Ryan Rossi, Franck Dernoncourt

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at ECIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07086 2025-01-14 cs.CL 79%

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models

Yongyu Mu, Hengyu Li, Junxin Wang, Xiaoxuan Zhou, Chenglong Wang, Yingfeng Luo, Qiaozhi He, Tong Xiao, Guocheng Chen, Jingbo Zhu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02167 2025-01-07 cs.CV 79%

Generating Multimodal Images with GAN: Integrating Text, Image, and Style

Chaoyi Tan, Wenqing Zhang, Zhen Qi, Kowei Shih, Xinshi Li, Ao Xiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20725 2024-12-31 cs.CV 79%

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

Min Zhang, Zilin Wang, Liyan Chen, Kunhong Liu, Juncong Lin

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08603 2024-12-31 cs.CL 79%

A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

Hanguang Xiao, Feizhong Zhou, Xingyue Liu, Tianqi Liu, Zhipeng Li, Xin Liu, Xiaoxuan Huang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Journal ref Information Fusion, 117 (2025) 102888

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18928 2024-12-30 cs.CV cs.LG 79%

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation

Lunhao Duan, Shanshan Zhao, Wenjun Yan, Yinglun Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Mingming Gong, Gui-Song Xia

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17225 2024-12-24 cs.CV 79%

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Lichen Ma, Tiezhu Yue, Pei Fu, Yujie Zhong, Kai Zhou, Xiaoming Wei, Jie Hu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13129 2024-12-24 cs.CV cs.LG 79%

M3T: Multi-Modal Medical Transformer to bridge Clinical Context with Visual Insights for Retinal Image Medical Description Generation

Nagur Shareef Shaik, Teja Krishna Cherukuri, Dong Hye Ye

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments This paper has been accepted for presentation at the IEEE International Conference on Image Processing (ICIP 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04363 2024-12-19 cs.CV 79%

Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs

Junhao Chen, Xiang Li, Xiaojun Ye, Chao Li, Zhaoxin Fan, Hao Zhao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by COLING 2025 (The 31st International Conference on Computational Linguistics) Project Page: https://idea23d.github.io/ Code: https://github.com/yisuanwang/Idea23D

详情

展开后加载摘要…

URL PDF HTML 收藏