arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2507.12761 2025-07-18 cs.CV cs.AI 62%

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

Hanlei Shi, Leyuan Qu, Yu Liu, Di Gao, Yuhua Zheng, Taihao Li

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(杭州高等研究院,中国科学院大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11694 2025-07-17 cs.CL cs.AI 62%

ExpliCIT-QA: Explainable Code-Based Image Table Question Answering

Maximiliano Hormazábal Lagos, Álvaro Bueno Sáez, Pedro Alonso Doval, Jorge Alcalde Vesteiro, Héctor Cerezo-Costas

机构 * Computer Vision Center, Universitat Autònoma de Barcelona(计算机视觉中心,巴塞罗那自治大学) Gradiant

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This work has been accepted for presentation at the 24nd Portuguese Conference on Artificial Intelligence (EPIA 2025) and will be published in the proceedings by Springer in the Lecture Notes in Computer Science (LNCS) series. Please cite the published version when available

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11152 2025-07-16 eess.IV cs.AI cs.CV 62%

Latent Space Consistency for Sparse-View CT Reconstruction

Duoyou Chen, Yunqing Chen, Can Zhang, Zhou Wang, Cheng Chen, Ruoxiu Xiao

机构 * University of Science and Technology Beijing(北京科技大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ACMMM2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11015 2025-07-16 cs.CV cs.AI 62%

Semantically Informed Salient Regions Guided Radiology Report Generation

Zeyi Hou, Zeqiang Wei, Ruixin Yan, Ning Lang, Xiuzhuang Zhou

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学) Department of Radiology, Peking University Third Hospital(北京大学第三医院放射科)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24085 2025-07-15 cs.CV cs.AI 62%

Imagine for Me: Creative Conceptual Blending of Real Images and Text via Blended Attention

Wonwoong Cho, Yanxia Zhang, Yan-Ying Chen, David I. Inouye

机构 * Elmore Family School of Electrical and Computer Engineering, Purdue University(埃尔摩家庭电气与计算机工程学院,普渡大学) Toyota Research Institute(丰田研究院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project website is available at https://imagineforme.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19021 2025-07-15 cs.CV cs.AI 62%

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

Tao Liu, Rongjie Li, Chongyu Wang, Xuming He

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08039 2025-07-14 cs.CV cs.AI cs.LG 62%

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models

Sujith Vemishetty, Advitiya Arora, Anupama Sharma

机构 * Synechron

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05621 2025-07-09 cs.CV cs.MM 62%

AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework

Suoxiang Zhang, Xiaxi Li, Hongrui Chang, Zhuoyan Hou, Guoxin Wu, Ronghua Ji

机构 * China Agricultural University(中国农业大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05302 2025-07-09 cs.CV cs.AI 62%

CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection

Binjia Zhou, Hengrui Lou, Lizhe Chen, Haoyuan Li, Dawei Luo, Shuai Chen, Jie Lei, Zunlei Feng, Yijun Bei

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Ant Group(蚂蚁集团) School of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19179 2025-07-08 cs.CV cs.AI cs.LG 62%

Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning

Dongwei Sun, Jing Yao, Wu Xue, Changsheng Zhou, Pedram Ghamisi, Xiangyong Cao

机构 * School of Computer Science and Technology and the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University(计算机科学与技术学院和教育部智能网络与网络安全重点实验室,西安交通大学) Aerospace Information Research Institute, Chinese Academy of Sciences(航天信息研究所,中国科学院) Space Engineering University(航天工程大学) School of Mathematics and Statistics, Guangdong University of Technology(数学与统计学院,广东工业大学) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍兹中心)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04522 2025-07-08 cs.CV cs.AI cs.RO 62%

Grounded Gesture Generation: Language, Motion, and Space

Anna Deichler, Jim O'Regan, Teo Guichoux, David Johansson, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院) Sorbonne University(索邦大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as a non-archival paper at the CVPR 2025 Humanoid Agents Workshop. Project page: https://groundedgestures.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04954 2025-07-08 cs.CV cs.CL cs.LG 62%

Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation

Xi Zhang, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * School of Computing Science, University of Glasgow(计算科学学院,格拉斯哥大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by BioNLP@ACL 2024

Journal ref Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, ACL 2024, pp. 624-634

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03313 2025-07-08 cs.CV cs.AI 62%

Personalized Image Generation from an Author Writing Style

Sagar Gandhi, Vishal Gandhi

机构 * Joyspace AI

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01055 2025-07-03 eess.IV cs.AI cs.CV 62%

Prompt Mechanisms in Medical Imaging: A Comprehensive Survey

Hao Yang, Xinlong Liang, Zhang Li, Yue Sun, Zheyu Hu, Xinghe Xie, Behdad Dashtbozorg, Jincheng Huang, Shiwei Zhu, Luyi Han, Jiong Zhang, Shanshan Wang, Ritse Mann, Qifeng Yu, Tao Tan

机构 * Faculty of Applied Sciences, Macao Polytechnic University(应用科学学院,澳门理工学院) Netherlands Cancer Institute, Department of Radiology(荷兰癌症研究所,放射科) Radboud University Medical Centre, Department of Radiology and Nuclear Medicine(拉德堡德大学医学中心,放射科和核医学科) College of Aerospace Science and Engineering, National University of Defense Technology(航空航天科学与工程学院,国防科技大学) Medical Department of Breast Cancer, Hunan Cancer Hospital(乳腺癌医学部,湖南癌症医院) the Affiliated Cancer Hospital of Xiangya School of Medicine, Central South University(湘雅医学院附属肿瘤医院,中南大学) Faculty of Biomedical Engineering, Eindhoven University of Technology(生物医学工程学院,埃因霍温理工大学) Laboratory of Advanced Theranostic Materials and Technology, University of Chinese Academy of Sciences(先进诊疗材料与技术实验室,中国科学院大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10493 2025-07-01 cs.CV cs.AI cs.LG 62%

AlignGuard: Scalable Safety Alignment for Text-to-Image Generation

Runtao Liu, I Chieh Chen, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, Fabio Pizzati

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23641 2025-07-01 cs.CV cs.AI 62%

VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation

Peng Huang, Junhu Fu, Bowen Guo, Zeju Li, Yuanyuan Wang, Yi Guo

机构 * College of Biomedical Engineering, Fudan University, Shanghai 200433, China(复旦大学生物医学工程学院) Key Laboratory of Medical Imaging Computing and Computer Assisted Intervention of Shanghai, Shanghai 200032, China(上海医学影像计算与计算机辅助干预重点实验室)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21655 2025-06-30 cs.LG cs.AI cs.CV 62%

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Minjie Hong, Zirun Guo, Yan Xia, Zehan Wang, Ziang Zhang, Tao Jin, Zhou Zhao

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18309 2025-06-27 cs.CV cs.AI 62%

MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement

Xiwei Deng, Xianchun He, Jianfeng Bao, Yudan Zhou, Shuhui Cai, Congbo Cai, Zhong Chen

机构 * Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院) Department of Magnetic Resonance Imaging, The First Affiliated Hospital of Zhengzhou University(郑州大学第一附属医院磁共振成像科) Department of Electronic Science, Xiamen University(厦门大学电子科学系)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20178 2025-06-26 cs.CL cs.AI cs.LG 62%

COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees

Zhiyuan Wang, Jinhao Duan, Qingni Wang, Xiaofeng Zhu, Tianlong Chen, Xiaoshuang Shi, Kaidi Xu

机构 * University of Electronic Science and Technology of China(电子科技大学) Drexel University(德雷塞尔大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17726 2025-06-26 cs.CV cs.CL 62%

VICCA: Visual Interpretation and Comprehension of Chest X-ray Anomalies in Generated Report Without Human Feedback

Sayeh Gholipour Picha, Dawood Al Chanti, Alice Caplier

机构 * Univ. Grenoble Alpes(格勒诺布尔阿尔卑斯大学) CNRS(法国国家科学研究中心) Grenoble INP(格勒诺布尔研究所) GIPSA-lab(GIPSA实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Journal ref Machine Learning with Applications, Volume 21, 2025, 100684, ISSN 2666-8270

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18946 2025-06-25 cs.CV cs.AI 62%

DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models

Zhe Dong, Yuzhe Sun, Tianzhu Liu, Yanfeng Gu

机构 * School of Electronics and Information Engineering, Harbin Institute of Technology(电子信息工程学院,哈尔滨工业大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09174 2025-06-24 cs.CV cs.AI 62%

DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training

Chen Xin, Andreas Hartel, Enkelejda Kasneci

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Corrected minor typos; no changes to results or conclusions

Journal ref Expert Systems with Applications 258 (2024): 125124

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17346 2025-06-24 cs.CV cs.AI 62%

A Novel Multi-layer Task-centric and Data Quality Framework for Autonomous Driving

Yuhan Zhou, Haihua Chen, Kewei Sha

机构 * Dept. of Information Science University of North Texas Denton, Texas, USA(信息科学系 诺克斯维尔大学 德顿 Texas USA)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21264 2025-06-18 cs.CV cs.AI 62%

LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Hanyu Wang, Saksham Suri, Yixuan Ren, Hao Chen, Abhinav Shrivastava

机构 * University of Maryland, College Park(马里兰大学 College Park 分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICLR 2025. Project page: https://hywang66.github.io/larp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08106 2025-06-17 cs.LG cs.AI cs.CV stat.ML 62%

PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation

Ziyan Wang, Sizhe Wei, Xiaoming Huo, Hao Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Rutgers University(罗格斯大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10955 2025-06-13 cs.LG cs.AI cs.CV 62%

ReGuidance: A Simple Diffusion Wrapper for Boosting Sample Quality on Hard Inverse Problems

Aayush Karan, Kulin Shah, Sitan Chen

机构 * Harvard SEAS(哈佛大学SEAS) UT Austin(得克萨斯大学奥斯汀分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 38 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03603 2025-06-13 cs.CV cs.MM 62%

A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation

S. Z. Zhou, Y. B. Wang, J. F. Wu, T. Hu, J. N. Zhang

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.MM

Comments revised

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17152 2025-06-12 cs.CV cs.AI 62%

XMeCap: Meme Caption Generation with Sub-Image Adaptability

Yuyan Chen, Songzhou Yan, Zhihong Zhu, Zhixu Li, Yanghua Xiao

机构 * Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University(复旦大学计算机学院数据科学实验室) Peking University(北京大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08189 2025-06-11 cs.CV cs.CL 62%

Open World Scene Graph Generation using Vision Language Models

Amartya Dutta, Kazi Sajeed Mehrab, Medha Sawhney, Abhilash Neog, Mridul Khurana, Sepideh Fatemi, Aanish Pradhan, M. Maruf, Ismini Lourentzou, Arka Daw, Anuj Karpatne

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted in CVPR 2025 Workshop (CVinW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07235 2025-06-10 cs.CV cs.CL 62%

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Tianyi Bai, Zengjie Hu, Fupeng Sun, Jiantao Qiu, Yizhen Jiang, Guangxin He, Bohan Zeng, Conghui He, Binhang Yuan, Wentao Zhang

机构 * HKUST(香港科技大学) Peking University(北京大学) Shanghai AI Lab(上海人工智能实验室) Imperial College London(伦敦帝国理工学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏