arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2511.00362 2025-11-04 cs.CV cs.AI cs.GR 62%

Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery

Momen Khandoker Ope, Akif Islam, Mohd Ruhul Ameen, Abu Saleh Musa Miah, Md Rashedul Islam, Jungpil Shin

机构 * University of Rajshahi(拉贾沙希大学) Marshall University(马歇尔大学) University of Aizu(御所大学) University of Asia Pacific(亚洲太平洋大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 6 Pages, 4 figures, 2 Tables, Submitted to ICECTE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00107 2025-11-04 cs.CV cs.AI cs.IR 62%

AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency

Piyushkumar Patel

机构 * Microsoft(微软)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23020 2025-10-28 cs.CV cs.CL 62%

M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark

Huixuan Zhang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(计算机技术研究所,北京大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06771 2025-10-27 cs.AI cs.CV cs.LG 62%

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang

机构 * Google DeepMind(谷歌DeepMind)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Journal ref International Conference on Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19641 2025-10-23 cs.CL cs.AI 62%

Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent

Yangshijie Zhang, Xinda Wang, Jialin Liu, Wenqiang Wang, Zhicong Ma, Xingxing Jia

机构 * Lanzhou University(兰州大学) Peking University(北京大学) Sun Yat-sen University(中山大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17519 2025-10-23 cs.CV cs.AI 62%

MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models

Yongshun Zhang, Zhongyi Fan, Yonghang Zhang, Zhangzikang Li, Weifeng Chen, Zhongwei Feng, Chaoyue Wang, Peng Hou, Anxiang Zeng

机构 * LLM Team, Shopee Pte. Ltd.(Shopee 股份有限公司语言模型团队)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Technical Report; Project Page: https://github.com/Shopee-MUG/MUG-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00939 2025-10-23 cs.CV cs.CL 62%

WikiVideo: Article Generation from Multiple Videos

Alexander Martin, Reno Kriz, William Gantt Walden, Kate Sanders, Hannah Recknor, Eugene Yang, Francis Ferraro, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学) Human Language Technology Center of Excellence(人机语言技术卓越中心) University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Repo can be found here: https://github.com/alexmartin1722/wikivideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02048 2025-10-22 eess.IV cs.AI cs.CV 62%

Regression is all you need for medical image translation

Sebastian Rassmann, David Kügler, Christian Ewert, Martin Reuter

机构 * German Center for Neurodegenerative Diseases (DZNE)(德国神经退行性疾病研究中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16844 2025-10-21 cs.CL cs.AI cs.CE 62%

FinSight: Towards Real-World Financial Deep Research

Jiajie Jin, Yuyao Zhang, Yimeng Xu, Hongjin Qian, Yutao Zhu, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学耿丽人工智能学院) BAAI(北京人工智能研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15176 2025-10-20 cs.CV cs.AI 62%

Methods and Trends in Detecting AI-Generated Images: A Comprehensive Review

Arpan Mahara, Naphtali Rishe

机构 * Knight Foundation School of Computing and Information Sciences, Florida International University(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 34 pages, 4 Figures, 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18668 2025-10-17 cs.CV cs.CL 62%

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation

Zhen Li, Duan Li, Yukai Guo, Xinyuan Guo, Bowen Li, Lanxi Xiao, Shenyu Qiao, Jiashu Chen, Zijian Wu, Hui Zhang, Xinhuan Shu, Shixia Liu

机构 * Tsinghua University(清华大学) Newcastle University(新castle大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 58 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08980 2025-10-14 cs.LG cs.AI cs.CV 62%

Learning Diffusion Models with Flexible Representation Guidance

Chenyu Wang, Cai Zhou, Sharut Gupta, Zongyu Lin, Stefanie Jegelka, Stephen Bates, Tommi Jaakkola

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025; Also Oral at ICML 2025 FM4LS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05976 2025-10-08 cs.CV cs.AI cs.LG 62%

Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis

Eashan Adhikarla, Yixin Liu, Brian D. Davison

机构 * Lehigh University(莱维大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15046 2025-10-08 cs.CL cs.AI 62%

ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding

Yifan Wu, Lutao Yan, Leixian Shen, Yinan Mei, Jiannan Wang, Yuyu Luo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Need to be revised

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 62%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04498 2025-10-07 cs.CL cs.AI 62%

GenQuest: An LLM-based Text Adventure Game for Language Learners

Qiao Wang, Adnan Labib, Robert Swier, Michael Hofmeyr, Zheng Yuan

机构 * Hosei University(立命馆大学) King’s College London(伦敦大学国王学院) Kindai University(_kindai大学) Tokyo Uni. of Science(东京科学大学) University of Sheffield(谢菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Workshop on Wordplay: When Language Meets Games, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 62%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24251 2025-10-07 cs.CV cs.CL 62%

Latent Visual Reasoning

Bangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang, Jialian Wu, Xiaodong Yu, Hao Chen, Emad Barsoum, Muhao Chen, Zicheng Liu

机构 * University of California, Davis(加州大学戴维斯分校) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16612 2025-10-07 cs.HC cs.AI cs.CV cs.CY 62%

Negative Shanshui: Real-time Interactive Ink Painting Synthesis

Aven-Le Zhou

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00046 2025-10-02 cs.CV cs.AI 62%

Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models

Xiaotian Zou

机构 * Xiaotian Zou

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25817 2025-10-01 cs.CL cs.CV 62%

Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer

Jaeyoung Kim, Jongho Lee, Hongjun Choi, Sion Jang

机构 * Teamreboott Inc.(Teamreboott公司) MIRI D.I.H Inc.(MIRI D.I.H公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22940 2025-09-30 cs.CL cs.CV 62%

LLMs Behind the Scenes: Enabling Narrative Scene Illustration

Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max Kreminski

机构 * Midjourney Northwestern University(西北大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21887 2025-09-29 cs.CV cs.MM 62%

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

Liyang Chen, Tianze Zhou, Xu He, Boshi Tang, Zhiyong Wu, Yang Huang, Yang Wu, Zhongqian Sun, Wei Yang, Helen Meng

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21375 2025-09-29 cs.CV cs.AI 62%

Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis

Aleksa Jelaca, Ying Jiao, Chang Tian, Marie-Francine Moens

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments text-to-image generation, automatic prompt, DPO, Counterfactual

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16663 2025-09-26 cs.CL cs.AI 62%

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

Yujin Han, Hao Chen, Andi Han, Zhiheng Wang, Xinyu Liu, Yingya Zhang, Shiwei Zhang, Difan Zou

专题命中 多模态生成 :MLLM(abstract);分类 cs.CL、cs.AI

Comments 31 pages, 16 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18907 2025-09-26 cs.AI cs.CV cs.RO 62%

EC-Diffuser: Multi-Object Manipulation via Entity-Centric Behavior Generation

Carl Qi, Dan Haramati, Tal Daniel, Aviv Tamar, Amy Zhang

机构 * UT Austin(得克萨斯大学) Technion, Israel Institute of Technology(技术学院,以色列技术学院) Brown University(布朗大学) Meta AI

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18179 2025-09-24 cs.CV cs.AI 62%

The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes

Sai Varun Kodathala, Rakesh Vunnam

机构 * Sports Vision, Inc.(体育视觉公司) Vizworld, Inc.(Vizworld公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13794 2025-09-22 cs.CV cs.AI 62%

LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

Yang Zhou, Shiyu Zhao, Yuxiao Chen, Zhenting Wang, Can Jin, Dimitris N. Metaxas

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15270 2025-09-22 cs.CV cs.AI 62%

PRISM: Phase-enhanced Radial-based Image Signature Mapping framework for fingerprinting AI-generated images

Emanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci, Roberto Di Pietro

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12888 2025-09-17 cs.CV cs.AI 62%

Runge-Kutta Approximation and Decoupled Attention for Rectified Flow Inversion and Semantic Editing

Weiming Chen, Zhihan Zhu, Yijia Wang, Zhihai He

机构 * Southern University of Science and Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏