arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4955 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4955 篇

2507.00983 2025-07-31 eess.IV cs.CV 70%

DMCIE: Diffusion Model with Concatenation of Inputs and Errors to Improve the Accuracy of the Segmentation of Brain Tumors in MRI Images

Sara Yavari, Rahul Nitin Pandya, Jacob Furst

机构 * School of Computing, DePaul University(计算学院,德保罗大学)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17046 2025-07-29 cs.CV 70%

Text-to-Image Generation Via Energy-Based CLIP

Roy Ganz, Michael Elad

机构 * Electrical Engineering Department Technion(技术学院电子工程系) Computer Science Department Technion(技术学院计算机科学系)

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07919 2025-07-28 cs.CL q-bio.BM 70%

Advancing biomolecular understanding and design following human instructions

Xiang Zhuang, Keyan Ding, Tianwen Lyu, Yinuo Jiang, Xiaotong Li, Zhuoyi Xiang, Zeyuan Wang, Ming Qin, Kehua Feng, Jike Wang, Qiang Zhang, Huajun Chen

专题命中 多模态生成 :multimodal(abstract);any-to-any(abstract);分类 cs.CL

Journal ref Nature Machine Intelligence volume 7, pages1154-1167 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15249 2025-07-22 cs.CV 70%

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

Yanbing Zhang, Zhe Wang, Qin Zhou, Mengping Yang

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18397 2025-07-16 cs.CV 70%

Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization

Kesen Zhao, Beier Zhu, Qianru Sun, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) Singapore Management University(新加坡管理大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05259 2025-07-08 cs.CV 70%

Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

Chun-Hsiao Yeh, Yilin Wang, Nanxuan Zhao, Richard Zhang, Yuheng Li, Yi Ma, Krishna Kumar Singh

机构 * UC Berkeley(伯克利大学) HKU(香港大学) Adobe(Adobe公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Project page: https://danielchyeh.github.io/x-planner/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04451 2025-07-08 cs.CV 70%

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step

Zheyuan Liu, Munan Ning, Qihui Zhang, Shuo Yang, Zhongrui Wang, Yiwei Yang, Xianzhe Xu, Yibing Song, Weihua Chen, Fan Wang, Li Yuan

机构 * School of Electrical and Computer Engineering, Peking University(北京大学电子工程学院) Hupan Lab(鸿篇实验室) DAMO Academy, Alibaba Group(阿里云达摩院) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23227 2025-07-01 cs.CV 70%

High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation

Lunhao Duan, Shanshan Zhao, Xingxing Weng, Jing Zhang, Gui-Song Xia

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) JD Explore Academy(京东探索研究院) School of Computer Science and the School of Artificial Intelligence, Wuhan University(武汉大学计算机学院和人工智能学院)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by TPAMI. Code: https://github.com/LHDuan/WSegPC

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21101 2025-06-27 cs.CV 70%

OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography

Caoshuo Li, Zengmao Ding, Xiaobin Hu, Bang Li, Donghao Luo, AndyPian Wu, Chaoyang Wang, Chengjie Wang, Taisong Jin, SevenShu, Yunsheng Wu, Yongge Liu, Rongrong Ji

机构 * Xiamen University(厦门大学) Tencent(腾讯) Anyang Normal University(安阳师范学院)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14936 2025-06-19 cs.AI 70%

CALM: Contextual Analog Logic with Multimodality

Maxwell J. Jacobson, Corey J. Maley, Yexiang Xue

机构 * Purdue University(普渡大学)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13326 2025-06-17 cs.CV cs.HC 70%

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation

Bo Pan, Yixiao Fu, Ke Wang, Junyu Lu, Lunke Pan, Ziyang Qian, Yuhan Chen, Guoliang Wang, Yitao Zhou, Li Zheng, Yinghao Tang, Zhen Wen, Yuchen Wu, Junhua Lu, Biao Zhu, Minfeng Zhu, Bo Zhang, Wei Chen

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03126 2025-06-04 cs.CV 70%

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation

Lu Qiu, Yizhuo Li, Yuying Ge, Yixiao Ge, Ying Shan, Xihui Liu

机构 * The University of Hong Kong(香港大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Project released at: https://qiulu66.github.io/animeshooter/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00591 2025-06-03 eess.IV cs.CV 70%

MR2US-Pro: Prostate MR to Ultrasound Image Translation and Registration Based on Diffusion Models

Xudong Ma, Nantheera Anantrasirichai, Stefanos Bolomytis, Alin Achim

机构 * Visual Information Laboratory, University of Bristol(布里斯托大学视觉信息实验室) Southmead Hospital, North Bristol NHS Trust(北布里斯托NHS信托南梅德医院)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01025 2025-06-03 cs.CV 70%

Modality Translation and Registration of MR and Ultrasound Images Using Diffusion Models

Xudong Ma, Nantheera Anantrasirichai, Stefanos Bolomytis, Alin Achim

机构 * Visual Information Laboratory, University of Bristol(布里斯托大学视觉信息实验室) Southmead Hospital, North Bristol NHS Trust(北布里斯托国家健康服务信托南梅德医院)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21911 2025-05-29 cs.CV 70%

AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment

Yiheng Lin, Shifang Zhao, Ting Liu, Xiaochao Qu, Luoqi Liu, Yao Zhao, Yunchao Wei

机构 * Institute of Information Science, Beijing Jiaotong University(信息科学学院,北京交通大学) MT Lab, Meitu Inc.(美图实验室,美图公司)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19554 2025-05-27 cs.CV 70%

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation

Jiongchao Jin, Shengchu Zhao, Dajun Chen, Wei Jiang, Yong Li

机构 * Ant Group(蚂蚁集团) Agency for Science, Technology and Research(科技研究局)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17223 2025-05-26 cs.CV 70%

REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge

Siyang Song, Micol Spitale, Xiangyu Kong, Hengde Zhu, Cheng Luo, Cristina Palmero, German Barquero, Sergio Escalera, Michel Valstar, Mohamed Daoudi, Tobias Baur, Fabien Ringeval, Andrew Howes, Elisabeth Andre, Hatice Gunes

机构 * University of Exeter(埃克塞特大学) Politecnico di Milano(米兰理工大学) University of Leicester(莱斯特大学) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学) King’s College London(伦敦国王学院) Universitat de Barcelona(巴塞罗那大学) University of Nottingham(诺丁汉大学) IMT Nord Europe(IMT北欧分校) University of Augsburg(奥格斯堡大学) Université Grenoble Alpes(格勒诺布尔阿尔卑斯大学) University of Cambridge(剑桥大学)

专题命中 多模态生成 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22679 2025-05-26 cs.CV 70%

Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Weiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang, Junlin Li, Li Zhang, Jian Zhang

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10562 2025-05-16 cs.CV 70%

End-to-End Vision Tokenizer Tuning

Wenxuan Wang, Fan Zhang, Yufeng Cui, Haiwen Diao, Zhuoyan Luo, Huchuan Lu, Jing Liu, Xinlong Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Dalian University of Technology(大连理工大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06176 2025-05-12 cs.GR cs.CV cs.LG 70%

MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills

Niladri Shekhar Dutt, Duygu Ceylan, Niloy J. Mitra

机构 * University College London(伦敦大学学院) Adobe Research UK(Adobe英国研究)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted at SIGGRAPH 2025 [ACM Transactions on Graphics]; Project website: https://monetgpt.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01857 2025-05-06 cs.CV 70%

DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion

Haoteng Li, Zhao Yang, Zezhong Qian, Gongpeng Zhao, Yuqi Huang, Jun Yu, Huazheng Zhou, Longjun Liu

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家级人机混合增强智能实验室) National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) Xi’an Jiaotong University(西安交通大学) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages, 6 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16386 2025-04-28 cs.SE cs.AI 70%

Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach

Yuxuan Wan, Chaozheng Wang, Yi Dong, Wenxuan Wang, Shuqing Li, Yintong Huo, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted by FSE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13074 2025-04-22 cs.CV 70%

SkyReels-V2: Infinite-length Film Generative Model

Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Junchen Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengcheng Ma, Weiming Xiong, Wei Wang, Nuo Pang, Kang Kang, Zhiheng Xu, Yuzhe Jin, Yupeng Liang, Yubing Song, Peng Zhao, Boyuan Xu, Di Qiu, Debang Li, Zhengcong Fei, Yang Li, Yahui Zhou

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments 31 pages,10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10291 2025-04-18 cs.CL cs.AI cs.CV cs.LG cs.MM 70%

Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective

Xiangru Zhu, Penglei Sun, Yaoxian Song, Yanghua Xiao, Zhixu Li, Chengyu Wang, Jun Huang, Bei Yang, Xiaoxiao Xu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06666 2025-04-10 cs.CV 70%

Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

Ruotian Peng, Haiying He, Yake Wei, Yandong Wen, Di Hu

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08603 2025-04-10 cs.GR cs.CV 70%

Design2GarmentCode: Turning Design Concepts to Tangible Garments Through Program Synthesis

Feng Zhou, Ruiyang Liu, Chen Liu, Gaofeng He, Yong-Lu Li, Xiaogang Jin, Huamin Wang

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments The IEEE/CVF Conference on Computer Vision and Pattern Recognition (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06256 2025-04-09 cs.CV 70%

Transfer between Modalities with MetaQueries

Xichen Pan, Satya Narayan Shukla, Aashu Singh, Zhuokai Zhao, Shlok Kumar Mishra, Jialiang Wang, Zhiyang Xu, Jiuhai Chen, Kunpeng Li, Felix Juefei-Xu, Ji Hou, Saining Xie

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Project Page: https://xichenpan.com/metaquery

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01218 2025-04-03 cs.LG cs.CV 70%

Prompting Forgetting: Unlearning in GANs via Textual Guidance

Piyush Nagasubramaniam, Neeraj Karamchandani, Chen Wu, Sencun Zhu

专题命中 多模态生成 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23715 2025-04-01 cs.CV 70%

HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation

Kun Liu, Qi Liu, Xinchen Liu, Jie Li, Yongdong Zhang, Jiebo Luo, Xiaodong He, Wu Liu

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17536 2025-03-25 cs.CV 70%

DermDiff: Generative Diffusion Model for Mitigating Racial Biases in Dermatology Diagnosis

Nusrat Munia, Abdullah-Al-Zubaer Imran

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Paper presented at ADSMI@MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏