arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

2503.16537 2025-03-24 cs.CL cs.CV 81%

Do Multimodal Large Language Models Understand Welding?

Grigorii Khvatskii, Yong Suk Lee, Corey Angst, Maria Gibbs, Robert Landers, Nitesh V. Chawla

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12927 2025-03-20 cs.CV cs.AI 81%

MMLNB: Multi-Modal Learning for Neuroblastoma Subtyping Classification Assisted with Textual Description Generation

Huangwei Chen, Yifei Chen, Zhenyu Yan, Mingyang Ding, Chenlei Li, Zhu Zhu, Feiwei Qin

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10324 2025-03-14 cs.CV cs.MM 81%

IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification

Yuhao Wang, Yongfeng Lv, Pingping Zhang, Huchuan Lu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments This work is accepted by CVPR2025. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08133 2025-03-12 cs.CV cs.AI 81%

MGHanD: Multi-modal Guidance for authentic Hand Diffusion

Taehyeon Eum, Jieun Choi, Tae-Kyun Kim

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15191 2025-03-12 cs.CV cs.LG cs.SD eess.AS 81%

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov, Alper Canberk, Kwot Sin Lee, Vicente Ordonez, Sergey Tulyakov

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments Project Page: snap-research.github.io/AVLink/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18459 2025-03-04 cs.CV cs.MM 81%

FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation

Yuki Imajuku, Yoko Yamakata, Kiyoharu Aizawa

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 15 pages, 5 figures. We found errors in the calculation of evaluation metrics, which were corrected in this version with $\color{blue}{\text{modifications highlighted in blue}}$. Please also see the Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00171 2025-03-04 cs.CV cs.AI 81%

PaliGemma-CXR: A Multi-task Multimodal Model for TB Chest X-ray Interpretation

Denis Musinguzi, Andrew Katumba, Sudi Murindanyi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20172 2025-02-28 cs.CV cs.CL 81%

Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think

Liang Chen, Shuai Bai, Wenhao Chai, Weichu Xie, Haozhe Zhao, Leon Vinci, Junyang Lin, Baobao Chang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 13 pages, 9 figures, codebase in https://github.com/chenllliang/DreamEngine

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11829 2025-02-18 cs.CL cs.AI cs.SE 81%

Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities

Hanbin Wang, Xiaoxuan Zhou, Zhipeng Xu, Keyuan Cheng, Yuxin Zuo, Kai Tian, Jingwei Song, Junting Lu, Wenhui Hu, Xueyang Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10536 2025-02-18 cs.CV cs.AI cs.LG 81%

PolyPath: Adapting a Large Multimodal Model for Multi-slide Pathology Report Generation

Faruk Ahmed, Lin Yang, Tiam Jaroensri, Andrew Sellergren, Yossi Matias, Avinatan Hassidim, Greg S. Corrado, Dale R. Webster, Shravya Shetty, Shruthi Prabhakara, Yun Liu, Daniel Golden, Ellery Wulczyn, David F. Steiner

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 main pages, 21 pages in total

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22108 2025-02-18 cs.CL cs.AI 81%

Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench

Zheyuan Liu, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng, Yongle Yuan, Meng Jiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments NAACL Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09843 2025-02-17 cs.AI cs.HC cs.MM 81%

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System

Karan Taneja, Ashok K. Goel

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments 5 pages, 3 figures, AAAI-MAKE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03163 2025-02-11 cs.CL cs.CV cs.CY 81%

Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Chenglei Si, Yanzhe Zhang, Ryan Li, Zhengyuan Yang, Ruibo Liu, Diyi Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments NAACL 2025; The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03429 2025-02-06 cs.CL cs.AI 81%

On Fairness of Unified Multimodal Large Language Model for Image Generation

Ming Liu, Hao Chen, Jindong Wang, Liwen Wang, Bhiksha Raj Ramakrishnan, Wensheng Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14170 2025-02-05 cs.IR cs.AI cs.MM 81%

Personalized Image Generation with Large Multimodal Models

Yiyan Xu, Wenjie Wang, Yang Zhang, Biao Tang, Peng Yan, Fuli Feng, Xiangnan He

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted for publication in WWW'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16698 2025-01-29 cs.CL cs.CV cs.RO 81%

3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow

Yueen Ma, Yuzheng Zhuang, Jianye Hao, Irwin King

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15393 2025-01-28 cs.AI cs.CL 81%

Diffusion-based Hierarchical Negative Sampling for Multimodal Knowledge Graph Completion

Guanglin Niu, Xiaowei Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments The version of a full paper accepted to DASFAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18216 2025-01-22 cs.CV cs.CL 81%

ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation

Mengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu, Zhongming Jiang, Xuehui Wang, Qinbin Li, Guangneng Hu, Shengchao Qin, Chi-Wing Fu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by the AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08187 2025-01-16 cs.CL cs.AI cs.CE cs.HC cs.LG q-bio.CB 81%

A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following

Yin Fang, Xinle Deng, Kangwei Liu, Ningyu Zhang, Jingyang Qian, Penghui Yang, Xiaohui Fan, Huajun Chen

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 37 pages; 13 figures; Code: https://github.com/zjunlp/Instructcell, Models: https://huggingface.co/zjunlp/Instructcell-chat, https://huggingface.co/zjunlp/InstructCell-instruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02523 2025-01-07 cs.CV cs.AI 81%

Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation

Dawei Dai, Mingming Jia, Yinxiu Zhou, Hang Xing, Chenghang Li

专题命中 多模态生成 :multimodal(title);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06550 2025-01-07 cs.CV cs.AI 81%

Multimodal Urban Areas of Interest Generation via Remote Sensing Imagery and Geographical Prior

Chuanji Shi, Yingying Zhang, Jiaotuan Wang, Xin Guo, Qiqi Zhu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 9 figures

Journal ref International Journal of Applied Earth Observation and Geoinformation, 136(2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16771 2024-12-31 cs.CV cs.AI 81%

ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving

Jiehui Huang, Xiao Dong, Wenhui Song, Zheng Chong, Zhenchao Tang, Jun Zhou, Yuhao Cheng, Long Chen, Hanhui Li, Yiqiang Yan, Shengcai Liao, Xiaodan Liang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Project page: https://ssugarwh.github.io/consistentid.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19009 2024-12-30 cs.CV cs.MM 81%

FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing

Wanglong Lu, Jikai Wang, Xiaogang Jin, Xianta Jiang, Hanli Zhao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Published at IEEE Transactions on Visualization and Computer Graphics; 21 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18775 2024-12-30 cs.CV cs.AI cs.LG 81%

ObitoNet: Multimodal High-Resolution Point Cloud Reconstruction

Apoorv Thapliyal, Vinay Lanka, Swathi Baskaran

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12206 2024-12-18 cs.MM cs.CR cs.CV 81%

Provably Secure Robust Image Steganography via Cross-Modal Error Correction

Yuang Qi, Kejiang Chen, Na Zhao, Zijin Yang, Weiming Zhang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments 7 pages. Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12165 2024-12-18 cs.CV cs.AI cs.CY cs.LG 81%

Multimodal Approaches to Fair Image Classification: An Ethical Perspective

Javon Hickmon

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Bachelor's thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10605 2024-12-17 cs.CV cs.AI 81%

MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration

Yanbo Ding, Shaobin Zhuang, Kunchang Li, Zhengrong Yue, Yu Qiao, Yali Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09008 2024-12-13 cs.CV cs.HC cs.MM 81%

MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments

Yuqi Tong, Yue Qiu, Ruiyang Li, Shi Qiu, Pheng-Ann Heng

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments IEEE AIxVR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08635 2024-12-12 cs.CL cs.CV cs.LG 81%

Multimodal Latent Language Modeling with Next-Token Diffusion

Yutao Sun, Hangbo Bao, Wenhui Wang, Zhiliang Peng, Li Dong, Shaohan Huang, Jianyong Wang, Furu Wei

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04707 2024-12-09 cs.AI cs.CE cs.CV cs.HC 81%

Parametric-ControlNet: Multimodal Control in Foundation Models for Precise Engineering Design Synthesis

Rui Zhou, Yanxia Zhang, Chenyang Yuan, Frank Permenter, Nikos Arechiga, Matt Klenk, Faez Ahmed

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏