arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

2507.14298 2025-07-22 cs.CL cs.AI cs.CV 78%

In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu, Alexander Jacobson, Lu Yuan, Leonid Sigal

机构 * UBC(不列颠哥伦比亚大学) Microsoft(微软) Vector Institute for AI(人工智能向量研究所) CIFAR AI Chair(卡尔·弗雷德里克人工智能主席)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2407.14506

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13737 2025-07-21 cs.AI cs.CL cs.HC cs.MM 78%

DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs

Ye Tian, Xiaoyuan Ren, Zihao Wang, Onat Gungor, Xiaofan Yu, Tajana Rosing

机构 * University of California San Diego, Computer Science and Engineering Department(加州大学圣地亚哥分校计算机科学与工程系)

专题命中 多模态生成 :multi-modal(title);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10338 2025-07-15 cs.SE cs.AR cs.LO 78%

AssertCoder: LLM-Based Assertion Generation via Multimodal Specification Extraction

Enyuan Tian, Yiwei Ci, Qiusong Yang, Yufeng Li, Zhichao Lyu

专题命中 多模态生成 :multimodal(title,abstract)

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05718 2025-07-09 cs.IT math.IT 78%

Cooperative Mapping, Localization, and Beam Management via Multi-Modal SLAM in ISAC Systems

Hang Que, Jie Yang, Tao Du, Shuqiang Xia, Chao-Kai Wen, Shi Jin

专题命中 多模态生成 :multi-modal(title,abstract)

Comments Accepted by IEEE Transactions on Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17428 2025-07-08 eess.IV 78%

Image Generation with Supervised Selection Based on Multimodal Features for Semantic Communications

Chengyang Liang, Dong Li

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20399 2025-06-26 cs.RO 78%

Multimodal Behaviour Trees for Robotic Laboratory Task Automation

Hatem Fakhruldeen, Arvind Raveendran Nambiar, Satheeshkumar Veeramani, Bonilkumar Vijaykumar Tailor, Hadi Beyzaee Juneghani, Gabriella Pizzuto, Andrew Ian Cooper

机构 * University of Liverpool(利兹大学)

专题命中 多模态生成 :multimodal(title,abstract)

Comments 7 pages, 5 figures, accepted and presented in ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14881 2025-06-26 cs.SE 78%

Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing

Zhi Tu, Liangkun Niu, Wei Fan, Tianyi Zhang

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 24 pages, 6 figures, Accepted to FSE 2025

Journal ref Proceedings of the ACM on Software Engineering, Volume 2, Issue FSE, Article No. FSE078 (July 2025), pp. 1733--1756

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18779 2025-06-24 cs.RO 78%

DefFusionNet: Learning Multimodal Goal Shapes for Deformable Object Manipulation via a Diffusion-based Probabilistic Model

Bao Thach, Siyeon Kim, Britton Jordan, Mohanraj Shanthi, Tanner Watts, Shing-Hei Ho, James M. Ferguson, Tucker Hermans, Alan Kuntz

机构 * Robotics Center and Kahlert School of Computing, University of Utah(大学计算机学院和机器人中心,犹他大学) NVIDIA Corporation(英伟达公司)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15873 2025-06-23 cs.HC 78%

DeckFlow: Iterative Specification on a Multimodal Generative Canvas

Gregory Croisdale, Emily Huang, John Joon Young Chung, Anhong Guo, Xu Wang, Austin Z. Henley, Cyrus Omar

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10311 2025-06-13 cs.DM 78%

The Freight Multimodal Transport Problem with Buses and Drones: An Integrated Approach for Last-Mile Delivery

E Su, Hu Qin, Jiliu Li, Rui Zhang

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07647 2025-06-10 eess.SP 78%

Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration

Xiang Cheng, Boxun Liu, Xuanyu Liu, Ensong Liu, Ziwei Huang

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14681 2025-06-02 cs.CR 78%

TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model

Ji Guo, Peihong Chen, Wenbo Jiang, Xiaolei Wen, Jiaming He, Jiachen Li, Guoming Lu, Aiguo Chen, Hongwei Li

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23693 2025-05-30 cs.CV cs.AI cs.CL 78%

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

Tingyu Song, Tongyan Hu, Guo Gan, Yilun Zhao

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学) Yale University(耶鲁大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22584 2025-05-29 cs.IR 78%

DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers

Navve Wasserman, Oliver Heinimann, Yuval Golbari, Tal Zimbalist, Eli Schwartz, Michal Irani

专题命中 多模态生成 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15973 2025-05-23 cs.HC 78%

An Exploratory Study on Multi-modal Generative AI in AR Storytelling

Hyungjun Doh, Jingyu Shi, Rahul Jain, Heesoo Kim, Karthik Ramani

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12293 2025-05-16 cs.SE cs.LG 78%

Unified Modeling Language Code Generation from Diagram Images Using Multimodal Large Language Models

Averi Bates, Ryan Vavricka, Shane Carleton, Ruosi Shao, Chongle Pan

专题命中 多模态生成 :multimodal(title,abstract)

Comments Published in the Journal of Machine Learning with Applications, Author Contributions: Averi Bates: Methodology, Development, Analysis, Data Curation, Drafting, Review. Ryan Vavricka: Data Curation, Development, Review. Shane Carleton: Supervision, Funding. Ruosi Shao: Review. Chongle Pan: Supervision, Review

Journal ref Mach. Learn. Appl. 20 (2025) 100660

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07825 2025-05-14 stat.ML cs.LG math.PR 78%

Diffusion-based supervised learning of generative models for efficient sampling of multimodal distributions

Hoang Tran, Zezhong Zhang, Feng Bao, Dan Lu, Guannan Zhang

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00215 2025-05-13 cs.LG 78%

Characterizing and Efficiently Accelerating Multimodal Generation Model Inference

Yejin Lee, Anna Sun, Basil Hosmer, Bilge Acun, Can Balioglu, Changhan Wang, Charles David Hernandez, Christian Puhrsch, Daniel Haziza, Driss Guessous, Francisco Massa, Jacob Kahn, Jeffrey Wan, Jeremy Reizenstein, Jiaqi Zhai, Joe Isaacson, Joel Schlosser, Juan Pino, Kaushik Ram Sadagopan, Leonid Shamis, Linjian Ma, Min-Jae Hwang, Mingda Chen, Mostafa Elhoushi, Pedro Rodriguez, Ram Pasunuru, Scott Yih, Sravya Popuri, Xing Liu, Carole-Jean Wu

机构 * Meta

专题命中 多模态生成 :multimodal(title);multi-modal(abstract)

Comments 13 pages including references. 8 Figures. Under review to HPCA 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02381 2025-05-06 eess.SP 78%

Multimodal Deep Learning-Empowered Beam Prediction in Future THz ISAC Systems

Kai Zhang, Wentao Yu, Hengtao He, Shenghui Song, Jun Zhang, Khaled B. Letaief

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18729 2025-04-29 cs.LG 78%

Multimodal graph representation learning for website generation based on visual sketch

Tung D. Vu, Chung Hoang, Truong-Son Hy

机构 * College of Engineering and Computer Science(工程与计算机科学学院) VinUniversity Hanoi(河内 Vin 大学) Department of Computer Science(计算机科学系) Hanoi University of Science and Technology(河内科学技术大学) The University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19449 2025-04-15 stat.ML cs.LG stat.CO 78%

Learned Reference-based Diffusion Sampling for multi-modal distributions

Maxence Noble, Louis Grenioux, Marylou Gabrié, Alain Oliviero Durmus

专题命中 多模态生成 :multi-modal(title,abstract)

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06543 2025-04-10 cs.IR 78%

DiffusionCom: Structure-Aware Multimodal Diffusion Model for Multimodal Knowledge Graph Completion

Wei Huang, Meiyu Liang, Peining Li, Xu Hou, Yawen Li, Junping Du, Zhe Xue, Zeli Guan

专题命中 多模态生成 :multimodal(title,abstract)

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04209 2025-04-01 cs.RO 78%

CALMM-Drive: Confidence-Aware Autonomous Driving with Large Multimodal Model

Ruoyu Yao, Yubin Wang, Haichao Liu, Rui Yang, Zengqi Peng, Lei Zhu, Jun Ma

专题命中 多模态生成 :multimodal(title,abstract)

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20913 2025-03-28 cs.CE cs.LG 78%

TransDiffSBDD: Causality-Aware Multi-Modal Structure-Based Drug Design

Xiuyuan Hu, Guoqing Liu, Can Chen, Yang Zhao, Hao Zhang, Xue Liu

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07975 2025-03-25 cs.CV cs.AI cs.CL 78%

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Haowei Zhang, Xingkai yu, Liang Zhao, Yisong Wang, Jiaying Liu, Chong Ruan

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08937 2025-03-13 eess.SP cs.LG 78%

Beam Selection in ISAC using Contextual Bandit with Multi-modal Transformer and Transfer Learning

Mohammad Farzanullah, Han Zhang, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 6 pages, 4 figures, 2 tables, IEEE International Conference on Communications 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06119 2025-03-11 cs.LG 78%

Unlocking Pretrained LLMs for Motion-Related Multimodal Generation: A Fine-Tuning Approach to Unify Diffusion and Next-Token Prediction

Shinichi Tanaka, Zhao Wang, Yoichi Kato, Jun Ohya

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03752 2025-03-07 cs.CY 78%

Multimodal Generative AI and Foundation Models for Behavioural Health in Online Gambling

Konrad Samsel, Mohammad Noaeen, Neil Seeman, Karim Keshavjee, Li-Jia Li, Zahra Shakeri

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11734 2025-03-04 q-bio.QM cs.LG q-bio.GN 78%

Multi-Modal and Multi-Attribute Generation of Single Cells with CFGen

Alessandro Palma, Till Richter, Hanyi Zhang, Manuel Lubetzki, Alexander Tong, Andrea Dittadi, Fabian Theis

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 41 pages, 22 figures

Journal ref The Thirteenth International Conference on Learning Representations (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11399 2025-02-19 cs.HC 78%

FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization

Yuki Tatsukawa, I-Chao Shen, Mustafa Doga Dogan, Anran Qi, Yuki Koyama, Ariel Shamir, Takeo Igarashi

专题命中 多模态生成 :multimodal(title,abstract)

Comments 14 pages

Journal ref CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏