arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2506.20155 2025-06-26 cs.CV 79%

Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java, Silky Singh, Tarun Ram Menta, Surgan Jandial, Balaji Krishnamurthy

机构 * Indian Institute of Technology, Bombay(印度理工学院班加罗尔分校) Indian Institute of Technology, Roorkee(印度理工学院罗尔基分校) Microsoft Research(微软研究院) Stanford University(斯坦福大学) Adobe MDSR(Adobe MDSR实验室) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ECCV 2024 (AI4VA Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00659 2025-06-26 cs.RO cs.AI 79%

Multimodal Coherent Explanation Generation of Robot Failures

Pradip Pramanick, Silvia Rossi

机构 * Interdepartmental Center for Advances in Robotic Surgery - ICAROS, University of Naples Federico II(跨部门先进机器人手术中心 - ICAROS,那不勒斯费德里科二世大学) Department of Electrical Engineering and Information Technologies - DIETI, University of Naples Federico II(电气工程与信息科技系 - DIETI,那不勒斯费德里科二世大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11299 2025-06-25 cs.CV 79%

MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching

Yepeng Liu, Zhichao Sun, Baosheng Yu, Yitian Zhao, Bo Du, Yongchao Xu, Jun Cheng

机构 * National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心) Institute of Artificial Intelligence(人工智能研究院) School of Computer Science(计算机科学学院) Hubei Key Laboratory of Multimedia and Network Communication Engineering(湖北省多媒体与网络通信工程重点实验室) Lee Kong Chian School of Medicine(李光耀医学院) Nanyang Technological University(南洋理工大学) Ningbo Institute of Materials Technology and Engineering(宁波材料技术与工程研究所) Chinese Academy of Sciences(中国科学院) Institute for Infocomm Research (I 2 R)(信息与通信研究所以(I 2 R)) Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR))

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accept by IEEE TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17645 2025-06-24 cs.CV 79%

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning

Shih-Wen Liu, Hsuan-Yu Fan, Wei-Ta Chu, Fu-En Yang, Yu-Chiang Frank Wang

机构 * National Cheng Kung University, Taiwan(国立成功大学) NVIDIA Research(NVIDIA研究)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to MIDL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05298 2025-06-23 cs.CL 79%

Coreference as an indicator of context scope in multimodal narrative

Nikolai Ilinykh, Shalom Lappin, Asad Sayeed, Sharid Loáiciga

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments 19 pages, 4 tables. Accepted to GEM2 Workshop: Generation, Evaluation & Metrics at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15372 2025-06-19 cs.CL 79%

COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline Generation

Raghvendra Kumar, S. A. Mohammed Salman, Aryan Sahu, Tridib Nandi, Pragathi Y. P., Sriparna Saha, Jose G. Moreno

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳分校计算机科学与工程系) Department of Metallurgical and Materials Engineering, National Institute of Technology Tiruchirappalli, India(印度理工学院 Tiruchirappalli 金属与材料工程系) Department of Computer Science and Information Systems, BITS Pilani – Goa Campus, India(比斯·帕尼学院 Goa 分校计算机科学与信息系统系) Department of Computer Science and Engineering, Indian Institute of Information Technology Vadodara, India(印度信息科技学院瓦达拉分校计算机科学与工程系) Department of Computer Science and Engineering, B.M.S. College of Engineering, Bangalore, India(班加罗尔 B.M.S. 工程学院计算机科学与工程系) Université de Toulouse, IRIT UMR 5505 CNRS, France(图卢兹大学 IRIT UMR 5505 CNRS 实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments ACL 2025 MAINs

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13667 2025-06-17 eess.IV cs.CV 79%

MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model

Bi Yuda, Jia Sihan, Gao Yutong, Abrol Anees, Fu Zening, Calhoun Vince

机构 * Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS)(跨机构神经影像与数据科学转化研究中心(TReNDS))

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11178 2025-06-16 cs.CV cs.LG cs.NE 79%

BrainMAP: Multimodal Graph Learning For Efficient Brain Disease Localization

Nguyen Linh Dan Le, Jing Ren, Ciyuan Peng, Chengyao Xie, Bowen Li, Feng Xia

机构 * School of Computing Technologies, RMIT University(计算技术学院,拉筹纳斯大学) Institute of Innovation, Science and Sustainability, Federation University Australia(创新、科学与可持续性研究所,联邦大学澳大利亚)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09182 2025-06-16 eess.IV cs.CV 79%

seg2med: a bridge from artificial anatomy to multimodal medical images

Zeyu Yang, Zhilin Chen, Yipeng Sun, Anika Strittmatter, Anish Raj, Ahmad Allababidi, Johann S. Rink, Frank G. Zöllner

机构 * Computer Assisted Clinical Medicine, Medical Faculty Mannheim, Heidelberg University(计算机辅助临床医学,曼海姆医学院,海德堡大学) Pattern Recognition Lab, Friedrich-Alexander-University Erlangen-Nuremberg(模式识别实验室,埃尔兰根-纽伦堡弗里德里希-亚历山大大学) Department of Radiology and Nuclear Medicine, University Medical Center Mannheim(放射学与核医学系,曼海姆大学医学中心) Mannheim Institute for Intelligent Systems in Medicine, Medical Faculty Mannheim, Heidelberg University(曼海姆智能医学研究所,曼海姆医学院,海德堡大学) Optical Bioimaging Laboratory, Department of Biomedical Engineering, College of Design and Engineering, National University of Singapore(光学生物成像实验室,生物医学工程系,设计与工程学院,新加坡国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages, 10 figures Web demo available at https://huggingface.co/spaces/Zeyu0601/frankenstein

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06733 2025-06-12 cs.CV 79%

RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation

Ruoxuan Zhang, Jidong Gao, Bin Wen, Hongxia Xie, Chenming Zhang, Hong-Han Shuai, Wen-Huang Cheng

机构 * Jilin University(吉林大学) Guangdong University of Technology(广东工业大学) National Chiao Tung University(交通大学) National Taiwan University(国立台湾大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments This is an extended version of arXiv:2503.05228

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01343 2025-06-11 cs.AI 79%

BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model Editing

Dongliang Guo, Mengxuan Hu, Zihan Guan, Thomas Hartvigsen, Sheng Li

机构 * University of Virginia(弗吉尼亚大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Journal ref Proceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05153 2025-06-10 cs.CV 79%

Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment

Minh-Quan Le, Gaurav Mittal, Tianjian Meng, A S M Iftekhar, Vishwas Suryanarayanan, Barun Patra, Dimitris Samaras, Mei Chen

机构 * Microsoft(微软公司) Stony Brook University(史蒂文尼森布鲁克大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICLR 2025. Project page with code release: https://roar-ai.github.io/hummingbird

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16656 2025-06-09 cs.CV 79%

Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Peiyu Wang, Yichen Wei, Yi Peng, Xiaokun Wang, Weijie Qiu, Wei Shen, Tianyidan Xie, Jiangbo Pei, Jianhao Zhang, Yunzhuo Hao, Xuchen Song, Yang Liu, Yahui Zhou

机构 * Skywork AI Kunlun Inc.(Kunlun公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10798 2025-06-05 cs.CV 79%

MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

Jian Yang, Dacheng Yin, Yizhou Zhou, Fengyun Rao, Wei Zhai, Yang Cao, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发式智能感知与认知大学科学与技术研究院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03007 2025-06-04 cs.CV 79%

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models

Jiarui Wang, Huiyu Duan, Juntong Wang, Ziheng Jia, Woo Yi Yang, Xiaorong Zhu, Yu Zhao, Jiaying Qian, Yuke Xing, Guangtao Zhai, Xiongkuo Min

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02467 2025-06-04 eess.IV cs.CV 79%

Multi-modal brain MRI synthesis based on SwinUNETR

Haowen Pang, Weiyan Guo, Chuyang Ye

机构 * School of Integrated Circuits and Electronics, Beijing Institute of Technology, Beijing, China(集成电路与电子学院,北京理工大学,北京,中国)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19901 2025-06-04 cs.CV 79%

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Peng Liu, Xiaoming Ren, Fengkai Liu, Qingsong Xie, Quanlong Zheng, Yanhao Zhang, Haonan Lu, Yujiu Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02433 2025-06-04 cs.CV 79%

OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking

Zhongjian Wang, Peng Zhang, Jinwei Qi, Guangyuan Wang, Chaonan Ji, Sheng Xu, Bang Zhang, Liefeng Bo

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CV

Comments Project Page https://humanaigc.github.io/omnitalker

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01853 2025-06-03 cs.CV 79%

ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

Junliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie, Jun Zhu

机构 * Tsinghua University(清华大学) Peking University(北京大学) ShengShu(盛舒)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://github.com/JAMESYJL/ShapeLLM-Omni

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02640 2025-06-03 cs.MM 79%

RoSMM: A Robust and Secure Multi-Modal Watermarking Framework for Diffusion Models

ZhongLi Fang, Yu Xie, Ping Chen

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03860 2025-06-03 cs.CV 79%

MDMP: Multi-modal Diffusion for supervised Motion Predictions with uncertainty

Leo Bringer, Joey Wilson, Kira Barton, Maani Ghaffari

机构 * University of Michigan(密歇根大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2025 - HuMoGen. Minor revisions made based on reviewer feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09724 2025-06-03 cs.CV 79%

MFCLIP: Multi-modal Fine-grained CLIP for Generalizable Diffusion Face Forgery Detection

Yaning Zhang, Tianyi Wang, Zitong Yu, Zan Gao, Linlin Shen, Shengyong Chen

机构 * Faculty of Computer Science and Technology, Qilu University of Technology (Shandong Academy of Sciences)(计算机科学与技术学院,齐鲁工业大学(山东科学院)) School of Computing, National University of Singapore(computing 学院,新加坡国立大学) School of Computing and Information Technology, Great Bay University(computing 与信息学院,大湾大学) National Engineering Laboratory for Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室,深圳大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Information Forensics and Security 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24260 2025-06-02 cs.AI 79%

Generative AI for Urban Design: A Stepwise Approach Integrating Human Expertise with Multimodal Diffusion Models

Mingyi He, Yuebing Liang, Shenhao Wang, Yunhan Zheng, Qingyi Wang, Dingyi Zhuang, Li Tian, Jinhua Zhao

机构 * Department of Civil and Environmental Engineering, University of California, Berkeley(加州大学伯克利分校土木与环境工程系) The Singapore-MIT Alliance for Research and Technology(新加坡-麻省理工联盟研究技术中心) Department of Urban and Regional Planning, University of Florida(佛罗里达大学城市与区域规划系) Department of Civil and Environmental Engineering, Massachusetts Institute of Technology(麻省理工学院土木与环境工程系) Department of Urban Planning, Tsinghua University(清华大学城市规划系) Department of Urban Studies and Planning, Massachusetts Institute of Technology(麻省理工学院城市研究与规划系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22948 2025-05-30 cs.AI 79%

Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages

Michael Sun, Weize Yuan, Gang Liu, Wojciech Matusik, Jie Chen

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) MIT Chemistry(麻省理工学院化学系) MIT-IBM Watson AI Lab, IBM Research(麻省理工-IBM Watson人工智能实验室,IBM研究院) University of Notre Dame(诺埃伯大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09117 2025-05-30 cs.CV 79%

Cross-Modal Causal Intervention for Medical Report Generation

Weixing Chen, Yang Liu, Ce Wang, Jiarui Zhu, Guanbin Li, Cheng-Lin Liu, Liang Lin

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室) School of Science, Sun Yat-sen University(中山大学理学院) Hong Kong Polytechnic University(香港理工大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE TIP 2025, 16 pages, 11 figures, 7 tables

Journal ref IEEE Transactions on Image Processing 34 (2025) 2970-2985

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16990 2025-05-27 cs.CV 79%

Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding

Runpeng Yu, Xinyin Ma, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20115 2025-05-27 cs.SE cs.AI 79%

AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers

Zijie Lin, Yiqing Shen, Qilin Cai, He Sun, Jinrui Zhou, Mingjun Xiao

机构 * University of Science and Technology of China(中国科学技术大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18341 2025-05-27 cs.RO cs.AI 79%

CrashAgent: Crash Scenario Generation via Multi-modal Reasoning

Miao Li, Wenhao Ding, Haohong Lin, Yiqi Lyu, Yihang Yao, Yuyou Zhang, Ding Zhao

机构 * Carnegie Mellon University(卡内基梅隆大学) NVIDIA Research(NVIDIA研究) Northwestern University(西北大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16602 2025-05-23 cs.CV 79%

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation

Bohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing Lu

机构 * School of Computer Science, Peking University(北京大学计算机学院) Department of Automation, Tsinghua University(清华大学自动化系) BeingBeyond

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12606 2025-05-20 cs.CV 79%

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking

Shiyu Xuan, Zechao Li, Jinhui Tang

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏