arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

2606.05950 2026-06-29 cs.AI 版本更新 77%

Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing

Edit-R2:面向多轮图像编辑的上下文感知强化学习

Yuxiao Ye, Haoran He, Fangyuan Kong, Xintao Wang, Pengfei Wan, Kun Gai, Ling Pan

机构 * Hong Kong University of Science and Technology(香港理工大学) Kuaishou Technology(快手科技)

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 提出Edit-R2框架,通过重构会话意图和联合优化推理与生成的强化学习,解决多轮图像编辑中的长上下文稀释和状态污染问题,并在MICE-Bench基准上取得领先性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09303 2026-06-09 cs.CV 新提交 77%

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning

再思考:通过候选发现与比较推理进行分割

Xinyan Gao, Haoran Hao, Xiangyu Yue

机构 * The Chinese University of Hong Kong(香港中文大学) Nanjing University(南京大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出两阶段框架Rea2Seg,先基于注意力图生成候选掩码,再用多模态大语言模型推理评分,将分割转化为候选发现与判别选择,并引入新基准ReasonSeg-SGDR全面评估感知、定位与推理能力。

Comments Project page: https://snowball521.github.io/Rea2Seg-Project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12983 2026-06-05 cs.CL 77%

ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation

ChartAttack: 测试大型语言模型在图表生成中对恶意提示的脆弱性

Jesus-German Ortiz-Barajas, Jonathan Tonglet, Vivek Gupta, Iryna Gurevych

机构 * INSAIT, Sofia University "St. Kliment Ohridski"(INSAIT索菲亚大学"圣克莱门特·欧赫里迪斯基") Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(无处不在知识处理实验室(UKP实验室)、计算机科学系、图腾达姆斯塔特大学和应用网络安全国家研究中心ATHENE) Arizona State University(亚利桑那州立大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CL

AI总结 本文提出ChartAttack框架,用于评估多模态大语言模型在生成误导性图表方面的能力,通过注入误导性元素来诱导错误解释,并引入AttackViz数据集来评估和改进模型的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03236 2026-06-03 cs.AI 77%

Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

先感知后推理:一种用于高效可靠主动移动代理的预推理感知框架

Zhijie Ding, Weinan Hong, Zicheng Zhu, Lei Li, Dezhi Kong, Hao Wang, Peng Zhou, Xuchu Jiang, Jiaming Xu

机构 * HyperAI Team, Xiaomi Corporation(HyperAI团队,小米公司) Zhongnan University of Economics and Law(中南财经政法大学) Jilin University(吉林大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 提出预推理感知框架(PRPF),通过轻量级多模态主动感知器(MPP)进行干预门控和上下文压缩,仅在需要时激活主动代理推理器(PAR),以解决主动移动代理中干预时机与方式决策的目标错位和冗余推理问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04114 2026-05-26 cs.CV 77%

Any2Any: Unified Arbitrary Modality Translation for Remote Sensing

Any2Any: 统一任意模态遥感翻译

Haoyang Chen, Jing Zhang, Hebaixu Wang, Shiqin Wang, Pohsun Huang, Jiayuan Li, Haonan Guo, Di Wang, Zheng Wang, Bo Du

机构 * National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University(多媒体软件国家工程研究中心、人工智能研究院、计算机科学学院、武汉大学) Hubei Key Laboratory of Multimedia(湖北省多媒体重点实验室) Zhongguancun Academy, Beijing, China. 100094(中关村学院,北京,中国。100094) School of Electronic Information, Wuhan University, Wuhan, China(电子信息学院,武汉大学,武汉,中国) School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan, China(测绘、制图与遥感信息工程国家重点实验室,武汉大学,武汉,中国)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);any-to-any(abstract);分类 cs.CV

AI总结 提出统一潜扩散框架Any2Any,通过共享潜空间和轻量残差适配器实现任意模态间的高效翻译,并在新数据集RST-1M上验证了其优于成对方法且具备零样本泛化能力。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16745 2026-05-19 cs.CV 77%

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers

EVA01: 通过混合变换器实现统一的原生3D理解和生成

Zongyuan Yang, Mingjing Yi, Wanli Ma, Chenzhuo Fan, Bocheng Li, Baolin Liu, Yuke Lou, Yingde Song, Yongping Xiong, Zhengdong Guo, Shimu Wang

机构 * SeeleAI Team(SeeleAI团队)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出EVA01框架,通过混合变换器架构扩展多模态大语言模型的模态边界,实现原生的3D网格理解和生成以及上下文感知编辑,提升文本到3D生成的保真度和多轮几何编辑能力。

Comments 28 pages, 10 figures, 6 tables. Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03239 2026-05-12 cs.CV 77%

COP-GEN: Latent Diffusion Transformer for Copernicus Earth Observation Data

COP-GEN:用于Copernicus地球观测数据的潜在扩散变换器

Miguel Espinosa, Eva Gmelich Meijling, Valerio Marsocci, Elliot J. Crowley, Mikolaj Czerkawski

机构 * School of Engineering University of Edinburgh(工程学院爱丁堡大学) European Space Agency (ESA)(欧洲航天局) Asterisk Labs(Asterisk实验室)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);any-to-any(abstract);分类 cs.CV

AI总结 COP-GEN通过建模异质地球观测模态的联合分布,实现了多模态潜在扩散变换器,支持灵活的任意到任意条件生成,包括无需重新训练的零样本模态翻译。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21164 2026-05-12 cs.AI 77%

Concise Geometric Description as a Bridge: Unleashing the Potential of LLM for Plane Geometry Problem Solving

简洁几何描述作为桥梁:释放LLM在平面几何问题求解中的潜力

Jingyun Wang, Dian Li, Xiaohan Wang, Gang Liu, Jiahong Yan, Guoliang Kang

机构 * Beihang University(北航大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出通过训练多模态大语言模型解释器生成几何描述,再利用通用LLM进行推理,以解决平面几何问题。通过设计CDL匹配奖励机制,提升生成几何描述的效率,并在新构建的Formalgeo7k-Rec-CoT数据集上验证了方法的有效性。

Comments CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20054 2026-03-17 cs.CV 77%

Marmot: Object-Level Self-Correction via Multi-Agent Reasoning

Marmot:通过多智能体推理实现对象级自校正

Jiayang Sun, Hongbo Wang, Jie Cao, Huaibo Huang, Ran He

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

AI总结 Marmot通过多智能体推理实现对象级自校正,提升图像文本对齐的准确性,解决多对象场景中计数、属性和空间关系的校正问题。

Journal ref Machine Intelligence Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02771 2026-01-07 cs.CV 77%

AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs

AbductiveMLLM: 提升多模态大语言模型中的视觉归纳推理

Boyu Chang, Qi Wang, Xi Guo, Zhixiong Nan, Yazhou Yao, Tianfei Zhou

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 AbductiveMLLM通过结合REASONER和IMAGINER组件,提升多模态大语言模型在视觉归纳推理中的性能。

Comments Accepted by AAAI 2026 as Oral. Code:https://github.com/ChangPtR/AbdMLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18131 2025-11-25 cs.CV 77%

Video4Edit: Viewing Image Editing as a Degenerate Temporal Process

Video4Edit: 将图像编辑视为一种退化的时间过程

Xiaofan Li, Yanpeng Sun, Chenming Wu, Fan Duan, YuAn Wang, Weihao Bo, Yumeng Zhang, Dingkang Liang

机构 * Baidu Inc.(百度公司)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 Video4Edit通过将图像编辑视为退化的时间过程,利用视频预训练的单帧演化先验,实现高效的数据微调,从而在性能上与主流模型相当,但仅需1%的监督数据。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14027 2025-11-19 cs.CL 77%

HiEAG: Evidence-Augmented Generation for Out-of-Context Misinformation Detection

Junjie Wu, Yumeng Fu, Nan Yu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23760 2025-09-30 cs.CV 77%

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

Xinyang Song, Libin Wang, Weining Wang, Shaozhen Liu, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun

机构 * Ant Group(蚂蚁集团)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18898 2025-06-24 cs.CV cs.AI cs.CL cs.MM 77%

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Jiaming Han, Hao Chen, Yang Zhao, Hanyu Wang, Qi Zhao, Ziyan Yang, Hao He, Xiangyu Yue, Lu Jiang

机构 * CUHK MMLab(香港中文大学MML实验室) ByteDance Seed(字节跳动种子)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project page: https://tar.csuhan.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11036 2025-06-16 cs.LG cs.MM 77%

Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification

Yang Qin, Chao Chen, Zhihang Fu, Dezhong Peng, Xi Peng, Peng Hu

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) National Key Laboratory of Fundamental Algorithms and Models for Engineering Simulation, Sichuan University(四川大学国家工程仿真基础算法与模型重点实验室) Sichuan National Innovation New Vision UHD Video Technology Co., Ltd(四川国家创新新视觉超高清视频技术有限公司) Tianfu Jincheng Laboratory(天府锦城实验室)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02980 2025-06-02 cs.CV 77%

GPT4Point: A Unified Framework for Point-Language Understanding and Generation

Zhangyang Qi, Ye Fang, Zeyi Sun, Xiaoyang Wu, Tong Wu, Jiaqi Wang, Dahua Lin, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06673 2024-12-10 cs.CV 77%

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Chunwei Wang, Guansong Lu, Junwei Yang, Runhui Huang, Jianhua Han, Lu Hou, Wei Zhang, Hang Xu

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06429 2023-03-23 cs.CV 77%

Linguistic Query-Guided Mask Generation for Referring Image Segmentation

Zhichao Wei, Xiaohao Chen, Mingqiang Chen, Siyu Zhu

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27595 2026-05-28 cs.CV cs.AI 76%

Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks

多模态大语言模型在农业图像解释与生成任务中的幻觉行为

Partho Ghose, Al Bashir, Prem Raj, Azlan Zahid

机构 * Texas A&M University System(德克萨斯大学系统)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本研究系统评估了多模态大语言模型在农业图像解释(图像到文本)和生成(文本到图像)任务中的幻觉行为,发现模型存在生物不一致、上下文不准确和农学不合理等错误模式,并通过少样本提示等方法分析了幻觉的残留影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13637 2026-04-21 cs.CV cs.MM 76%

Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation

探索互惠的跨模态注意力以实现情境感知的人体可供性生成

Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal, Michael Blumenstein

机构 * University of Technology Sydney(技术大学悉尼大学) Indian Institute of Technology, Kharagpur(印度理工学院哈里科特) Indian Statistical Institute, Kolkata(印度统计研究所科契)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.MM

AI总结 本文提出一种互惠的跨模态注意力机制,通过不同模态的空间特征图互相关注,以提升2D场景中人体可供性预测的准确性与效率。

Comments Accepted in The IEEE Transactions on Artificial Intelligence (TAI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14477 2026-04-14 cs.CV cs.AI eess.IV 76%

XD-MAP: Cross-Modal Domain Adaptation via Semantic Parametric Maps for Scalable Training Data Generation

XD-MAP:通过语义参数地图实现跨模态领域适应以实现可扩展的训练数据生成

Frank Bieder, Hendrik Königshof, Haohao Hu, Fabian Immel, Yinzhe Shen, Jan-Hendrik Pauls, Christoph Stiller

机构 * KIT Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) FZI Research Center for Information Technology(FZI信息技术研究中心)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出XD-MAP,通过语义参数地图将图像数据集中的传感器特定知识转移到LiDAR领域,无需手动标注即可提升多模态领域适应性能。

Comments 10 pages, 7 figures, 3 tables, accepted at CVPRW

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22274 2026-02-11 cs.CV cs.CL 76%

Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication

常见物体脱离上下文(COOCo):研究多模态上下文和语义场景违规在指代通信中的作用

Filippo Merlo, Ece Takmaz, Wenkai Chen, Albert Gatt

机构 * Utrecht University(乌特雷赫大学) University of Trento(特伦托大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

AI总结 COOCo研究了VLMs在指代生成中如何利用场景上下文,发现模型根据语义相关性和噪声水平动态平衡局部与上下文信息。

Comments Accepted to TACL (pre-MIT Press publication version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15392 2026-01-23 cs.AI cs.CV cs.LG 76%

GeMM-GAN: A Multimodal Generative Model Conditioned on Histopathology Images and Clinical Descriptions for Gene Expression Profile Generation

GeMM-GAN: 一种基于组织病理图像和临床描述的多模态生成模型,用于基因表达谱生成

Francesca Pia Panaccione, Carlo Sgaravatti, Pietro Pinoli

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

AI总结 GeMM-GAN通过结合图像和文本信息生成逼真的基因表达谱,提升疾病预测准确性。

Comments 12 pages, 2 figures. Published at Image Analysis and Processing - ICIAP 2025 Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13729 2025-12-17 cs.LG cs.AI cs.CV 76%

Composite Classifier-Free Guidance for Multi-Modal Conditioning in Wind Dynamics Super-Resolution

多模态条件下的风动力超分辨率复合分类器引导方法

Jacob Schnell, Aditya Makkar, Gunadi Gani, Aniket Srinivasan Ashok, Darren Lo, Mike Optis, Alexander Wong, Yuhao Chen

机构 * University of Waterloo(滑铁卢大学) Veer Renewables

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出复合分类器引导方法,用于多模态条件下的风动力超分辨率重建,实现高保真度与低成本的风数据生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13689 2025-11-19 cs.CL cs.CV 76%

Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation

Sofia Jamil, Kotla Sai Charan, Sriparna Saha, Koustava Goswami, Joseph K J

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02046 2025-11-05 cs.CV cs.AI 76%

Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis

Soham Joshi, Shwet Kamal Mishra, Viswanath Gopalakrishnan

机构 * International Institute of Information Technology Bangalore(国际信息科技学院班加罗尔)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

Comments First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21562 2025-08-05 cs.CL cs.AI cs.AR 76%

FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction

Jun Yin, Pengyu Zeng, Jing Zhong, Peilin Li, Miao Zhang, Ran Luo, Shuai Lu

专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19149 2025-06-12 cs.CV cs.AI cs.CR cs.LG 76%

Multimodal Pragmatic Jailbreak on Text-to-image Models

Tong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang, Shuo Chen, Philip Torr, Vera Demberg, Volker Tresp, Jindong Gu

机构 * LMU Munich, Germany(慕尼黑大学) Munich Center for Machine Learning, Germany(慕尼黑机器学习中心) Saarland University, Germany(萨尔兰大学) Max Planck Institute for Informatics, Germany(马克斯·普朗克信息研究所) Cornell University, USA(康奈尔大学) University of Oxford, UK(牛津大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08838 2025-05-20 eess.IV cs.AI cs.CV 76%

Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts

Peixuan Ge, Tongkun Su, Faqin Lv, Baoliang Zhao, Peng Zhang, Chi Hong Wong, Liang Yao, Yu Sun, Zenan Wang, Pak Kin Wong, Ying Hu

机构 * Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) University of Macau(澳门大学) Chinese PLA General Hospital(中国人民解放军总医院) Macau University of Science and Technology(澳门科学技术大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10453 2025-05-06 cs.CV cs.GR cs.MM 76%

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Liu He, Yizhi Song, Hejun Huang, Pinxin Liu, Yunlong Tang, Daniel Aliaga, Xin Zhou

机构 * Purdue University(普渡大学) Baidu USA(百度美国公司) University of Rochester(罗切斯特大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.MM

Comments Accepted by CVPR 2025 AI4CC Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏