arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4651 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2606.09360 2026-06-09 cs.CV 新提交 74%

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification

ExDet: 基于跨模态外推与校正的开放域开放词汇检测

Yupeng Zhang, Yuzhong Feng, Ruize Han, Zhiwei Chen, Wei Feng, Liang Wan

机构 * College of Intelligence and Computing, Tianjin University(天津大学智能与计算学部) Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology(深圳理工大学计算机科学与人工智能学院) School of Artificial Intelligence, Nanchang University(南昌大学人工智能学院)

专题命中 图文多模态 :cross-modal(title);分类 cs.CV

AI总结 提出ExDet框架,通过文本引导外推(TGE)和检测器兼容校正(DCR)模块,无需额外训练即可增强开放域开放词汇检测的跨类别和跨域泛化能力,在多个基准上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00400 2026-05-13 cs.AI 74%

KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA

KEPO:基于知识的偏好优化用于多模态推理及其在医学VQA中的应用

Fan Yang, Rui Meng, Trudi Di Qi, Ali Ezzati, Yuxin Wen

机构 * Chapman University(查普曼大学) Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室) University of California, Irvine(加州大学伊文斯分校)

专题命中 图文多模态 :multimodal(title);分类 cs.AI

AI总结 KEPO通过质量门控的在线蒸馏和知识增强的探索策略,提升多模态推理任务的训练稳定性与推理一致性,优于强化学习和在线蒸馏基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04564 2026-04-07 cs.RO cs.CV 74%

Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs

基于视觉提示的越野制图推理使用多模态大语言模型

Abdelmoamen Nasser, Yousef Baba'a, Murad Mebrahtu, Nadya Abdel Madjid, Jorge Dias, Majid Khonji

机构 * Khalifa University(哈利法大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出一种零样本方法,利用SAM2进行环境分割和多模态大语言模型(VLM)进行可行驶区域推理,通过整合分割图像和原始图像实现越野制图推理,无需专门地形模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27898 2026-03-31 cs.CV 74%

SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation

SAGE:面向sink的 grounded 解码用于多模态幻觉缓解

Tripti Shukla, Zsolt Kira

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 SAGE通过动态调节自注意力机制,实时监控 grounding 可靠性,有效缓解多模态幻觉问题,实验显示在MSCOCO和AMBER基准上取得显著提升。

Comments 25 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21010 2026-03-24 cs.CV 74%

SkinCLIP-VL: Consistency-Aware Vision-Language Learning for Multimodal Skin Cancer Diagnosis

SkinCLIP-VL: 一种面向多模态皮肤癌诊断的一致性感知视觉-语言学习框架

Zhixiang Lu, Shijie Xu, Kaicheng Yan, Xuyue Cai, Chong Zhang, Yulong Li, Angelos Stefanidis, Anh Nguyen, Jionglong Su

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出SkinCLIP-VL,通过冻结感知与自适应推理范式,结合CLIP编码器与轻量量化Qwen2.5-VL,引入一致性聚焦对齐损失,实现高效皮肤癌诊断,优于13B参数基线模型,参数更少且临床信任度更高。

Comments Accepted by 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06263 2026-03-10 cs.CV 74%

iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models

iLLaVA: 图像的价值少于三分之一的输入标记在大型多模态模型中

Lianyu Hu, Liqing Gao, Fanhua Shang, Liang Wan, Wei Feng

机构 * School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院) School of Computer Science and Technology, Tiangong University(天津工大学计算机科学与技术学院) Key Research Center for Surface Monitoring and Analysis of Relics, State Administration of Cultural Heritage(文物表面监测与分析关键研究中心,国家文物局)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 iLLaVA通过联合优化图像编码器和LLM,减少冗余标记并提升吞吐量,实现更高效的多模态模型性能。

Comments Accepted by ICLR2026,code is released at https://github.com/hulianyuyy/iLLaVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00251 2026-03-05 cs.SE cs.AI cs.SY eess.SY 74%

GENAI WORKBENCH: AI-Assisted Analysis and Synthesis of Engineering Systems from Multimodal Engineering Data

GenAI Workbench: 人工智能辅助从多模态工程数据中分析和合成工程系统

H. Sinan Bank, Daniel R. Herber

机构 * Colorado State University(科罗拉多州立大学)

专题命中 图文多模态 :multimodal(title);分类 cs.AI

AI总结 GenAI Workbench通过整合AI技术,实现从多模态工程数据中辅助分析和合成工程系统,促进更集成和数据驱动的工程设计方法。

Comments 7 pages, 3 figures, accepted to be presented at IISE Annual Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22120 2026-02-06 cs.CV 74%

See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning

看得少,看得对:双向感知塑造用于多模态推理

Shuoshuo Zhang, Yizhen Zhang, Jingjing Fu, Lei Song, Jiang Bian, Yujiu Yang, Rui Wang

机构 * Microsoft Research Asia(微软亚洲研究院) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出BiPS,通过双向感知塑造提升多模态推理性能,有效增强细粒度视觉依赖并提升跨领域泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19267 2026-01-28 cs.CL 74%

DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models

DiaDem: 促进音频视频字幕中的对话描述以提升多模态大语言模型

Xinlong Chen, Weihong Lin, Jingyun Hua, Linli Yao, Yue Ding, Bozhou Li, Bohan Zeng, Yang Shi, Qiang Liu, Yuanxing Zhang, Pengfei Wan, Liang Wang, Tieniu Tan

机构 * New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Peking University(北京大学) Nanjing University(南京大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

AI总结 DiaDem通过合成高质量数据集和难度分区的两阶段GRPO策略,提升了音频视频字幕中的对话描述准确性,并在多种基准测试中表现出色。

Comments Project webpage: https://diadem-captioner.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14233 2026-01-28 cs.CV 74%

Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception

通过视觉属性增强描述性标题以实现多模态感知

Yanpeng Sun, Jing Hao, Ke Zhu, Jiang-Jiang Liu, Yuxiang Zhao, Xiaofan Li, Na Zhao, Zechao Li, Jingdong Wang

机构 * NJUST(南京理工大学) Baidu VIS(百度视觉部) HKU(香港大学) NJU(南京大学) SUTD(新加坡科技设计大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出EDC方法,通过整合视觉专家的低级和细粒度属性,提升图像标题的描述质量,以增强多模态感知能力。

Comments An open-source Agent for generating detailed image captions

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16973 2026-01-26 cs.CV 74%

VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents

VisGym: 多模态代理的多样化、可定制化、可扩展环境

Zirui Wang, Junyi Zhang, Jiaxin Ge, Long Lian, Letian Fu, Lisa Dunlap, Ken Goldberg, XuDong Wang, Ion Stoica, David M. Chan, Sewon Min, Joseph E. Gonzalez

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 VisGym通过多样化环境评估多模态代理在多步视觉决策中的表现,揭示模型在长上下文处理和视觉化任务中的局限性。

Comments Project page: https://visgym.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16108 2026-01-23 cs.AI 74%

Multimodal Climate Disinformation Detection: Integrating Vision-Language Models with External Knowledge Sources

多模态气候虚假信息检测:整合视觉-语言模型与外部知识源

Marzieh Adeli Shamsabad, Hamed Ghodrati

机构 * CRIM

专题命中 图文多模态 :multimodal(title);分类 cs.AI

AI总结 本文提出整合视觉-语言模型与外部知识源,以提升多模态气候虚假信息检测的准确性与实时性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00597 2025-12-02 cs.CV 74%

Scaling Down to Scale Up: Towards Operationally-Efficient and Deployable Clinical Models via Cross-Modal Low-Rank Adaptation for Medical Vision-Language Models

缩小规模以扩大规模:通过跨模态低秩适应实现操作高效且可部署的临床模型

Thuraya Alzubaidi, Farhad R. Nezami, Muzammil Behzad

机构 * King Fahd University of Petroleum(国王法赫德石油与矿物大学) Institute for Medical Engineering(医学工程研究所) Science, Massachusetts Institute of Technology, US(科学,麻省理工学院,美国) Harvard Medical School, Harvard University, US(哈佛医学院,哈佛大学,美国) SDAIA-KFUPM Joint Research Center for Artificial Intelligence, Saudi Arabia(SDAIA-KFUPM人工智能联合研究中心,沙特阿拉伯)

专题命中 图文多模态 :cross-modal(title);分类 cs.CV

AI总结 通过跨模态低秩适应,MedCT-VLM在零样本分类中实现了对CT影像的高效适应,显著提升了病理分类的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17103 2025-11-24 cs.CV 74%

Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs

弥合视觉情感差距:通过学习噪声图像-文本对借用文本知识

Daiqing Wu, Dongbao Yang, Yu Zhou, Can Ma

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) TMCC, College of Computer Science, Nankai University(TMCC,南开大学计算机学院)

专题命中 图文多模态 :image-text(title);分类 cs.CV

AI总结 本文提出通过学习噪声图像-文本对中的文本知识,弥合视觉情感识别中的情感差距,提升预训练视觉模型的情感感知能力。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06009 2025-11-19 cs.CV 74%

Continual Learning for Image Captioning through Improved Image-Text Alignment

Bertram Taetz, Gal Bordelius

机构 * IT & Engineering International University of Applied Sciences(IT与工程国际应用科学大学)

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02044 2025-11-11 cs.CV 74%

A multi-modal vision-language model for generalizable annotation-free pathology localization

Hao Yang, Hong-Yu Zhou, Jiarun Liu, Weijian Huang, Cheng Li, Zhihuan Li, Yuanxu Gao, Qiegen Liu, Yong Liang, Qi Yang, Song Wu, Tao Tan, Hairong Zheng, Kang Zhang, Shanshan Wang

机构 * Paul C. Lauterbur Research Center for Biomedical Imaging(生物医学成像研究中心) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Pengcheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Chinese Medicine Guangdong Laboratory(广东中医药实验室) Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院)

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22838 2025-10-28 cs.CV 74%

Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models

Aya Nakayama, Brian Wong, Yuji Nishimura, Kaito Tanaka

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20625 2025-10-28 cs.CV 74%

T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting

Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound

机构 * University of Nottingham(诺丁汉大学) University of St Andrews(圣安德鲁大学) City University of Hong Kong(香港城市大学) Nanyang Technology University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23827 2025-09-30 cs.CV cs.LG 74%

Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models

Efthymios Tsaprazlis, Tiantian Feng, Anil Ramakrishna, Rahul Gupta, Shrikanth Narayanan

机构 * University of Southern California(南加州大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11247 2025-09-16 cs.CV 74%

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States

Robert Long, Rongxin Jiang, Mingrui Yan

机构 * University of Padua(帕多瓦大学) Heilongjiang University of Science and Technology(黑龙江科技大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04678 2025-08-20 eess.IV cs.CV 74%

RadGPT: Constructing 3D Image-Text Tumor Datasets

Pedro R. A. S. Bassi, Mehmet Can Yavuz, Kang Wang, Xiaoxi Chen, Wenxuan Li, Sergio Decherchi, Andrea Cavalli, Yang Yang, Alan Yuille, Zongwei Zhou

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Bologna(博洛尼亚大学) Italian Institute of Technology(意大利理工学院) University of California, San Francisco(加州大学旧金山分校) Istanbul Medipol University(伊斯坦布尔Medipol大学) University of Zurich(苏黎世大学) ETH AI Center(ETH人工智能中心) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

专题命中 图文多模态 :image-text(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.01955 2025-08-19 cs.CV 74%

Adaptively Clustering Neighbor Elements for Image-Text Generation

Zihua Wang, Xu Yang, Hanwang Zhang, Haiyang Xu, Ming Yan, Fei Huang, Yu Zhang

机构 * School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学) DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments This work has been accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10339 2025-08-15 cs.CV cs.LG 74%

Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models

Andrew Bai, Justin Cui, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15621 2025-08-01 cs.CV cs.AI cs.CL cs.MM 74%

LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学) IIT-CNR(意大利国家研究 council(IIT))

专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

Comments ICCV 2025 Workshop on What is Next in Multimodal Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19370 2025-07-28 cs.CV 74%

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Felix Brandstaetter, Erik Schuetz, Katharina Winter, Fabian Flohr

机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14953 2025-07-18 cs.CV 74%

Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

Yang Liu, Wentao Feng, Zhuoyao Liu, Shudong Huang, Jiancheng Lv

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Engineering Research Center of Machine Learning and Industry Intelligence(机器学习与工业智能工程研究中心)

专题命中 图文多模态 :image-text(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12236 2025-07-17 cs.CV 74%

Generate to Ground: Multimodal Text Conditioning Boosts Phrase Grounding in Medical Vision-Language Models

Felix Nützel, Mischa Dombrowski, Bernhard Kainz

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗赖堡-亚历山大大学埃尔兰根-纽伦堡) Imperial College London(伦敦帝国理工学院)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments 20 pages, 6 figures. To appear in Proc. MIDL 2025 (PMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18491 2025-06-12 cs.CL 74%

MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering

Shuo Yang, Siwen Luo, Soyeon Caren Han, Eduard Hovy

机构 * The University of Melbourne(墨尔本大学) The University of Western Australia(西澳大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

Comments Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05441 2025-06-09 eess.IV cs.CV cs.LG 74%

Deep histological synthesis from mass spectrometry imaging for multimodal registration

Kimberley M. Bird, Xujiong Ye, Alan M. Race, James M. Brown

机构 * University of Lincoln(林肯大学) University of Exeter(埃克塞特大学) AstraZeneca Computational Pathology GmbH(阿斯利康计算病理学 GmbH)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments Medical Image Understanding and Analysis (MIUA) 2025 Extended Abstract Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20937 2025-05-28 cs.CL 74%

On VLMs for Diverse Tasks in Multimodal Meme Classification

Deepesh Gavit, Debajyoti Mazumder, Samiran Das, Jasabanta Patro

机构 * Peng Wang and Shuai Bai and Sinan Tan and Shijie Wang and Zhihao Fan and Jinze Bai and Keqin Chen and Xuejing Liu and Jialin Wang and Wenbin Ge and Yang Fan and Kai Dang and Mengfei Du and Xuancheng Ren and Rui Men and Dayiheng Liu and Chang Zhou and Jingren Zhou and Junyang Lin(研究人员)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏