arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45986 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2603.21010 2026-03-24 cs.CV 74%

SkinCLIP-VL: Consistency-Aware Vision-Language Learning for Multimodal Skin Cancer Diagnosis

SkinCLIP-VL: 一种面向多模态皮肤癌诊断的一致性感知视觉-语言学习框架

Zhixiang Lu, Shijie Xu, Kaicheng Yan, Xuyue Cai, Chong Zhang, Yulong Li, Angelos Stefanidis, Anh Nguyen, Jionglong Su

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出SkinCLIP-VL,通过冻结感知与自适应推理范式,结合CLIP编码器与轻量量化Qwen2.5-VL,引入一致性聚焦对齐损失,实现高效皮肤癌诊断,优于13B参数基线模型,参数更少且临床信任度更高。

Comments Accepted by 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06263 2026-03-10 cs.CV 74%

iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models

iLLaVA: 图像的价值少于三分之一的输入标记在大型多模态模型中

Lianyu Hu, Liqing Gao, Fanhua Shang, Liang Wan, Wei Feng

机构 * School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院) School of Computer Science and Technology, Tiangong University(天津工大学计算机科学与技术学院) Key Research Center for Surface Monitoring and Analysis of Relics, State Administration of Cultural Heritage(文物表面监测与分析关键研究中心,国家文物局)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 iLLaVA通过联合优化图像编码器和LLM,减少冗余标记并提升吞吐量,实现更高效的多模态模型性能。

Comments Accepted by ICLR2026,code is released at https://github.com/hulianyuyy/iLLaVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00251 2026-03-05 cs.SE cs.AI cs.SY eess.SY 74%

GENAI WORKBENCH: AI-Assisted Analysis and Synthesis of Engineering Systems from Multimodal Engineering Data

GenAI Workbench: 人工智能辅助从多模态工程数据中分析和合成工程系统

H. Sinan Bank, Daniel R. Herber

机构 * Colorado State University(科罗拉多州立大学)

专题命中 图文多模态 :multimodal(title);分类 cs.AI

AI总结 GenAI Workbench通过整合AI技术,实现从多模态工程数据中辅助分析和合成工程系统,促进更集成和数据驱动的工程设计方法。

Comments 7 pages, 3 figures, accepted to be presented at IISE Annual Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22120 2026-02-06 cs.CV 74%

See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning

看得少,看得对:双向感知塑造用于多模态推理

Shuoshuo Zhang, Yizhen Zhang, Jingjing Fu, Lei Song, Jiang Bian, Yujiu Yang, Rui Wang

机构 * Microsoft Research Asia(微软亚洲研究院) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出BiPS,通过双向感知塑造提升多模态推理性能,有效增强细粒度视觉依赖并提升跨领域泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19267 2026-01-28 cs.CL 74%

DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models

DiaDem: 促进音频视频字幕中的对话描述以提升多模态大语言模型

Xinlong Chen, Weihong Lin, Jingyun Hua, Linli Yao, Yue Ding, Bozhou Li, Bohan Zeng, Yang Shi, Qiang Liu, Yuanxing Zhang, Pengfei Wan, Liang Wang, Tieniu Tan

机构 * New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Peking University(北京大学) Nanjing University(南京大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

AI总结 DiaDem通过合成高质量数据集和难度分区的两阶段GRPO策略,提升了音频视频字幕中的对话描述准确性,并在多种基准测试中表现出色。

Comments Project webpage: https://diadem-captioner.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14233 2026-01-28 cs.CV 74%

Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception

通过视觉属性增强描述性标题以实现多模态感知

Yanpeng Sun, Jing Hao, Ke Zhu, Jiang-Jiang Liu, Yuxiang Zhao, Xiaofan Li, Na Zhao, Zechao Li, Jingdong Wang

机构 * NJUST(南京理工大学) Baidu VIS(百度视觉部) HKU(香港大学) NJU(南京大学) SUTD(新加坡科技设计大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出EDC方法,通过整合视觉专家的低级和细粒度属性,提升图像标题的描述质量,以增强多模态感知能力。

Comments An open-source Agent for generating detailed image captions

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16973 2026-01-26 cs.CV 74%

VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents

VisGym: 多模态代理的多样化、可定制化、可扩展环境

Zirui Wang, Junyi Zhang, Jiaxin Ge, Long Lian, Letian Fu, Lisa Dunlap, Ken Goldberg, XuDong Wang, Ion Stoica, David M. Chan, Sewon Min, Joseph E. Gonzalez

专题命中 图文多模态 :multimodal(title);分类 cs.CV

AI总结 VisGym通过多样化环境评估多模态代理在多步视觉决策中的表现,揭示模型在长上下文处理和视觉化任务中的局限性。

Comments Project page: https://visgym.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16108 2026-01-23 cs.AI 74%

Multimodal Climate Disinformation Detection: Integrating Vision-Language Models with External Knowledge Sources

多模态气候虚假信息检测:整合视觉-语言模型与外部知识源

Marzieh Adeli Shamsabad, Hamed Ghodrati

机构 * CRIM

专题命中 图文多模态 :multimodal(title);分类 cs.AI

AI总结 本文提出整合视觉-语言模型与外部知识源,以提升多模态气候虚假信息检测的准确性与实时性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00597 2025-12-02 cs.CV 74%

Scaling Down to Scale Up: Towards Operationally-Efficient and Deployable Clinical Models via Cross-Modal Low-Rank Adaptation for Medical Vision-Language Models

缩小规模以扩大规模:通过跨模态低秩适应实现操作高效且可部署的临床模型

Thuraya Alzubaidi, Farhad R. Nezami, Muzammil Behzad

机构 * King Fahd University of Petroleum(国王法赫德石油与矿物大学) Institute for Medical Engineering(医学工程研究所) Science, Massachusetts Institute of Technology, US(科学,麻省理工学院,美国) Harvard Medical School, Harvard University, US(哈佛医学院,哈佛大学,美国) SDAIA-KFUPM Joint Research Center for Artificial Intelligence, Saudi Arabia(SDAIA-KFUPM人工智能联合研究中心,沙特阿拉伯)

专题命中 图文多模态 :cross-modal(title);分类 cs.CV

AI总结 通过跨模态低秩适应,MedCT-VLM在零样本分类中实现了对CT影像的高效适应,显著提升了病理分类的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17103 2025-11-24 cs.CV 74%

Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs

弥合视觉情感差距:通过学习噪声图像-文本对借用文本知识

Daiqing Wu, Dongbao Yang, Yu Zhou, Can Ma

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) TMCC, College of Computer Science, Nankai University(TMCC,南开大学计算机学院)

专题命中 图文多模态 :image-text(title);分类 cs.CV

AI总结 本文提出通过学习噪声图像-文本对中的文本知识,弥合视觉情感识别中的情感差距,提升预训练视觉模型的情感感知能力。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06009 2025-11-19 cs.CV 74%

Continual Learning for Image Captioning through Improved Image-Text Alignment

Bertram Taetz, Gal Bordelius

机构 * IT & Engineering International University of Applied Sciences(IT与工程国际应用科学大学)

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02044 2025-11-11 cs.CV 74%

A multi-modal vision-language model for generalizable annotation-free pathology localization

Hao Yang, Hong-Yu Zhou, Jiarun Liu, Weijian Huang, Cheng Li, Zhihuan Li, Yuanxu Gao, Qiegen Liu, Yong Liang, Qi Yang, Song Wu, Tao Tan, Hairong Zheng, Kang Zhang, Shanshan Wang

机构 * Paul C. Lauterbur Research Center for Biomedical Imaging(生物医学成像研究中心) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Pengcheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Chinese Medicine Guangdong Laboratory(广东中医药实验室) Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院)

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22838 2025-10-28 cs.CV 74%

Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models

Aya Nakayama, Brian Wong, Yuji Nishimura, Kaito Tanaka

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20625 2025-10-28 cs.CV 74%

T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting

Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound

机构 * University of Nottingham(诺丁汉大学) University of St Andrews(圣安德鲁大学) City University of Hong Kong(香港城市大学) Nanyang Technology University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23827 2025-09-30 cs.CV cs.LG 74%

Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models

Efthymios Tsaprazlis, Tiantian Feng, Anil Ramakrishna, Rahul Gupta, Shrikanth Narayanan

机构 * University of Southern California(南加州大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11247 2025-09-16 cs.CV 74%

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States

Robert Long, Rongxin Jiang, Mingrui Yan

机构 * University of Padua(帕多瓦大学) Heilongjiang University of Science and Technology(黑龙江科技大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04678 2025-08-20 eess.IV cs.CV 74%

RadGPT: Constructing 3D Image-Text Tumor Datasets

Pedro R. A. S. Bassi, Mehmet Can Yavuz, Kang Wang, Xiaoxi Chen, Wenxuan Li, Sergio Decherchi, Andrea Cavalli, Yang Yang, Alan Yuille, Zongwei Zhou

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Bologna(博洛尼亚大学) Italian Institute of Technology(意大利理工学院) University of California, San Francisco(加州大学旧金山分校) Istanbul Medipol University(伊斯坦布尔Medipol大学) University of Zurich(苏黎世大学) ETH AI Center(ETH人工智能中心) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

专题命中 图文多模态 :image-text(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.01955 2025-08-19 cs.CV 74%

Adaptively Clustering Neighbor Elements for Image-Text Generation

Zihua Wang, Xu Yang, Hanwang Zhang, Haiyang Xu, Ming Yan, Fei Huang, Yu Zhang

机构 * School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学) DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments This work has been accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10339 2025-08-15 cs.CV cs.LG 74%

Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models

Andrew Bai, Justin Cui, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15621 2025-08-01 cs.CV cs.AI cs.CL cs.MM 74%

LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学) IIT-CNR(意大利国家研究 council(IIT))

专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

Comments ICCV 2025 Workshop on What is Next in Multimodal Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19370 2025-07-28 cs.CV 74%

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Felix Brandstaetter, Erik Schuetz, Katharina Winter, Fabian Flohr

机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14953 2025-07-18 cs.CV 74%

Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

Yang Liu, Wentao Feng, Zhuoyao Liu, Shudong Huang, Jiancheng Lv

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Engineering Research Center of Machine Learning and Industry Intelligence(机器学习与工业智能工程研究中心)

专题命中 图文多模态 :image-text(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12236 2025-07-17 cs.CV 74%

Generate to Ground: Multimodal Text Conditioning Boosts Phrase Grounding in Medical Vision-Language Models

Felix Nützel, Mischa Dombrowski, Bernhard Kainz

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗赖堡-亚历山大大学埃尔兰根-纽伦堡) Imperial College London(伦敦帝国理工学院)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments 20 pages, 6 figures. To appear in Proc. MIDL 2025 (PMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18491 2025-06-12 cs.CL 74%

MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering

Shuo Yang, Siwen Luo, Soyeon Caren Han, Eduard Hovy

机构 * The University of Melbourne(墨尔本大学) The University of Western Australia(西澳大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

Comments Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05441 2025-06-09 eess.IV cs.CV cs.LG 74%

Deep histological synthesis from mass spectrometry imaging for multimodal registration

Kimberley M. Bird, Xujiong Ye, Alan M. Race, James M. Brown

机构 * University of Lincoln(林肯大学) University of Exeter(埃克塞特大学) AstraZeneca Computational Pathology GmbH(阿斯利康计算病理学 GmbH)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments Medical Image Understanding and Analysis (MIUA) 2025 Extended Abstract Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20937 2025-05-28 cs.CL 74%

On VLMs for Diverse Tasks in Multimodal Meme Classification

Deepesh Gavit, Debajyoti Mazumder, Samiran Das, Jasabanta Patro

机构 * Peng Wang and Shuai Bai and Sinan Tan and Shijie Wang and Zhihao Fan and Jinze Bai and Keqin Chen and Xuejing Liu and Jialin Wang and Wenbin Ge and Yang Fan and Kai Dang and Mengfei Du and Xuancheng Ren and Rui Men and Dayiheng Liu and Chang Zhou and Jingren Zhou and Junyang Lin(研究人员)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11576 2025-03-17 cs.CV 74%

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A. Said Gurbuz, Michele Dolfi, Miquel Farré, Peter W. J. Staar

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17251 2024-12-24 cs.CV cs.LG eess.IV 74%

GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning

Teja Krishna Cherukuri, Nagur Shareef Shaik, Jyostna Devi Bodapati, Dong Hye Ye

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments This paper has been accepted for presentation at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01725 2024-12-03 cs.CV 74%

Attacks on multimodal models

Viacheslav Iablochnikov, Alexander Rogachev

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments 19 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19296 2024-11-19 cs.AI 74%

Multi-Modal CLIP-Informed Protein Editing

Mingze Yin, Hanjing Zhou, Yiheng Zhu, Miao Lin, Yixuan Wu, Jialu Wu, Hongxia Xu, Chang-Yu Hsieh, Tingjun Hou, Jintai Chen, Jian Wu

专题命中 图文多模态 :multi-modal(title);分类 cs.AI

Comments 13 pages, 7 figures, 5 tables

Journal ref Health Data Science, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏