arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45832 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4634 篇

2602.11636 2026-02-13 cs.CV cs.AI 81%

ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning

ScalSelect: 可扩展的无训练多模态数据选择用于高效的视觉指令微调

Changti Wu, Jiahuai Mao, Yuzhuo Miao, Shijie Lian, Bin Yu, Xiaopeng Lin, Cong Huang, Lei Zhang, Kai Chen

机构 * East China Normal University(华东师范大学) Zhongguancun Academy(中关村学院) The Hong Kong Polytechnic University(香港理工大学) Harbin Institute of Technology(哈尔滨工业大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 ScalSelect提出一种无训练、可扩展的多模态数据选择方法,通过线性时间复杂度实现高效视觉指令微调,实验显示其性能接近甚至超越全数据训练。

Comments The code is available at \href{https://github.com/ChangtiWu/ScalSelect}{ScalSelect}

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07017 2026-02-10 cs.CV cs.AI 81%

XAI-CLIP: ROI-Guided Perturbation Framework for Explainable Medical Image Segmentation in Multimodal Vision-Language Models

XAI-CLIP: 通过区域感兴趣引导扰动框架实现多模态视觉-语言模型中可解释的医学图像分割

Thuraya Alzubaidi, Sana Ammar, Maryam Alsharqi, Islem Rekik, Muzammil Behzad

机构 * King Fahd University of Petroleum and Minerals(国王法赫德石油和矿物大学) Massachusetts Institute of Technology(麻省理工学院) Imperial College London(伦敦帝国学院) KFUPM-SDAIA Joint Research Centre for Artificial Intelligence(KFUPM-SDAIA联合人工智能研究中心)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 XAI-CLIP通过多模态视觉-语言模型嵌入实现医学图像分割的可解释性和效率提升,减少计算开销并提高分割精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07209 2026-02-06 cs.CV cs.AI cs.GR 81%

SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model

SIRR-LMM:通过大多模态模型实现单图像反射去除

Yu Guo, Zhiqiang Lao, Xiyun Song, Yubin Zhou, Heather Yu

机构 * Futurewei Technologies(未来科技公司)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出SIRR-LMM方法,通过大模型和合成数据集生成技术,提升单图像反射去除的性能。

Comments 12 pages, 14 figures, accepted in WACVW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03007 2026-02-04 cs.CV cs.AI cs.LG 81%

VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering

VOILA:基于信息价值的保真度选择框架用于成本感知的多模态问答

Rahul Atul Bhope, K. R. Jayaram, Vinod Muthusamy, Ritesh Kumar, Vatche Isahagian, Nalini Venkatasubramanian

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 VOILA通过基于信息价值的自适应保真度选择,在多模态问答中实现成本优化与高精度的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20042 2026-02-03 cs.CV cs.AI 81%

Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval

超越视觉:基于多模态检索的上下文丰富图像描述

Nguyen Lam Phu Quy, Pham Phu Hoa, Tran Chi Nguyen, Dao Sy Duy Minh, Nguyen Hoang Minh Ngoc, Huynh Trung Kiet

机构 * University of Science - VNUHCM(越南胡志明市大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种多模态检索方法,通过整合外部文本知识生成更丰富的上下文图像描述,提升视觉-文本理解能力。

Comments 7 pages, 5 figures. System description for the EVENTA Grand Challenge (Track 1) at ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05609 2026-02-03 cs.CV cs.AI 81%

HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection

HOI-R1: 探索多模态大语言模型在人类-物体交互检测中的潜力

Junwen Chen, Peilin Xiong, Keiji Yanai

机构 * Department of Informatics, The University of Electro-Communications(信息学系,东京电通大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 HOI-R1利用多模态大语言模型的推理能力,无需额外模块实现人类-物体交互检测任务的提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09851 2026-01-28 cs.CV cs.AI cs.HC 81%

ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning

ViSIL:多模态视频描述信息损失的统一评估

Po-han Li, Shenghui Chen, Ufuk Topcu, Sandeep Chinchali

机构 * The University of Texas at Austin, Texas, USA(德克萨斯大学奥斯汀分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 ViSIL通过信息论框架量化多模态视频摘要的信息损失,实现跨格式的统一评估,并在VQA任务中提升准确率7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11020 2026-01-27 cs.CV cs.AI 81%

GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation

GeoVLMath: 通过跨模态奖励增强视觉-语言模型中的几何推理以辅助线创建

Shasha Guo, Liang Pang, Xi Wang, Yanling Wang, Huawei Shen, Jing Zhang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Renmin University of China(中国人民大学) Zhipu AI(智谱AI)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 GeoVLMath通过跨模态奖励模型提升视觉-语言模型在复杂立体几何问题中的几何推理能力。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16218 2026-01-26 cs.CL cs.AI 81%

M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models

M3Kang: 评估视觉-语言模型在多语言多模态数学推理中的表现

Aleix Torres-Camps, Nathaniel Mitrani Hadida, Víctor Conchello Vendrell, Àlex Batlle Casellas, Arnau Padrés Masdemont, Jordi Ros-Giralt

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 M3Kang是一个大规模多语言多模态数学推理数据集,用于评估视觉-语言模型在多语言数学推理中的表现,通过基准测试揭示模型在基础数学和图表推理上的不足,并展示多语言技术在多模态设置中的有效性。

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10094 2026-01-16 cs.CV cs.AI cs.LG 81%

V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation

V-Zero:基于零标注的自改进多模态推理

Han Wang, Yi Yang, Jingyuan Hu, Minfeng Zhu, Wei Chen

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) Zhejiang University(浙江大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 V-Zero通过自改进机制在无需人工标注的情况下提升多模态模型的视觉数学推理和通用视觉能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21518 2026-01-15 cs.CV cs.CL cs.LG 81%

Head Pursuit: Probing Attention Specialization in Multimodal Transformers

头部追踪:探究多模态转换器中的注意力专业化

Lorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello, Alberto Cazzaniga

机构 * Area Science Park(面积科学公园) Sapienza University of Rome(罗马萨皮恩扎大学) Institute of Science and Technology(科学与技术研究所)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本研究通过分析多模态转换器中注意力头的专业化,揭示了模型内部可控的结构,并展示了通过编辑少量头部以增强或抑制特定概念的可行性。

Comments Accepted at NeurIPS 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09116 2026-01-15 cs.CV cs.AI 81%

LP-LLM: End-to-End Real-World Degraded License Plate Text Recognition via Large Multimodal Models

LP-LLM:基于大多模态模型的端到端真实世界退化车牌文本识别

Haoyan Gong, Hongbin Liu

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 LP-LLM通过端到端结构感知多模态推理框架,结合字符感知模块和LoRA微调策略,提升真实世界退化车牌文本识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06010 2026-01-14 cs.CV cs.MM 81%

Latent Reconstruction from Generated Data for Multimodal Misinformation Detection

从生成数据中进行潜在重建用于多模态虚假信息检测

Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis

机构 * Information Technology Institute, Centre for Research & Technology, Hellas(信息科技研究所,研究中心,希腊) Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,亚里士多德大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 本研究提出MisCaption This!框架和LAMAR网络,通过生成高保真度的误标数据和潜在重建技术,提升多模态虚假信息检测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23545 2025-12-30 cs.CV cs.AI 81%

PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis

PathFound: 一种促进证据寻求病理诊断的代理多模态模型

Shengyi Hua, Jianfeng Wu, Tianle Shen, Kangzhe Hu, Zhongzhen Huang, Shujuan Ni, Zhihong Zhang, Yuan Li, Zhe Wang, Xiaofan Zhang

机构 * Qing Yuan Research Institute, Shanghai Jiao Tong University(上海交通大学庆元研究院) Shanghai Innovation Institute(上海创新研究院) Department of Pathology, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学生命科学与医学学院病理科) Intelligent Pathology Institute, Division of Life Sciences and Medicine(生命科学与医学学院智能病理研究所) Department of Pathology, Fudan University Shanghai Cancer Center(复旦大学上海癌症中心病理科) Department of Oncology, Shanghai Medical College, Fudan University(复旦大学上海医学院肿瘤科) Institute of Pathology, Fudan University(复旦大学病理研究所) Department of Pathology, The First Affiliated Hospital with Nanjing Medical University(南京医科大学第一附属医院病理科)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 PathFound是一种通过主动信息获取和诊断完善来提升病理诊断准确性的代理多模态模型,其在多种临床场景中表现出卓越的诊断性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17097 2025-12-11 cs.CV cs.CL 81%

Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning

让LVLMs聚焦:基于上下文的注意力调节以提升多模态上下文学习

Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li, Ligong Han, Hongyang He, Zhengtao Yao, Yingjie Victor Chen, Songlin Fei, Dongfang Liu, Ruixiang Tang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出CAMA,一种无需训练的注意力调节方法,通过动态调整注意力logits提升多模态上下文学习性能。

Comments 14 pages, 8 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19094 2025-12-04 cs.CV cs.AI 81%

SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring

SATORI-R1: 通过显式视觉锚定激励多模态推理

Chuming Shen, Wei Wei, Xiaoye Qu, Yu Cheng

机构 * Huazhong University of Science and Technology(华中科技大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SATORI-R1通过显式视觉锚定提升多模态推理,采用三个可验证阶段和VQA-Verify数据集,在VQA任务中实现15.7%的准确率提升。

Comments 21 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23375 2025-12-01 cs.CL cs.CV 81%

Optimizing Multimodal Language Models through Attention-based Interpretability

通过基于注意力的可解释性优化多模态语言模型

Alexander Sergeev, Evgeny Kotelnikov

机构 * European University at Saint Petersburg(圣彼得堡欧洲大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 通过基于注意力的可解释性方法优化多模态语言模型,识别关键对象相关的注意力头以提高效率和性能。

Comments Accepted for ICAI-2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21947 2025-12-01 cs.CV cs.AI cs.LG 81%

WalkCLIP: Multimodal Learning for Urban Walkability Prediction

WalkCLIP: 城市步行性预测的多模态学习

Shilong Xiang, JangHyeon Lee, Min Namgung, Yao-Yi Chiang

机构 * University of Minnesota Twin Cities(明尼苏达大学双城分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 WalkCLIP通过整合视觉和行为数据,实现了对城市步行性的高精度预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06665 2025-11-11 cs.CV cs.AI 81%

Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks

Lingran Song, Yucheng Zhou, Jianbing Shen

机构 * Lingran Song, Yucheng Zhou, Jianbing Shen(作者)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05577 2025-11-11 cs.LG cond-mat.mtrl-sci cs.AI cs.CL 81%

Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction

An Vuong, Minh-Hao Van, Prateek Verma, Chen Zhao, Xintao Wu

机构 * Department of EECS University of Arkansas(电子工程与科学系 亚拉荷加大学) Department of CS Baylor University(计算机科学系 基尔默大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17902 2025-11-10 cs.CV cs.CL 81%

TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection

Girish A. Koushik, Helen Treharne, Aditya Joshi, Diptesh Kanojia

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to Special Track on AI for Social Impact (AISI) at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15418 2025-11-10 cs.CL cs.AI 81%

Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs

Lee Qi Zun, Mohamad Zulhilmi Bin Abdul Halim, Goh Man Fye

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00095 2025-11-04 cs.CV cs.AI 81%

SpinalSAM-R1: A Vision-Language Multimodal Interactive System for Spine CT Segmentation

Jiaming Liu, Dingwei Fan, Junyong Zhao, Chunlin Li, Haipeng Si, Liang Sun

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学) Department of Orthopedics, Qilu Hospital, Shandong University(骨科部,齐鲁医院,山东大学) Key Laboratory of Qingdao in Medicine and Engineering(医学与工程青岛重点实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 2 Tables,5 Figures,16 Equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12149 2025-11-04 cs.CL cs.MM cs.SI 81%

Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models

Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang, Hang Yu

机构 * Anhui Polytechnic University (AHPU)(安徽工程大学) Shanghai University (SHU)(上海大学) Shanghai Jiao Tong University (SJTU)(上海交通大学) Sichuan University (SCU)(四川大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17098 2025-10-22 cs.CL cs.CV 81%

TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration

Yanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng, Jinfa Huang, Ruixiang Tang

机构 * Brown University(布朗大学) University of Bristol(布里斯托大学) Columbia University(哥伦比亚大学) University of Rochester(罗切斯特大学) Rutgers University(罗格斯大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments EMNLP2025 Main, 28 pages, 11 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14138 2025-10-16 cs.CV cs.AI 81%

ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom

Jingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu, Lei Li, Jiahui Gao, Jiyue Jiang, Lingpeng Kong, Chuan Wu

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11115 2025-10-14 cs.CV cs.MM 81%

Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning

Hao Tang, Shengfeng He, Jing Qin

机构 * Centre for Smart Health, The Hong Kong Polytechnic University(香港理工大学智能健康中心) School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算机与信息系统学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23250 2025-10-08 cs.AI cs.CV 81%

Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned

Brandon Ong, Tej Deep Pala, Vernon Toh, William Chandra Tjhi, Soujanya Poria

机构 * AI Singapore(AI新加坡) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04145 2025-10-07 cs.CV cs.CL cs.IR 81%

Automating construction safety inspections using a multi-modal vision-language RAG framework

Chenxin Wang, Elyas Asadi Shamsabadi, Zhaohui Chen, Luming Shen, Alireza Ahmadian Fard Fini, Daniel Dias-da-Costa

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments 33 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03295 2025-10-07 cs.CV cs.CL cs.LG 81%

Multimodal Arabic Captioning with Interpretable Visual Concept Integration

Passant Elchafei, Amany Fashwan

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏