arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45832 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4634 篇

2602.00135 2026-02-03 cs.CV 83%

LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models

LLaVA-FA: 通过傅里叶近似压缩大型多模态模型

Pengcheng Zheng, Chaoning Zhang, Jiarong Mo, GuoHui Li, Jiaquan Zhang, Jiahao Zhang, Sihan Cao, Sheng Zheng, Caiyan Qin, Guoqing Wang, Yang Yang

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 LLaVA-FA通过频域联合低秩和量化近似,实现高效压缩大型多模态模型,采用极坐标量化和对角校准方案,提升压缩效果与效率。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05895 2026-01-28 cs.CV 83%

BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model

BTCChat: 通过多模态大语言模型推进遥感双时相变化描述

Yujie Li, Wenjia Xu, Yuanben Zhang, Zhiwei Wei, Mugen Peng

机构 * State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室) Beijing University of Posts and Telecommunications(北京邮电大学) Aerospace Information Research Institute(航天信息研究所) Chinese Academy of Sciences(中国科学院) School of Geographical Sciences(地理科学学院) Hunan Normal University(湖南师范大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 BTCChat通过多模态大语言模型提升遥感双时相变化描述能力,引入变化提取模块和提示增强机制,实现更精确的视觉-语义对齐和更优的性能表现。

Comments 5 pages, 2 figures; Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11808 2026-01-27 cs.CV cs.AI cs.CL cs.CY cs.MM 83%

Labels or Input? Rethinking Augmentation in Multimodal Hate Detection

标签还是输入?重新思考多模态仇恨检测中的增强

Sahajpreet Singh, Kokil Jaidka, Subhayan Mukerjee

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出通过提示优化、微调和自动化数据增强改进小型模型,开发多模态增强框架以提升隐含仇恨检测性能。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23243 2026-01-12 cs.CV 83%

Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism

遥感图像的多模态解释:动态分辨率输入策略与多尺度视觉-语言对齐机制

Siyu Zhang, Lianlei Shan, Runhe Qiu

机构 * Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出动态分辨率输入策略与多尺度视觉-语言对齐机制,提升遥感图像多模态解释的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10969 2026-01-01 cs.CV 83%

Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation

弥合一致性差距:用于交错图像-文本生成的显式结构化记忆

Zeteng Lin, Xingxing Li, Wen You, Xiaoyang Li, Zehan Lu, Yujun Cai, Jing Tang

机构 * Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州)) University of Queensland(昆士兰大学)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 IUT-Plug通过显式结构化记忆机制解决多模态生成中的上下文漂移问题,提升长序列一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18033 2025-12-29 cs.RO cs.AI 83%

OpenNav: Open-World Navigation with Multimodal Large Language Models

OpenNav: 基于多模态大语言模型的开放世界导航

Mingfeng Yuan, Letian Wang, Steven L. Waslander

机构 * University of Toronto Institute for Aerospace Studies(多伦多大学航空航天研究 institute) University of Toronto Robotics Institute(多伦多大学机器人研究所)

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 OpenNav利用多模态大语言模型实现开放世界导航,通过生成价值图增强机器人空间理解,并在真实场景中验证其鲁棒性。

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21583 2025-12-29 cs.AI 83%

A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning

一种整合视觉-语言模型和逻辑树推理的医学多模态诊断框架

Zelin Zang, Wenyi Gu, Siqi Ma, Dan Yang, Yue Shen, Zhu Zhang, Guohui Fan, Wing-Kuen Ling, Fuji Yang

机构 * Tsientang Institute of Advanced Study (TIAS)(钱塘先进研究所) Westlake University(西湖大学) Ant Group(蚂蚁集团) China-Japan Friendship Hospital(中日友好医院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种整合视觉-语言模型与逻辑树推理的医学诊断框架,旨在提升多模态医学AI的可信度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20257 2025-12-24 cs.CV 83%

LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation

LADLE-MM:基于有限标注的多模态虚假信息检测器

Daniele Cardullo, Simone Teglia, Irene Amerini

机构 * Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 LADLE-MM是一种基于有限标注的多模态虚假信息检测器,通过学习的集成方法在有限资源下实现高效检测,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17178 2025-12-22 cs.CV cs.IR 83%

ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching

ABE-CLIP: 无训练属性绑定增强用于组合图像-文本匹配

Qi Zhang, Yuxu Chen, Lei Deng, Lili Shen

机构 * School of Mathematics, Sichuan University(四川大学数学学院)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 ABE-CLIP通过语义细化和局部对齐策略提升CLIP模型的属性-物体绑定性能,无需额外训练。

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12107 2025-12-16 cs.CV 83%

EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography

EchoVLM:基于测量的多模态学习用于超声心动图

Yuheng Li, Yue Zhang, Abdoul Aziz Amadou, Yuxiang Lai, Jike Zhong, Tiziano Passerini, Dorin Comaniciu, Puneet Sharma

机构 * Georgia Institute of Technology(佐治亚理工学院) Siemens Healthineers(西门子医疗) Siemens Healthcare Limited(西门子医疗有限公司) Emory University(埃默里大学) University of Southern California(南加州大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 EchoVLM通过引入基于测量的多模态学习方法,实现了超声心动图的端到端解读,提升了疾病分类和视图识别的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02278 2025-12-09 cs.CV 83%

SCMM: Calibrating Cross-modal Representations for Text-Based Person Search

SCMM:基于文本的人脸搜索的跨模态表示校准

Jing Liu, Donglai Wei, Yang Liu, Sipeng Zhang, Tong Yang, Wei Zhou, Weiping Ding, Victor C. M. Leung

机构 * College of Future Information Technology, Fudan University(复旦大学未来信息技术学院) Department of Electrical and Computer Engineering, The University of British Columbia(不列颠哥伦比亚大学电气与计算机工程系) College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院) MEGVII Technology(MEGVII技术) School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) School of AI and CS, Nantong University(南通大学人工智能与计算机科学学院) Academy of Artificial Intelligence, SMBU(SMBU人工智能学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 SCMM通过缝校准和掩码建模方法,提升跨模态表示学习在文本驱动的人脸搜索中的性能。

Comments 11 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00396 2025-11-27 cs.CV 83%

Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning

Saliency-R1: 通过置信度引导的强化学习激励多模态大语言模型统一的显著性推理能力

Long Li, Shuichen Ji, Ziyang Luo, Zhihui Li, Dingwen Zhang, Junwei Han, Nian Liu

机构 * Northwestern Polytechnical University(西北工业大学) University of Science and Technology of China(中国科学技术大学)

专题命中 图文多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Saliency-R1通过置信度引导的强化学习,提升多模态大语言模型在显著性推理任务中的统一能力,实现显著物体检测、显著实例分割和共显著物体检测的高效处理。

Comments Main text (excluding references): 8 pages, 4 figures; Supplementary Materials (excluding references): 9 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12215 2025-11-18 cs.CV 83%

FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention

Peng Zhang, Zhihui Lai, Wenting Chen, Xu Wu, Heng Kong

机构 * Peng Zhang 1 , Zhihui Lai 1 , Wenting Chen 2 1 1 footnotemark: 1 , Xu Wu 1 , Heng Kong 3(某机构)

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11984 2025-11-18 cs.CV 83%

From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology

Zhenhao Guo, Rachit Saluja, Tianyuan Yao, Quan Liu, Junchao Zhu, Haibo Wang, Daniel Reisenbüchler, Yuankai Huo, Benjamin Liechty, David J. Pisapia, Kenji Ikemura, Steven Salvatoree, Surya Seshane, Mert R. Sabuncu, Yihe Yang, Ruining Deng

机构 * New York University(纽约大学) Cornell Tech(康奈尔科技) Vanderbilt University(范德比大学) Carnegie Mellon University(卡内基梅隆大学) University of Regensburg(莱茵河畔大学) Weill Cornell Medicine(韦尔·科恩医学中心) Northwell Health(北well健康)

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21375 2025-11-05 cs.CV 83%

GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution

Fengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang, Haotian Wang, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang, Hongzhen Wang, Wenjing Yang, Bo Du, Jing Zhang

机构 * College of Computer Science and Technology, National University of Defense Technology, China(中国国防科技大学计算机科学与技术学院) Beijing University of Posts and Telecommunications, China(北京邮电大学) School of Computer Science, Wuhan University, China(武汉大学计算机学院) Zhongguancun Academy, China(中关村学院) Tsinghua University, China(清华大学) Beihang University, China(北航大学)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments NeurlPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20392 2025-10-31 cs.CV 83%

Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning

Zhifang Zhang, Shuo He, Haobo Wang, Bingquan Shen, Lei Feng

机构 * Southeast University(东南大学) University of Queensland(昆士兰大学) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25303 2025-10-30 cs.CL 83%

Teaching Sarcasm: Few-Shot Multimodal Sarcasm Detection via Distillation to a Parameter-Efficient Student

Soumyadeep Jana, Sanasam Ranbir Singh

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Indian Institute of Technology Guwahati(印度理工学院古瓦哈蒂)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13497 2025-10-16 cs.LG cs.AI 83%

DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation

Zexin Wang, Lin Shi, Haoyu Wu, Junru Luo, Xiangzeng Kong, Jun Qi

机构 * Aliyun School of Big Data, Changzhou University(阿里云大数据学院,长洲大学) Department of Computing, Xi’an JiaoTong-Liverpool University(计算系,西安交通大学-利物浦大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Center for Artificial Intelligence in Agriculture, Fujian Agriculture and Forestry University(农业人工智能中心,福建农林大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10466 2025-10-14 cs.CV 83%

When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance

Jinjin Cao, Zhiyang Chen, Zijun Wang, Liyuan Ma, Weijian Luo, Guojun Qi

机构 * MAPLE Lab, Westlake University(西溪大学MAPLE实验室)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18369 2025-10-13 cs.CV 83%

RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models

Yeongtak Oh, Dohyun Chung, Juhyeon Shin, Sangha Park, Johan Barthelemy, Jisoo Mok, Sungroh Yoon

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07567 2025-10-10 cs.CV 83%

Cross-Modal Attention Guided Unlearning in Vision-Language Models

Karuna Bhaila, Aneesh Komanduri, Minh-Hao Van, Xintao Wu

专题命中 图文多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07216 2025-10-08 cs.CV 83%

Bridging Semantic Logic Gaps: A Cognition Inspired Multimodal Boundary Preserving Network for Image Manipulation Localization

Songlin Li, Zhiqing Guo, Yuanman Li, Zeyu Li, Yunfeng Diao, Gaobo Yang, Liejun Wang

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) School of Electronic and Information Engineering, Shenzhen University(深圳大学电子与信息工程学院) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02608 2025-10-07 cs.AI 83%

Mitigating Modal Imbalance in Multimodal Reasoning

Chen Henry Wu, Neil Kale, Aditi Raghunathan

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 10 pages, 10 figures, CoLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23673 2025-09-30 cs.CV cs.AI cs.CL cs.MM 83%

RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks

Amit Agarwal, Hitesh Laxmichand Patel, Srikant Panda, Hansa Meghwani, Jyotika Singh, Karan Dua, Paul Li, Tao Sheng, Sujith Ravi, Dan Roth

机构 * Oracle AI

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12587 2025-09-25 cs.CV 83%

Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models

Tan-Hanh Pham, Chris Ngo

机构 * Harvard Medical School, Harvard University(哈佛医学院、哈佛大学) Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital(阿提尼乌拉A.马丁努斯生物医学成像中心、麻省总医院) Knovel Engineering Lab, Singapore(Knovel工程实验室、新加坡)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08742 2025-09-11 q-fin.CP cs.AI 83%

FinZero: Launching Multi-modal Financial Time Series Forecast with Large Reasoning Model

Yanlong Wang, Jian Xu, Fei Ma, Hongkang Zhang, Hang Yu, Tiantian Gao, Yu Wang, Haochen You, Shao-Lun Huang, Danny Dongning Sun, Xiao-Ping Zhang

机构 * Tsinghua University(清华大学) Pengcheng Laboratory(鹏城实验室) Guangming Laboratory(光明实验室) Ant Group(蚂蚁集团) Columbia University(哥伦比亚大学) Southern University of Science and Technology(南方科技大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08715 2025-09-11 cs.CV 83%

BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion

Sike Xiang, Shuang Chen, Amir Atapour-Abarghouei

机构 * Department of Computer Science, Durham University(计算机科学系,杜伦大学)

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01644 2025-09-03 cs.CV 83%

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Yanqing Liu, Xianhang Li, Letian Zhang, Zirui Wang, Zeyu Zheng, Yuyin Zhou, Cihang Xie

机构 * University of California Santa Cruz(加州大学圣克鲁兹分校) Apple(苹果公司) University of California Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00039 2025-09-03 cs.CV 83%

AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models

Yuqi Li, Chuanguang Yang, Junhao Dong, Zhengtao Yao, Haoyan Xu, Zeyu Dong, Hansheng Zeng, Zhulin An, Yingli Tian

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11793 2025-08-27 cs.CV 83%

MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise

Ruiqi Wu, Chenran Zhang, Jianle Zhang, Yi Zhou, Tao Zhou, Huazhu Fu

机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China(教育部新一代人工智能技术及其交叉应用重点实验室) Nanjing University of Science and Technology, Nanjing, China(南京理工大学) Agency for Science, Technology and Research (A*STAR), Singapore(新加坡科技研究局)

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Early Accepted by The International Conference on Medical Image Computing and Computer Assisted Intervention(MICCAI)2024

详情

展开后加载摘要…

URL PDF HTML 收藏