arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4878 篇

2412.20206 2026-08-05 cs.CV 版本更新 70%

Toward Visual Grounding: A Survey

视觉定位:一项综述

Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang, Changsheng Xu

专题命中 其他多模态 :multimodal(abstract,abstract_cn);分类 cs.CV

AI总结 本综述梳理视觉定位的发展与背景,总结近年进展与新挑战,定义规范研究设置,介绍相关数据集与应用,提出未来方向,是该领域最全面的综述,适合不同阶段研究者。

Comments Accepted by TPAMI 2025. We keep tracing related works at https://github.com/linhuixiao/Awesome-Visual-Grounding, article publication page: https://ieeexplore.ieee.org/abstract/document/11235566

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 3, pp. 2749-2771, March 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24187 2026-07-28 cs.AI 新提交 70%

Myopia Prevention and Control 3.0: Artificial Intelligence--Driven Risk Stratification, Proactive Monitoring, and Personalized Intervention

近视防控3.0:人工智能驱动的风险分层、主动监测和个性化干预

Tieniu Wang, Cangzhu Huang, Qianhui Li

机构 * Shanghai Nile Intelligent Technology Co., Ltd.(上海尼罗智能科技有限公司) Beijing Tanyuan Academy of Intelligent Sensing(北京潭渊智能传感研究院)

专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 研究借助人工智能、数字传感等技术推动近视防控从被动变主动精准模式。通过多模态数据机器学习预测风险、可穿戴等监测及个性化干预形成闭环。评估各阶段证据,讨论相关挑战并给出未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28920 2026-06-30 cs.CV 70%

ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images

ExACT: 基于示例驱动的校准精化用于遥感图像中免训练的视觉定位

Zixiao Zhang, Lingling Li, Pei He, Xu Liu, Licheng Jiao

机构 * Xidian University(西安电子科技大学)

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出ExACT框架,通过一次性视觉提示机制弥合多模态大语言模型在遥感视觉定位中的模态差距,实现免训练的精确像素级定位。

Comments 11 pages, 8 figures, supplementary material included

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23356 2026-06-23 cs.CV cs.LG 新提交 70%

Changing Modalities: Adapting Remote Sensing Models to New Satellites and Sensors

改变模态:将遥感模型适应新卫星和传感器

Tim G. Zhou, Anthony Fuller, Geoff Pleiss, Evan Shelhamer

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Carleton University(卡尔顿大学)

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 针对遥感模型在新卫星和传感器上的部署问题,提出DeluluNet架构,通过模态幻觉实现模态迁移、添加和子集三种场景下的模型适应,无需重新标注。

Comments 17 pages, 7 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00143 2026-06-23 cs.LG cs.CV 版本更新 70%

Is Oracle Pruning the True Oracle?

Oracle剪枝真的是真正的Oracle吗?

Sicheng Feng, Keda Tao, Huan Wang

机构 * Westlake University(西湖大学) Nankai University(南开大学) ENCODE Lab, Westlake University(西湖大学ENCODE实验室)

专题命中 其他多模态 :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文通过大规模实验(37K模型)发现,对于中等规模以上的深度学习模型,Oracle剪枝选择的权重在重训练后性能与重训练前几乎无关,质疑了Oracle剪枝作为剪枝方法基础的有效性。

Comments TMLR, Webpage: https://fscdc.github.io/Oracle-Pruning-Sanity-Check/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17953 2026-06-17 cs.CV 新提交 70%

MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias

MLLMs 先正确后错误:追踪并纠正后层文本偏见

Xingming Li, Ao Cheng, Qiyao Sun, Xixiang He, Xuanyu Ji, Runke Huang, Qingyong Hu

机构 * National University of Defense Technology(国防科技大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Intelligent Game and Decision Lab(智能博弈与决策实验室)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract_cn);分类 cs.CV

AI总结 发现多模态大语言模型在中间层形成正确视觉预测,但最终输出时被文本覆盖,通过检测预测方向变化(85%失败转向文本,89%成功转向视觉)提出无训练方法CALRD,在冲突基准上提升高达9.4%。

Comments Accepted at IJCAI 2026. 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00240 2026-06-02 cs.AI cs.MA 70%

MindZero: Learning Online Mental Reasoning With Zero Annotations

MindZero:零标注的在线心智推理学习

Shunchi Zhang, Jin Lu, Chuanyang Jin, Yichao Zhou, Zhining Zhang, Tianmin Shu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 提出MindZero框架,通过自监督强化学习训练多模态大语言模型,实现高效鲁棒的在线心智推理,无需显式心智状态标注。

Comments ICML 2026. Website: https://scai.cs.jhu.edu/MindZero

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27179 2026-03-31 cs.CV 70%

Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision

基于推理的异常检测与定位:图像级监督

Yizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang, Qingjie Liu, Yunhong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室) Hangzhou Innovation Institute, Beihang University(北京航空航天大学杭州创新研究院)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出基于图像级监督的异常检测与定位方法,利用大语言模型的推理能力实现像素级定位,无需额外组件或标注。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17712 2026-03-19 cs.RO cs.CV 70%

AERR-Nav: Adaptive Exploration-Recovery-Reminiscing Strategy for Zero-Shot Object Navigation

AERR-Nav:面向零样本物体导航的自适应探索-恢复-回忆策略

Jingzhi Huang, Junkai Huang, Haoyang Yang, Haoang Li, Yi Wang

机构 * Hong Kong Polytechnic University(香港理工大学) Institute of automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出AERR-Nav框架,通过自适应探索-恢复-回忆策略和自适应探索状态,解决零样本物体导航中探索与利用的平衡问题,在HM3D和MP3D基准测试中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02541 2026-03-10 cs.CV 70%

Mix-modal Federated Learning for MRI Image Segmentation

多模态联邦学习用于磁共振成像图像分割

Guyue Hu, Siyuan Song, Jingpeng Sun, Zhe Jin, Chenglong Li, Jin Tang

机构 * School of Artificial Intelligence(人工智能学院) School of Computer Science and Technology(计算机科学与技术学院) State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) Anhui Provincial Key Laboratory of Security Artificial Intelligence(安徽省安全人工智能重点实验室) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室)

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种多模态联邦学习框架,用于解决MRI图像分割中的模态异质性和数据异质性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18869 2026-02-24 cs.CV 70%

Enhancing 3D LiDAR Segmentation by Shaping Dense and Accurate 2D Semantic Predictions

通过塑造密集且准确的2D语义预测来增强3D激光雷达分割

Xiaoyu Dong, Tiankui Xian, Wanshui Gan, Naoto Yokoya

机构 * The University of Tokyo(东京大学)

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MM2D3D模型,通过多模态引导滤波和动态跨伪监督提升2D预测质量,从而增强3D激光雷达分割的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17315 2026-02-17 cs.CL 70%

HIPPO: Enhancing the Table Understanding Capability of LLMs through Hybrid-Modal Preference Optimization

HIPPO:通过混合模态偏好优化增强大语言模型的表格理解能力

Haolan Wang, Zhenghao Liu, Xinze Li, Xiaocui Yang, Yu Gu, Yukun Yan, Qi Shi, Fangfang Li, Chong Chen, Ge Yu

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院) Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 其他多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CL

AI总结 HIPPO 通过混合模态偏好优化提升大语言模型的表格理解能力,实现 4% 的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18240 2026-01-27 cs.CV 70%

V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering

V-Loop:用于医学视觉问答中幻觉检测的视觉逻辑循环验证

Mengyuan Jin, Zehui Liao, Yong Xia

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 V-Loop通过双向推理和视觉逻辑循环验证,提升医学视觉问答中幻觉检测的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16584 2025-12-19 cs.CV 70%

Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs

Sketch-in-Latents: 在潜在空间中实现多模态统一推理

Jintao Tong, Jiaqi Gu, Yujing Lou, Lubin Fan, Yixiong Zou, Yue Wu, Jieping Ye, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) Alibaba Cloud Computing(阿里巴巴云计算)

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 SkiLa通过在潜在空间中实现多模态统一推理,扩展MLLMs的自回归能力,生成连续视觉嵌入,提升视觉任务性能和多模态泛化能力。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11099 2025-12-15 cs.CV 70%

VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction

VGent: 通过模块化设计实现视觉 grounding 的解耦推理与预测

Weitai Kang, Jason Kuen, Mengwei Ren, Zijun Wei, Yan Yan, Kangning Liu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Adobe(Adobe公司)

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 VGent 通过模块化设计实现视觉 grounding 的解耦推理与预测,利用冻结 MLLM 和解码器交叉注意力机制,提升多目标识别性能。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17686 2025-10-21 cs.CV 70%

Towards 3D Objectness Learning in an Open World

Taichi Liu, Zhenyu Wang, Ruofeng Liu, Guang Wang, Desheng Zhang

机构 * Rutgers University(罗格斯大学) Tsinghua University(清华大学) Michigan State University(密歇根州立大学) Florida State University(佛罗里达州立大学)

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17087 2025-09-23 cs.AI 70%

Governing Automated Strategic Intelligence

Nicholas Kruus, Madhavendra Thakur, Adam Khoja, Leonhard Nagel, Maximilian Nicholson, Abeer Sharma, Jason Hausenloy, Alberto KoTafoya, Aliya Mukhanova, Alli Katila-Miikkulainen, Harish Chandran, Ivan Zhang, Jessie Chen, Joel Raj, Jord Nguyen, Lai Hsien Hao, Neja Jayasundara, Soham Sen, Sophie Zhang, Ashley Dora Kokui Tamaklo, Bhavya Thakur, Henry Close, Janghee Lee, Nina Sefton, Raghavendra Thakur, Shiv Munagala, Yeeun Kim

专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03094 2025-08-06 cs.CV 70%

Augmenting Continual Learning of Diseases with LLM-Generated Visual Concepts

Jiantao Tan, Peixian Ma, Kanghao Chen, Zhiming Dai, Ruixuan Wang

机构 * Sun Yat-sen University(中山大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Peng Cheng Laboratory(鹏城实验室) Key Laboratory of Machine Intelligence and Advanced Computing, MOE(教育部机器智能与先进计算重点实验室)

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03158 2025-06-05 cs.LG cs.CV 70%

DUAL: Dynamic Uncertainty-Aware Learning

Jiahao Qin, Bei Peng, Feng Liu, Guangliang Cheng, Lu Zong

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07159 2025-05-13 eess.IV cs.CV 70%

Skull stripping with purely synthetic data

Jong Sung Park, Juhyung Ha, Siddhesh Thakur, Alexandra Badea, Spyridon Bakas, Eleftherios Garyfallidis

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Oral at ISMRM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14171 2025-04-22 cs.AI 70%

Adaptation Method for Misinformation Identification

Yangping Chen, Weijie Shi, Mengze Li, Yue Cui, Hao Chen, Jia Zhu, Jiajie Xu

机构 * Soochow University(苏霍沃大学) Hong Kong University of Science and Technology(香港科技大学) Tencent(腾讯) Zhejiang Normal University(浙江师范大学)

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13341 2025-04-15 cs.CV 70%

Multi-aspect Knowledge Distillation with Large Language Model

Taegyeong Lee, Jinsik Bang, Soyeong Kwon, Taehwan Kim

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accept to CVPRW2025 (FGVC12)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06147 2025-03-13 eess.SP cs.AI 70%

Multiclass Arrhythmia Classification using Smartwatch Photoplethysmography Signals Collected in Real-life Settings

Dong Han, Jihye Moon, Luís Roberto Mercado Díaz, Darren Chen, Devan Williams, Eric Y. Ding, Khanh-Van Tran, David D. McManus, Ki H. Chon

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09218 2025-01-17 q-bio.QM cs.AI 70%

Interpretable Droplet Digital PCR Assay for Trustworthy Molecular Diagnostics

Yuanyuan Wei, Yucheng Wu, Fuyang Qu, Yao Mu, Yi-Ping Ho, Ho-Pui Ho, Wu Yuan, Mingkun Xu

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15761 2025-01-15 cs.CV 70%

MambaTrack: Exploiting Dual-Enhancement for Night UAV Tracking

Chunhui Zhang, Li Liu, Hao Wen, Xi Zhou, Yanfeng Wang

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09027 2024-12-02 cs.CV cs.AI cs.CL cs.MM 70%

CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension

Zhi Zhang, Helen Yannakoudakis, Xiantong Zhen, Ekaterina Shutova

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref EACL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15881 2024-10-24 cs.CV 70%

LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Fangxun Shu, Yue Liao, Le Zhuo, Chenning Xu, Lei Zhang, Guanghao Zhang, Haonan Shi, Long Chen, Tao Zhong, Wanggui He, Siming Fu, Haoyuan Li, Bolin Li, Zhelun Yu, Si Liu, Hongsheng Li, Hao Jiang

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03062 2024-10-07 cs.AI 70%

Image First or Text First? Optimising the Sequencing of Modalities in Large Language Model Prompting and Reasoning Tasks

Grant Wardle, Teo Susnjak

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16312 2024-09-26 q-bio.QM cs.AI eess.SP 70%

SEE: Semantically Aligned EEG-to-Text Translation

Yitian Tao, Yan Liang, Luoyu Wang, Yongqing Li, Qing Yang, Han Zhang

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02695 2024-08-07 cs.LG cs.AI 70%

Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding Confusion

Shaoxu Cheng, Kanglei Geng, Chiyuan He, Zihuan Qiu, Linfeng Xu, Heqian Qiu, Lanxiao Wang, Qingbo Wu, Fanman Meng, Hongliang Li

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏