arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-10 至 2026-03-10 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 14 篇

2509.11165 2026-03-10 cs.CV 83%

Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning

Traffic-MLLM: 基于好奇心正则化的交通场景案例推理监督学习

Waikit Xiu, Qiang Lu, Bingchen Liu, Chen Sun, Xiying Li

机构 * The University of Hong Kong(香港大学) Sun Yat-sen University(中山大学) University of Bristol(布里斯托大学)

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Traffic-MLLM通过好奇心驱动的结构化案例学习,提升多模态交通场景推理的鲁棒性和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08202 2026-03-10 cs.CV cs.AI 81%

MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data

MM-TS: 多模态温度和边距调度用于长尾数据的对比学习

Siarhei Sheludzko, Dhimitrios Duka, Bernt Schiele, Hilde Kuehne, Anna Kukleva

机构 * University of Bonn(波恩大学) MPI for Informatics, SIC(信息研究所) Tuebingen AI Center/University of Tuebingen(图宾根人工智能中心/图宾根大学) MIT-IBM Watson AI Lab(麻省理工-IBM Watson人工智能实验室)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 MM-TS通过动态温度和边距调度提升多模态对比学习在长尾数据上的性能,实现InfoNCE和最大边距方法的统一。

Comments 18 pages, 11 figures. Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07997 2026-03-10 cs.AI 79%

CMMR-VLN: Vision-and-Language Navigation via Continual Multimodal Memory Retrieval

CMMR-VLN:通过持续多模态记忆检索实现视觉与语言导航

Haozhou Li, Xiangyu Dong, Huiyan Jiang, Yaoming Zhou, Xiaoguang Ma

机构 * Foshan Graduate School of Innovation at Northeastern University(东北大学创新研究生院) Faculty of Robot Science and Engineering at Northeastern University(东北大学机器人科学与工程学院) College of Software at Northeastern University(东北大学软件学院) School of Aeronautic Science and Engineering at Beihang University(北航航空科学与工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 CMMR-VLN通过引入持续多模态记忆检索机制,提升视觉与语言导航任务中对先前经验的选择性利用能力,显著提高导航成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07039 2026-03-10 cs.AI 79%

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

具有4D空间-时间嵌入的自监督多模态世界模型

Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang

机构 * Ecological Intelligence Lab(生态智能实验室) School of Complex Adaptive Systems(复杂适应系统学院) University of Houston(休斯顿大学) Geosensing Systems Engineering & Sciences Lab(传感系统工程与科学实验室) Stanford University(斯坦福大学) Allen Institute for Artificial Intelligence(人工智能研究院) Spatial Intelligence Lab(空间智能实验室) Department of Computer Science(计算机科学系) Georgia Institute of Technology(佐治亚理工学院) Florida Museum of Natural History(佛罗里达自然历史博物馆) University of Florida(佛罗里达大学) NSF Institute for Geospatial Understanding(国家科学基金会地理理解研究所) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 DeepEarth通过4D空间-时间嵌入实现自监督多模态世界模型,在生态预测中取得最佳性能。

Comments 8 pages, 5 figures, 1 table. Presented at 2026 World Modeling Workshop, Mila Quebec

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06982 2026-03-10 cs.CV cs.IR 79%

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

优化多模态模型用于基于图像的形状检索:预对齐和硬对比学习的作用

Paul Julius Kühn, Cedric Spengler, Michael Weinmann, Arjan Kuijper, Saptarshi Neil Sinha

机构 * Fraunhofer IGD(弗劳恩霍夫图像研究中心) Delft University of Technology(代尔夫特理工大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出通过预对齐和硬对比学习优化多模态模型,提升基于图像的形状检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06704 2026-03-10 cs.CV cs.LG 70%

On the Generalization Capacities of MLLMs for Spatial Intelligence

关于多模态大语言模型在空间智能中的泛化能力

Gongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang, Deli Zhao, Shijian Lu, Ran Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) HuPan Lab(虎朋实验室) Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出Camera-Aware MLLM框架,通过注入相机内参、数据增强和蒸馏几何先验,提升多模态大语言模型在空间任务中的泛化能力与鲁棒性。

Comments ICLR 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21194 2026-03-10 cs.CV cs.AI 62%

BotaCLIP: Contrastive Learning for Botany-Aware Representation of Earth Observation Data

BotaCLIP:基于对比学习的植物学感知地球观测数据表示

Selene Cerna, Sara Si-Moussi, Wilfried Thuiller, Hadrien Hendrikx, Vincent Miele

机构 * univ-grenoble-alpes(格勒诺布尔阿尔卑斯大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 BotaCLIP通过对比学习将植物学知识注入地球观测数据表示,提升生物多样性建模任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11148 2026-03-10 cs.CL cs.AI 62%

Improving the Efficiency of Visually Augmented Language Models

提升视觉增强语言模型的效率

Paula Ontalvilla, Aitor Ormazabal, Gorka Azkune

机构 * HiTZ Center - Ixa, University of the Basque Country (UPV/EHU)(巴斯克大学HiTZ中心 - Ixa,巴斯克大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出BLIND-VALM模型,通过使用CLIP获取的视觉 grounding 文本表示,提升视觉增强语言模型的效率和性能。

Comments COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07497 2026-03-10 cs.CV 57%

AMR-CCR: Anchored Modular Retrieval for Continual Chinese Character Recognition

AMR-CCR: 以锚定模块检索实现连续中文字符识别

Yuchuan Wu, Yinglian Zhu, Haiyang Yu, Ke Niu, Bin Li, Xiangyang Xue

机构 * Fudan University(复旦大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 AMR-CCR通过锚定模块检索框架实现连续中文字符识别,解决持续类增长和脚本多样性问题,提出轻量级脚本注入模块和多原型字典以提升识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07493 2026-03-10 cs.CV 57%

RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object Detection

RayD3D: 沿射线蒸馏深度知识以实现鲁棒的多视角3D目标检测

Rui Ding, Zhaonian Kuang, Zongwei Zhou, Meng Yang, Xinhu Zheng, Gang Hua

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 RayD3D通过沿射线蒸馏深度知识,提升多视角3D目标检测的鲁棒性,无需增加计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19112 2026-03-10 cs.CV 57%

Universal 3D Shape Matching via Coarse-to-Fine Language Guidance

通过粗到细的语言引导实现通用3D形状匹配

Qinfeng Xiao, Guofeng Mei, Bo Yang, Liying Zhang, Jian Zhang, Kit-lun Yick

机构 * Hong Kong Polytechnic University, HK SAR(香港理工大学) Fondazione Bruno Kessler, Italy(布鲁诺·凯斯勒基金会) University of Technology Sydney, Australia(悉尼科技大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 UniMatch通过粗到细的语言引导方法,实现跨类别非等距形状的通用3D匹配。

Comments Accepted by CVPR 2026

Journal ref CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06722 2026-03-10 cs.LG cs.AI 57%

ProtAlign: Contrastive learning paradigm for Sequence and structure alignment

ProtAlign: 用于序列和结构对齐的对比学习范式

Aditya Ranganath, Hasin Us Sami, Kowshik Thopalli, Bhavya Kailkhura, Wesam Sakla

机构 * Center for Applied Scientific Computing, Lawrence Livermore National Laboratory(应用科学计算中心,劳伦斯利弗莫尔国家实验室)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

AI总结 ProtAlign通过对比学习范式,统一了蛋白质序列与结构的表示,提升跨模态检索和下游预测性能。

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07611 2026-03-10 cs.HC 50%

Beyond Semantic Similarity: Open Challenges for Embedding-Based Creative Process Analysis Across AI Design Tools

超越语义相似性:基于嵌入的AI设计工具创意过程分析的开放挑战

Seung Won Lee, Semin Jin, Kyung Hoon Hyun

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文探讨了基于嵌入的AI设计工具创意过程分析中超越语义相似性的开放挑战,提出通过上下文感知干预提升分析敏感性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06927 2026-03-10 cs.RO 50%

A Contrastive Fewshot RGBD Traversability Segmentation Framework for Indoor Robotic Navigation

一种用于室内外机器人导航的对比式少样本RGBD可 traversability 分割框架

Qiyuan An, Tuan Dang, Fillia Makedon

机构 * Department of Computer Science and Engineering, University of Texas at Arlington(计算机科学与工程系,德克萨斯大学阿灵顿分校) Uber(优步) Cognitive Robotics Lab, Department of Electrical Engineering and Computer Science, University of Arkansas(认知机器人实验室,电气工程与计算机科学系,阿肯色大学)

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 本文提出了一种基于对比学习的少样本RGBD分割框架,利用负例原型和稀疏深度信息提升室内机器人导航的可 traversability 分割性能。

Journal ref IEEE International Conference on Robotics & Automation 2026 (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏