arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2606.28344 2026-06-30 cs.IR cs.AI cs.CL cs.CV cs.LG 67%

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

PIXELRAG:网页截图在检索增强生成中优于文本

Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出PixelRAG方法,以网页截图替代文本进行检索和阅读,利用视觉嵌入模型和对比学习,在30M截图库上实现端到端RAG,在多项任务中优于文本基线,准确率提升最高18.1%,并通过图像压缩降低3倍token成本。

Comments Our code is available at https://github.com/StarTrail-org/PixelRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10020 2026-06-05 cs.CL cs.AI cs.CV 67%

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

性能提升的幻象:为何对比解码无法减轻多模态大语言模型中的对象幻觉?

Hao Yin, Guangzong Si, Zilei Wang

机构 * University of Science and Technology of China(中国科学技术大学) Eastern Institute of Technology, Ningbo(宁波东部技术研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文研究了对比解码方法在减轻多模态大语言模型(MLLMs)中对象幻觉方面的有效性,发现其性能提升主要源于两个误导性因素,挑战了对比解码策略的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27009 2026-05-27 cs.LG 67%

SCENT: Aligning Mass Spectra with Molecular Structure for Olfactory Perception

SCENT: 将质谱与分子结构对齐用于嗅觉感知

Ziqi Zhang, Eunyeong Jin, Miguel Vasco, Farzaneh Taleb, Nona Rajabi, Alexandra Gutmann, Jonathan Williams, Antônio H. Ribeiro, Danica Kragic

机构 * Dept. of Intelligent Systems, KTH Royal Institute of Technology(智能系统系,皇家理工学院) Atmospheric Chemistry Dept., Max Planck Institute for Chemistry(大气化学部,马克斯·普朗克研究所) Dept. of Information Technology, Uppsala University(信息科技系,乌普萨拉大学) Science for Life Laboratory (SciLifeLab), Uppsala(生命科学实验室(SciLifeLab),乌普萨拉)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract)

AI总结 提出SCENT多模态对比学习框架,通过将电子电离质谱表示与预训练化学结构嵌入对齐,在无需分子结构的情况下实现与结构模型相当的嗅觉预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22715 2026-05-26 cs.CV cs.AI cs.CL cs.HC 67%

AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

AnyMo:野外人体运动的几何感知与设置无关建模

Baiyu Chen, Zechen Li, Wilson Wongso, Lihuan Li, Xiachong Lin, Hao Xue, Benjamin Tag, Flora Salim

机构 * The University of New South Wales(新南威尔士大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出AnyMo框架,通过物理模拟生成多样化IMU信号、图编码器预训练和LLM对齐,实现跨设备/数据集的零样本活动识别、跨模态检索和运动描述,性能显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26467 2026-04-30 cs.CR 67%

Differentially Private Contrastive Learning via Bounding Group-level Contribution

差分隐私对比学习通过限制群体级贡献

Kecen Li, Chen Gong, Zinan Lin, Tianhao Wang, Xiaokui Xiao

专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract)

AI总结 本文提出DP-GCL框架,通过限制梯度依赖提升差分隐私对比学习效果,实验显示在多个数据集上均取得最佳性能,图像分类准确率提升5.6%,图像-文本检索准确率提升20.1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18123 2026-04-03 cs.CV cs.AI cs.CL cs.LG 67%

Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models

偏差是子空间,而非坐标:视觉-语言模型中事后去偏的几何重思

Dachuan Zhao, Weiyue Li, Zhenda Shen, Yushu Qiu, Bowen Xu, Haoyu Chen, Yongchao Chen

机构 * Harvard University(哈佛大学) MIT(麻省理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出SPD框架,通过几何方法识别并去除线性可解码的偏子空间,提升视觉-语言模型的公平性与任务性能。

Comments Accepted at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16737 2026-03-18 cs.CV cs.AI cs.CL 67%

Retrieving Counterfactuals Improves Visual In-Context Learning

检索反事实改进视觉上下文学习

Guangzhi Xiong, Sanchit Sinha, Zhenghao He, Aidong Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出CIRCLES框架,通过主动检索反事实样例提升视觉语言模型的因果推理能力,实验表明其在多个数据集上优于现有方法,尤其在小规模模型和信息稀缺场景下表现突出。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02556 2026-03-04 cs.CV cs.AI cs.CL cs.LG 67%

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

通过对比的视角:VLMs中的自改进视觉推理

Zhiyu Pan, Yizheng Wu, Jiashen Hua, Junyi Feng, Shaotian Yan, Bing Deng, Zhiguo Cao, Jieping Ye

机构 * Huazhong University of Science and Technology(华中科技大学) Alibaba Cloud(阿里云)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 通过视觉对比提升VLMs的推理能力,提出VC-STaR框架,有效减少推理幻觉并提升多种VLMs的视觉推理性能。

Comments 19 pages, 9 figures, accepted to ICLR 2026 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01990 2026-03-03 cs.AI cs.CL cs.CV 67%

According to Me: Long-Term Personalized Referential Memory QA

根据我:长期个性化参照记忆问答

Jingbiao Mei, Jinghong Chen, Guangyu Yang, Xinyu Hou, Margaret Li, Bill Byrne

机构 * Department of Engineering, University of Cambridge, United Kingdom(剑桥大学工程系) Department of Physics, University of Cambridge, United Kingdom(剑桥大学物理系) Independent Researcher(独立研究者)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出ATM-Bench,首个多模态多来源个性化参照记忆问答基准,并提出Schema-Guided Memory方法提升记忆推理性能

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01493 2026-03-03 cs.IR cs.AI cs.CV cs.MM 67%

PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval

PhotoBench: 超越视觉匹配,迈向个性化意图驱动的图片检索

Tianyi Xu, Rong Shan, Junjie Wu, Jiadeng Huang, Teng Wang, Jiachen Zhu, Wenteng Chen, Minxin Tu, Quantao Dou, Zhaoxiang Wang, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 PhotoBench 是首个基于真实个人相册构建的基准,旨在通过多源意图驱动推理提升个性化图片检索能力。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04451 2026-02-06 cs.IR 67%

SDR-CIR: Semantic Debias Retrieval Framework for Training-Free Zero-Shot Composed Image Retrieval

SDR-CIR:一种基于链式推理的语义去偏检索框架用于无训练零样本复合图像检索

Yi Sun, Jinyu Xu, Qing Xie, Jiachen Li, Yanchun Ma, Yongjian Liu

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract)

AI总结 SDR-CIR通过语义去偏排名方法提升无训练零样本复合图像检索性能。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19606 2026-01-28 cs.CV cs.AI cs.LG cs.SD eess.AS 67%

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining

GMS-CAVP:通过多尺度对比和生成预训练提升音频视频对应关系

Shentong Mo, Zehua Chen, Jun Zhu

机构 * Carnegie Mellon University(卡内基梅隆大学) MBZUAI Tsinghua University(清华大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI、eess.AS

AI总结 GMS-CAVP通过多尺度对比和生成预训练方法提升音频视频对应关系建模,实现更深入的跨模态理解和高保真生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18850 2026-01-28 cs.SE 67%

Towards Safety-Compliant Transformer Architectures for Automotive Systems

面向汽车系统的安全合规Transformer架构

Sven Kirchner, Nils Purschke, Chengdong Wu, Alois Knoll

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract)

AI总结 本文提出了一种安全合规的Transformer架构,通过多模态基础模型提升汽车系统的容错性和鲁棒性,实现自动驾驶中的可认证AI系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14607 2026-01-22 physics.geo-ph 67%

SeisBind: Physics-Aware Tri-Modal Representation Binding for Seismic Data via Contrastive Learning

SeisBind:通过对比学习实现地震数据的物理感知三模态表示绑定

Chaohua Liang, Jun Matsushima

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract)

AI总结 SeisBind通过对比学习将地震数据与物理描述符结合,实现更可解释的地下速度模型构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20145 2025-12-24 cs.CL cs.AI cs.CV cs.IR cs.LG 67%

Retrieval-augmented Prompt Learning for Pre-trained Foundation Models

基于检索的提示学习用于预训练基础模型

Xiang Chen, Yixin Ou, Quan Feng, Lei Li, Piji Li, Haibo Ye, Sheng-Jun Huang, Shuofei Qiao, Shumin Deng, Huajun Chen, Ningyu Zhang

机构 * MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(信息产业部模式分析与机器智能重点实验室,计算机科学与技术学院,南京航空航天大学) Zhejiang University(浙江大学) Hunan Vanguard Group Corporation Limited(湖南先锋集团有限公司) National University of Singapore(新加坡国立大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 RetroPrompt通过整合检索机制和知识库,提升预训练基础模型在少样本和零样本场景下的泛化能力与记忆平衡。

Comments IEEE/ACM Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17537 2025-12-10 cs.AI cs.CL cs.CV 67%

CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale

CLIBD:融合视觉与基因组学用于大规模生物多样性监测

ZeMing Gong, Austin T. Wang, Xiaoliang Huo, Joakim Bruslund Haurum, Scott C. Lowe, Graham W. Taylor, Angel X. Chang

机构 * Simon Fraser University(西蒙弗雷泽大学) Aalborg University(奥胡斯大学) Vector Institute(向量研究所) University of Guelph(圭尔夫大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 CLIBD通过融合视觉与基因组学数据,利用对比学习实现对昆虫物种的高效分类,提升生物多样性监测的准确性。

Comments Add Variations of DNA encoding

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20854 2025-11-27 cs.CV cs.AI cs.CL 67%

Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries

从舌尖上的遗忘检索查询中无监督建模可记忆性

Sree Bhattacharyya, Yaman Kumar Singla, Sudhir Yarram, Somesh Kumar Singh, Harini S, James Z. Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Adobe Media and Data Science Research(Adobe媒体与数据科学研究)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出首个大规模无监督数据集,通过舌尖上的遗忘检索查询建模视觉内容的可记忆性,并展示了其在回忆生成和ToT检索任务中的优越表现。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14229 2025-11-19 cs.LG 67%

EBind: a practical approach to space binding

Jim Broadbent, Felix Cohen, Frederik Hvilshøj, Eric Landau, Eren Sasoglu

机构 * Encord London, UK(Encord伦敦)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17415 2025-10-21 cs.CL cs.AI cs.MA cs.MM cs.SE 67%

BenCao: An Instruction-Tuned Large Language Model for Traditional Chinese Medicine

Jiacheng Xie, Yang Yu, Yibo Chen, Hanyao Zhang, Lening Zhao, Jiaxuan He, Lei Jiang, Xiaoting Tang, Guanghui An, Dong Xu

机构 * Community Health Service Center Shanghai Pudong New Area(上海浦东新区社区卫生服务中心) School of Acupuncture-Moxibustion and Tuina, Shanghai University of Traditional Chinese Medicine(上海中医药大学针灸推拿学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03309 2025-10-07 cs.LG q-bio.BM 67%

Thin Bridges for Drug Text Alignment: Lightweight Contrastive Learning for Target Specific Drug Retrieval

Mallikarjuna Tupakula

机构 * Rochester Institute of Technology(罗切斯特技术研究所)

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19054 2025-07-28 cs.CV cs.AI cs.CL cs.IR cs.LG 67%

Closing the Modality Gap for Mixed Modality Search

Binxu Li, Yuhui Zhang, Xiaohan Wang, Weixin Liang, Ludwig Schmidt, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project page: https://yuhui-zh15.github.io/MixedModalitySearch/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16216 2025-06-23 cs.LG 67%

From Pixels to CSI: Distilling Latent Dynamics For Efficient Wireless Resource Management

Charbel Bou Chaaya, Abanoub M. Girgis, Mehdi Bennis

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18017 2025-06-04 cs.CV cs.AI cs.CL cs.IR 67%

ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Qiuchen Wang, Ruixue Ding, Zehui Chen, Weiqi Wu, Shihang Wang, Pengjun Xie, Feng Zhao

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, USTC(脑启发智能感知与认知联合实验室,中国科学技术大学) Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) Shanghai Jiao Tong University(上海交通大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20361 2025-05-27 cs.CR 67%

From ML to LLM: Evaluating the Robustness of Phishing Webpage Detection Models against Adversarial Attacks

Aditya Kulkarni, Vivek Balachandran, Dinil Mon Divakaran, Tamal Das

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19900 2025-03-26 cs.CV cs.AI cs.CL 67%

CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

Hao Yu, Zhuokai Zhao, Shen Yan, Lukasz Korycki, Jianyu Wang, Baosheng He, Jiayi Liu, Lizhu Zhang, Xiangjun Fan, Hanchao Yu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18495 2025-03-05 cs.MM cs.AI cs.CV cs.IR 67%

A Comprehensive Survey on Composed Image Retrieval

Xuemeng Song, Haoqiang Lin, Haokun Wen, Bohan Hou, Mingzhu Xu, Liqiang Nie

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07546 2025-02-18 cs.CV cs.AI cs.CL 67%

Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection

YeongHyeon Park, Myung Jin Kim, Hyeong Seok Kim

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 4 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09059 2025-01-20 cs.CL cs.AI cs.CV 67%

Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval

Delong Liu, Haiwen Li, Zhicheng Zhao, Yuan Dong

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The paper was withdrawn due to a dispute among the authors regarding the content of the article

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03513 2024-12-10 cs.AI cs.CL cs.CV cs.LG 67%

Enhancing CLIP Conceptual Embedding through Knowledge Distillation

Kuei-Chun Kao

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10821 2024-11-19 cs.LG q-bio.BM 67%

GeomCLIP: Contrastive Geometry-Text Pre-training for Molecules

Teng Xiao, Chao Cui, Huaisheng Zhu, Vasant G. Honavar

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract)

Comments BIBM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏