arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2601.16582 2026-01-26 cs.CV 57%

X-Aligner: Composed Visual Retrieval without the Bells and Whistles

X-Aligner:无需炫技的组合视觉检索

Yuqian Zheng, Mariana-Iuliana Georgescu

机构 * Technical University of Munich(慕尼黑技术大学) Helmholtz Munich(亥姆霍兹慕尼黑)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 X-Aligner通过结合视觉和文本查询,提升视频检索性能,采用多阶段训练和交叉注意力模块实现更优的多模态表示对齐。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00243 2026-01-26 cs.CV 57%

VOCAL: Visual Odometry via ContrAstive Learning

基于对比学习的视觉里程计:VOCAL

Chi-Yao Huang, Zeel Bhatt, Yezhou Yang

机构 * Arizona State University(亚利桑那州立大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 VOCAL通过对比学习将视觉里程计重新定义为标签排序问题,提升可解释性和灵活性,推动更通用的空间智能发展。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11393 2026-01-23 cs.CV 57%

Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic Learning

异质不确定性引导的复合图像检索与细粒度概率学习

Haomiao Tang, Jinpeng Wang, Minyi Zhao, Guanghao Meng, Ruisheng Luo, Long Chen, Shu-Tao Xia

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出异质不确定性引导范式,通过细粒度概率学习提升复合图像检索的鲁棒性和准确性。

Comments Accepted for publication and oral presentation at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14121 2026-01-21 cs.CL 57%

NewsRECON: News article REtrieval for image CONtextualization

NewsRECON: 用于图像上下文化的新闻文章检索

Jonathan Tonglet, Iryna Gurevych, Tinne Tuytelaars, Marie-Francine Moens

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(通用知识处理实验室(UKP实验室)、计算机科学系、图恩大学(TU Darmstadt)和应用网络安全国家研究中心ATHENE) Department of Electrical Engineering, KU Leuven(电气工程系、鲁汶大学) Department of Computer Science, KU Leuven(计算机科学系、鲁汶大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 NewsRECON通过链接图像与相关新闻文章,利用元数据推断图像日期和位置,优于现有方法并结合多模态大语言模型实现新的SOTA结果。

Comments Preprint under review. Code available at https://github.com/jtonglet/arxiv2025-newsrecon

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14060 2026-01-21 cs.CV 57%

Fine-Grained Zero-Shot Composed Image Retrieval with Complementary Visual-Semantic Integration

细粒度零样本组合图像检索与互补视觉-语义整合

Yongcong Ye, Kai Zhang, Yanghai Zhang, Enhong Chen, Longfei Li, Jun Zhou

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学) Zhejiang University(浙江大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出CVSI方法,通过互补视觉-语义整合提升细粒度零样本组合图像检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13476 2026-01-21 cs.LG cs.AI 57%

A Unified Variational Imputation Framework for Electric Vehicle Charging Data Using Retrieval-Augmented Language Model

基于检索增强语言模型的统一变分填补框架用于电动汽车充电数据

Jinhao Li, Hao Wang

机构 * Department of Data Science and AI, Faculty of IT and Monash Energy Institute, Monash University(数据科学与人工智能系,信息科技学院和莫纳什能源研究所,莫纳什大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 本文提出PRAIM框架,利用检索增强语言模型和变分神经架构,提升电动汽车充电数据填补的准确性和统计分布保留能力,从而提高预测性能。

Comments 15 pages

Journal ref IEEE Transactions on Smart Grid, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23770 2026-01-21 cs.CV 57%

GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning

GenView++:统一自适应生成增强和质量驱动监督以实现对比表征学习

Xiaojie Li, Bei Wang, Wei Liu, Jianlong Wu, Yue Yu, Liqiang Nie, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

AI总结 GenView++通过自适应生成增强和质量驱动监督,提升对比学习在视觉和视觉-语言任务中的性能。

Comments The code is available at \url{https://github.com/xiaojieli0903/GenViewPlusPlus}

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12010 2026-01-21 cs.CV 57%

SMc2f: Robust Scenario Mining for Robotic Autonomy from Coarse to Fine

SMc2f: 从粗到细的机器人自主性鲁棒场景挖掘

Yifei Chen, Ross Greer

机构 * Department of Computer Science at Xi’an University of Technology(西安理工大学计算机科学系) department of Computer Science & Engineering at the University of California, Merced(加州大学默塞德分校计算机科学与工程系)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

AI总结 SMc2f通过从粗到细的流程,利用视觉语言模型和文本-轨迹对比学习提升机器人场景挖掘的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06743 2026-01-16 cs.MM cs.IR 57%

The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24

生命周期检索的最新进展:ACM生命周期检索挑战工作坊2022-24年进展综述

Allie Tran, Werner Bailer, Duc-Tien Dang-Nguyen, Graham Healy, Steve Hodges, Björn Þór Jónsson, Luca Rossetto, Klaus Schoeffmann, Minh-Triet Tran, Lucia Vadicamo, Cathal Gurrin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM

AI总结 本文综述了ACM生命周期检索挑战工作坊2022-24年在交互式生命周期检索领域的最新进展,重点分析了嵌入式检索方法、大语言模型整合及多模态界面的创新,并探讨了系统性能与用户界面设计的平衡问题。

Journal ref 2025 IEEE Access, 13. pp. 216340-216363. ISSN 2169-3536

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07333 2026-01-13 cs.CV cs.RO 57%

OSCAR: Open-Set CAD Retrieval from a Language Prompt and a Single Image

OSCAR: 从语言提示和单张图像进行开放集CAD检索

Tessa Pulli, Jean-Baptiste Weibel, Peter Hönig, Matthias Hirschmanner, Markus Vincze, Andreas Holzinger

机构 * BOKU University, Human-Centered AI Lab, FTEC, Department for Ecosystem Management, Climate(博乐大学,以人为中心的人工智能实验室,FTEC,生态系统管理部,气候) Institute for Human Centered Computing, Faculty of Informatics(以人为中心的计算研究所,信息学院) Biomedical Engineering, TU Graz, Graz, Austria(生物医学工程,格拉茨技术大学,格拉茨,奥地利)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 OSCAR通过语言提示和单张图像实现开放集CAD检索,无需训练即可在未标记数据库中检索匹配对象模型,提升6D物体姿态估计的自动化水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06496 2026-01-13 cs.CV 57%

3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence

3D CoCa v2:基于测试时间搜索的对比学习用于通用空间智能

Hao Tang, Ting Huang, Zeyu Zhang

机构 * School of Computer Science, Peking University(北京大学计算机学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 3D CoCa v2通过测试时间搜索提升3D场景描述的泛化能力,实现对比学习与描述生成的统一,提高鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05792 2026-01-12 cs.LG cs.AI q-bio.BM 57%

Tensor-DTI: Enhancing Biomolecular Interaction Prediction with Contrastive Embedding Learning

张量-DTI:通过对比嵌入学习增强生物分子相互作用预测

Manel Gil-Sorribes, Júlia Vilalta-Mor, Isaac Filella-Mercè, Robert Soliva, Álvaro Ciudad, Víctor Guallar, Alexis Molina

机构 * Nostrum Biodiscovery Barcelona Supercomputing Center(巴塞罗那超级计算中心) Faculty of Pharmacy and Food Sciences, University of Barcelona(巴塞罗那大学药学与食品科学系) Catalan Institution for Research and Advanced Studies (ICREA)(加泰罗尼亚研究与高级研究机构) Data Science Dpt., Almirall S.A.(Almirall S.A.数据科学部门)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 Tensor-DTI通过整合多模态嵌入与对比学习,提升生物分子相互作用预测的准确性与可靠性。

Comments Accepted at the Generative and Experimental Perspectives for Biomolecular Design Workshop at ICLR 2025 and at the Learning Meaningful Representations of Life Workshop at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07344 2026-01-08 cs.DC cs.AI 57%

Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding

Venus: 一种高效的边缘内存与检索系统用于基于VLM的在线视频理解

Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng, Mu Yuan, Xiaowen Chu, Weijie Hong, Xu Chen

机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China(计算机科学与工程学院,中山大学,广州,中国) Department of Information Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China(信息工程系,香港中文大学,香港特别行政区,中国) Shenzhen Smart City Communications Co., Ltd., China(深圳智慧城市通信有限公司,中国)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 Venus提出了一种高效的边缘内存与检索系统,通过边缘-云解耦架构实现快速的在线视频理解,显著降低系统开销并提升响应速度。

Comments Accepted by IEEE International Conference on Computer Communications 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07227 2026-01-08 cs.CV cs.IR 57%

Which Country Is This? Automatic Country Ranking of Street View Photos

这是哪个国家?街景照片的自动国家排名

Tim Menzner, Jochen L. Leidner, Florian Mittag

机构 * Coburg University of Applied Sciences and Arts(科堡应用科学与艺术大学) University of Sheffield(谢菲尔德大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出Country Guesser系统,通过结合计算机视觉、机器学习和文本检索方法,对街景照片中的地点进行国家排名预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03262 2026-01-08 cs.IR cs.CL 57%

Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey

多模态大语言模型在视觉丰富文档检索中的作用:一项综述

Xiantao Zhang

机构 * Beihang University(北航大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文综述了多模态大语言模型在视觉丰富文档检索中的应用,探讨了三种关键角色及其在RAG中的作用与权衡。

Comments 18 pages; accepted at AACL-IJCNLP 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21121 2026-01-07 cs.IR cs.AI 57%

Beyond Patch Aggregation: 3-Pass Pyramid Indexing for Vision-Enhanced Document Retrieval

超越补丁聚合:面向视觉增强文档检索的三阶段金字塔索引

Anup Roy, Rishabh Gyanendra Upadhyay, Animesh Rameshbhai Panara, Robin Mills, Aidan Millar

机构 * Inception AI Mubadala

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 VisionRAG是一种无OCR、模型无关的多模态检索系统,通过三阶段金字塔索引提升文档检索效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02660 2026-01-05 cs.CV cs.IR 57%

Spatially-Grounded Document Retrieval via Patch-to-Region Relevance Propagation

通过补丁到区域相关性传播实现空间感知的文档检索

Athos Georgiou

机构 * Independent Researcher(独立研究者)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 通过结合ColPali的补丁相似性与OCR区域,实现更精确的文档检索,提升RAG任务中的上下文相关性

Comments 21 pages, 6 figures, 8 tables. Includes ancillary files with full benchmark results and ablation studies. Code available at https://github.com/athrael-soju/Snappy

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09147 2026-01-05 cs.NI cs.AI cs.DC cs.LG cs.MA 57%

Agentic TinyML for Intent-aware Handover in 6G Wireless Networks

代理型TinyML用于6G无线网络中的意图感知切换

Alaa Saleh, Roberto Morabito, Sasu Tarkoma, Anders Lindgren, Susanna Pirttikangas, Lauri Lovén

机构 * Center for Ubiquitous Computing(无处不在计算中心) University of Oulu(奥卢大学) Department of Communication Systems(通信系统系) EURECOM Department of Computer Science(计算机科学系) University of Helsinki(赫尔辛基大学) RISE Research Institutes of Sweden(瑞典研究机构) Department of Computer Science, Electrical and Space Engineering(计算机科学、电气与空间工程系) Luleå University of Technology(卢勒奥技术大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 本文提出WAAN框架,利用TinyML代理实现6G网络中的意图感知主动切换,通过多模态环境控制案例验证其在保持用户体验方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22251 2026-01-01 cs.LG cs.AI 57%

Interpretable Perturbation Modeling Through Biomedical Knowledge Graphs

通过生物医学知识图谱进行可解释的扰动建模

Pascal Passigan, Kevin Zhu, Angelina Ning

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于生物医学知识图谱的可解释扰动建模方法,通过整合多模态嵌入和图注意力网络,提升对药物-细胞对基因表达扰动的预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22780 2025-12-30 cs.CV eess.IV 57%

Plug In, Grade Right: Psychology-Inspired AGIQA

插入选项,正确评分:心理学启发的AGIQA

Zhicheng Liao, Baoliang Chen, Hanwei Zhu, Lingyu Zhu, Shiqi Wang, Weisi Lin

机构 * School of Computer Science, South China Normal University(南方科技大学计算机科学学院) School of Computer Science and Engineering, Nanyang Technological University(南洋理工大学计算机科学与工程学院) School of Computer Science, City University of Hong Kong(香港城市大学计算机科学学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 基于心理学的AGIQA模型通过改进的分级响应模型提升图像质量评分的准确性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05195 2025-12-29 cs.LG cs.AI 57%

HistoKernel: Whole Slide Image Level Maximum Mean Discrepancy Kernels for Pan-Cancer Predictive Modelling

HistoKernel:全切片图像层面的最大均值差异核用于跨癌症预测建模

Piotr Keller, Muhammad Dawood, Brinder Singh Chohan, Fayyaz ul Amir Afsar Minhas

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

AI总结 HistoKernel通过最大均值差异核提升全切片图像层面的跨癌症预测性能,实现多模态数据整合与斑块级可解释性。

Comments 28 pages, 5 figures, 1 Table. Preprint for article in review at Nature Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21683 2025-12-29 cs.CV 57%

Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image Segmentation

对比图模型用于跨域少样本医学图像分割

Yuntian Bo, Tao Zhou, Zechao Li, Haofeng Zhang, Ling Shao

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出对比图模型,通过结构一致性先验和子图匹配解码机制,在跨域少样本医学图像分割中实现高精度与强源域准确性。

Comments Accepted to IEEE Transactions on Medical Imaging (T-MI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20858 2025-12-25 cs.CV 57%

ALIVE: An Avatar-Lecture Interactive Video Engine with Content-Aware Retrieval for Real-Time Interaction

ALIVE: 一种具有内容感知检索的虚拟形象-讲座交互视频引擎,用于实时交互

Md Zabirul Islam, Md Motaleb Hossen Manik, Ge Wang

机构 * Department of Computer Science Rensselaer Polytechnic Institute(计算机科学系罗切斯特理工学院) Department of Biomedical Engineering Rensselaer Polytechnic Institute(生物医学工程系罗切斯特理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 ALIVE通过本地部署和内容感知检索,实现基于虚拟形象的实时互动学习,提升讲座的教育价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14856 2025-12-25 cs.CL 57%

T5Gemma 2: Seeing, Reading, and Understanding Longer

T5Gemma 2: 视觉、阅读与理解更长内容

Biao Zhang, Paul Suganthan, Gaël Liu, Ilya Philippov, Sahil Dua, Ben Hora, Kat Black, Gus Martins, Omar Sanseviero, Shreya Pathak, Cassidy Hardin, Francesco Visin, Jiageng Zhang, Kathleen Kenealy, Qin Yin, Xiaodan Song, Olivier Lacombe, Armand Joulin, Tris Warkentin, Adam Roberts

机构 * Google(谷歌)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 T5Gemma 2通过改进的编码器-解码器架构,实现了多模态和长上下文处理能力,同时在预训练和微调性能上优于Gemma 3。

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20029 2025-12-24 cs.CV 57%

$\text{H}^2$em: Learning Hierarchical Hyperbolic Embeddings for Compositional Zero-Shot Learning

H²em:学习分层双曲嵌入用于组合零样本学习

Lin Li, Jiahui Li, Jiaming Lei, Jun Xiao, Feifei Shao, Long Chen

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 H²em通过双曲几何学习分层嵌入,解决组合零样本学习中的层次结构建模问题,提升细粒度辨别能力和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18853 2025-12-23 cs.CV cs.HC 57%

VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference

VizDefender:通过主动定位和意图推断揭示可视化篡改

Sicheng Song, Yanjie Zhang, Zixin Chen, Huamin Qu, Changbo Wang, Chenhui Li

机构 * East China Normal University(华东师范大学) Hong Kong University of Science and Technology(香港科学与技术大学) School of Computer Science and Technology(计算机科学与技术学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 VizDefender通过半脆弱水印和意图分析模块,主动定位和推断可视化篡改,有效检测和分析数据篡改与视觉编码篡改。

Comments IEEE Transactions on Visualization and Computer Graphics (IEEE PacificVis'26 TVCG Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27261 2025-12-23 cs.CV 57%

RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding

RegionRAG: 基于区域级别的检索增强生成用于视觉文档理解

Yinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun, Chuanbin Liu, Hongtao Xie

机构 * Yinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun, Chuanbin Liu, Hongtao Xie(作者)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 RegionRAG通过区域级别检索增强生成,提升视觉文档理解的效率和准确性,实现更高的检索和问答性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18407 2025-12-23 cs.CV 57%

Through the PRISm: Importance-Aware Scene Graphs for Image Retrieval

通过PRISm:面向图像检索的重要性感知场景图

Dimitrios Georgoulopoulos, Nikolaos Chaidos, Angeliki Dimitriou, Giorgos Stamou

机构 * National Technical University of Athens(国家技术大学雅典)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 PRISm通过重要性感知场景图提升图像检索的语义准确性与人类感知一致性。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17667 2025-12-22 cs.CR cs.AI cs.NI 57%

STAR: Semantic-Traffic Alignment and Retrieval for Zero-Shot HTTPS Website Fingerprinting

STAR:语义-交通对齐与检索用于零样本HTTPS网站指纹识别

Yifei Cheng, Yujia Zhu, Baiyang Li, Xinhao Deng, Yitong Cai, Yaochen Ren, Qingyun Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络与信息安全学院) Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与空间研究院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

AI总结 STAR通过零样本跨模态检索方法,实现对未见过网站的高准确率指纹识别,提升HTTPS加密流量的隐私保护能力。

Comments Accepted by IEEE INFOCOM 2026. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08976 2025-12-19 cs.LG cs.AI 57%

Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs

peek-a-boo推理:多模态大语言模型中的对比区域遮挡

Isha Chaturvedi, Anjana Nair, Yushen Li, Adhitya Rajendra Kumar, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Vasu Sharma

机构 * Algoverse AI Research(Algoverse AI研究机构) Princeton University(普林斯顿大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 对比区域遮挡通过遮挡视觉区域并对比推理轨迹,揭示多模态大语言模型在推理过程中的依赖模式和失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏