arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2501.14369 2025-01-27 cs.CV 70%

Low-rank Prompt Interaction for Continual Vision-Language Retrieval

Weicai Yan, Ye Wang, Wang Lin, Zirun Guo, Zhou Zhao, Tao Jin

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13277 2025-01-24 cs.CV 70%

MEDFORM: A Foundation Model for Contrastive Learning of CT Imaging and Clinical Numeric Data in Multi-Cancer Analysis

Daeun Jung, Jaehyeok Jang, Sooyoung Jang, Yu Rang Park

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13372 2024-12-19 cs.CV 70%

Restore Anything Model via Efficient Degradation Adaptation

Bin Ren, Eduard Zamfir, Zongwei Wu, Yawei Li, Yidi Li, Danda Pani Paudel, Radu Timofte, Ming-Hsuan Yang, Nicu Sebe

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Efficient Any Image Restoration

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11793 2024-12-16 cs.AI 70%

Leveraging Chemistry Foundation Models to Facilitate Structure Focused Retrieval Augmented Generation in Multi-Agent Workflows for Catalyst and Materials Design

Nathaniel H. Park, Tiffany J. Callahan, James L. Hedrick, Tim Erdmann, Sara Capponi

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08715 2024-12-12 cs.CV 70%

Retrieval Augmented Recipe Generation

Guoshan Liu, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments ACCEPT on IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05756 2024-12-10 cs.CV 70%

Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Wenliang Zhong, Weizhi An, Feng Jiang, Hehuan Ma, Yuzhi Guo, Junzhou Huang

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00436 2024-10-02 cs.RO cs.CV 70%

Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations

Miyu Goko, Motonari Kambara, Daichi Saito, Seitaro Otsuki, Komei Sugiura

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted for presentation at CoRL2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19439 2024-10-01 cs.CV 70%

Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery

Andy V. Huynh, Lauren E. Gillespie, Jael Lopez-Saucedo, Claire Tang, Rohan Sikand, Moisés Expósito-Alonso

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17779 2024-09-30 cs.CV 70%

DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction

Chaofan Gan, Yuanpeng Tu, Yuxi Li, Weiyao Lin

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15810 2024-09-25 cs.CV 70%

Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification

Naiwen Hu, Haozhe Cheng, Yifan Xie, Pengcheng Shi, Jihua Zhu

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at IROS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11795 2024-09-11 cs.CV 70%

PoseScript: Linking 3D Human Poses and Natural Language

Ginger Delmas, Philippe Weinzaepfel, Thomas Lucas, Francesc Moreno-Noguer, Grégory Rogez

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments TPAMI 2024, extended version of the ECCV 2022 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18520 2024-08-30 cs.CV 70%

Text-Region Matching for Multi-Label Image Recognition with Missing Labels

Leilei Ma, Hongxing Xie, Lei Wang, Yanping Fu, Dengdi Sun, Haifeng Zhao

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to ACM International Conference on Multimedia (ACM MM) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09527 2024-08-23 cs.AI 70%

ALS-HAR: Harnessing Wearable Ambient Light Sensors to Enhance IMU-based Human Activity Recogntion

Lala Shakti Swarup Ray, Daniel Geißler, Mengxi Liu, Bo Zhou, Sungho Suh, Paul Lukowicz

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16264 2024-07-24 cs.CV 70%

Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring

Shreyank N Gowda, David A. Clifton

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted in MICCAI-24

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12710 2024-07-17 cs.CV 70%

Text-Video Retrieval with Global-Local Semantic Consistent Learning

Haonan Zhang, Pengpeng Zeng, Lianli Gao, Jingkuan Song, Yihang Duan, Xinyu Lyu, Hengtao Shen

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments The author has withdrawn this paper due to a critical definitional error in concept learning for global/local-interaction learning during training. This error led to an alignment issue with the definition of the text-video retrieval task, causing an unfair comparison with state-of-the-art (SOTA) methods. Consequently, this hindered the accurate evaluation of the paper's contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04255 2024-07-08 cs.CV 70%

Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge

Xiangyu Wu, Zhouyang Chi, Yang Yang, Jianfeng Lu

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Second Place of WSDM2023 Toloka Visual Question Answering Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07450 2024-06-12 cs.CV cs.LG 70%

Benchmarking Vision-Language Contrastive Methods for Medical Representation Learning

Shuvendu Roy, Yasaman Parhizkar, Franklin Ogidi, Vahid Reza Khazaie, Michael Colacci, Ali Etemad, Elham Dolatabadi, Arash Afkanpour

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18167 2024-05-29 eess.IV cs.CV 70%

Confidence-aware multi-modality learning for eye disease screening

Ke Zou, Tian Lin, Zongbo Han, Meng Wang, Xuedong Yuan, Haoyu Chen, Changqing Zhang, Xiaojing Shen, Huazhu Fu

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments 27 pages, 7 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18339 2024-03-29 eess.IV cs.CV 70%

H2ASeg: Hierarchical Adaptive Interaction and Weighting Network for Tumor Segmentation in PET/CT Images

Jinpeng Lu, Jingyun Chen, Linghan Cai, Songhan Jiang, Yongbing Zhang

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02227 2024-03-18 cs.LG cs.AI 70%

SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training

Kazem Meidani, Parshin Shojaee, Chandan K. Reddy, Amir Barati Farimani

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

Comments ICLR 2024 Spotlight Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07272 2024-03-07 cs.CV 70%

Zero-shot Composed Text-Image Retrieval

Yikun Liu, Jiangchao Yao, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07196 2024-02-22 cs.CV 70%

Retrieval-Enhanced Contrastive Vision-Text Models

Ahmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia Schmid

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03689 2023-11-06 cs.CV 70%

COLA: A Benchmark for Compositional Text-to-image Retrieval

Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan A. Plummer, Ranjay Krishna, Kate Saenko

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2023. Webpage: https://cs-people.bu.edu/array/research/cola/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00566 2023-11-02 cs.CV 70%

CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders

Anthony Fuller, Koreen Millard, James R. Green

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments NeurIPS 2023 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05268 2023-10-31 cs.LG cs.AI cs.CL cs.CV cs.MM 70%

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Paul Pu Liang, Zihao Deng, Martin Ma, James Zou, Louis-Philippe Morency, Ruslan Salakhutdinov

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2023. Code available at: https://github.com/pliang279/FactorCL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09199 2023-10-19 cs.CV 70%

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Xi Chen, Xiao Wang, Lucas Beyer, Alexander Kolesnikov, Jialin Wu, Paul Voigtlaender, Basil Mustafa, Sebastian Goodman, Ibrahim Alabdulmohsin, Piotr Padlewski, Daniel Salz, Xi Xiong, Daniel Vlasic, Filip Pavetic, Keran Rong, Tianli Yu, Daniel Keysers, Xiaohua Zhai, Radu Soricut

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01358 2023-10-03 cs.CV 70%

NEUCORE: Neural Concept Reasoning for Composed Image Retrieval

Shu Zhao, Huijuan Xu

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12134 2023-09-22 cs.SD cs.IR cs.LG eess.AS 70%

Self-Supervised Contrastive Learning for Robust Audio-Sheet Music Retrieval Systems

Luis Carvalho, Tobias Washüttl, Gerhard Widmer

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

Journal ref Proceedings of the 14th ACM Multimedia Systems Conference (MMSys '23), June 7-10, 2023, Vancouver, BC, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06866 2023-08-15 cs.CV 70%

Improving Face Recognition from Caption Supervision with Multi-Granular Contextual Feature Aggregation

Md Mahedi Hasan, Nasser Nasrabadi

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments This article has been accepted for publication in the IEEE International Joint Conference on Biometrics (IJCB), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07615 2023-05-24 cs.CL 70%

UGIF: UI Grounded Instruction Following

Sagar Gubbi Venkatesh, Partha Talukdar, Srini Narayanan

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏