arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2401.13613 2024-01-25 cs.CV cs.AI 62%

Enhancing Image Retrieval : A Comprehensive Study on Photo Search using the CLIP Mode

Naresh Kumar Lahajal, Harini S

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02979 2024-01-09 cs.CL cs.AI cs.IR 62%

Are we describing the same sound? An analysis of word embedding spaces of expressive piano performance

Silvan David Peter, Shreyan Chowdhury, Carlos Eduardo Cancino-Chacón, Gerhard Widmer

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、cs.AI

Journal ref Proceedings of the Forum for Information Retrieval Evaluation, FIRE, 2023, Panjim, India

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02309 2024-01-08 cs.CV cs.MM 62%

TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection

Hao Sun, Mingyao Zhou, Wenjing Chen, Wei Xie

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by AAAI-24

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06854 2023-12-21 cs.CV cs.CL cs.CR cs.LG 62%

Robust Contrastive Language-Image Pre-training against Data Poisoning and Backdoor Attacks

Wenhan Yang, Jingdong Gao, Baharan Mirzasoleiman

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18890 2023-12-19 cs.CV cs.AI 62%

Towards Generalized Multi-stage Clustering: Multi-view Self-distillation

Jiatai Wang, Zhiwei Xu, Xin Wang, Tao Li

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06497 2023-12-15 cs.CV cs.AI 62%

DRUformer: Enhancing the driving scene Important object detection with driving relationship self-understanding

Yingjie Niu, Ming Ding, Keisuke Fujii, Kento Ohtani, Alexander Carballo, Kazuya Takeda

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18273 2023-12-01 cs.CV cs.MM 62%

HKUST at SemEval-2023 Task 1: Visual Word Sense Disambiguation with Context Augmentation and Visual Assistance

Zhuohao Yin, Xin Huang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16464 2023-11-29 cs.CV cs.AI 62%

Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

Yicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma, Hengwei Bian, Yatai Ji, Yujiu Yang, Xiu Li

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13752 2023-11-27 cs.CV cs.AI 62%

3D-MIR: A Benchmark and Empirical Study on 3D Medical Image Retrieval in Radiology

Asma Ben Abacha, Alberto Santamaria-Pang, Ho Hin Lee, Jameson Merkow, Qin Cai, Surya Teja Devarakonda, Abdullah Islam, Julia Gong, Matthew P. Lungren, Thomas Lin, Noel C Codella, Ivan Tarapov

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10859 2023-11-15 cs.LG cs.AI cs.CV stat.ML 62%

Visualizing the Diversity of Representations Learned by Bayesian Neural Networks

Dennis Grinwald, Kirill Bykov, Shinichi Nakajima, Marina M. -C. Höhne

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 16 pages, 18 figures

Journal ref Published in Transactions on Machine Learning Research (11/2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00282 2023-11-13 cs.CV cs.IR cs.MM 62%

(Un)likelihood Training for Interpretable Embedding

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Zhijian Hou

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments accepted in ACM Transactions on Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16999 2023-10-31 cs.CV cs.AI cs.LG 62%

Three Towers: Flexible Contrastive Learning with Pretrained Image Models

Jannik Kossen, Mark Collier, Basil Mustafa, Xiao Wang, Xiaohua Zhai, Lucas Beyer, Andreas Steiner, Jesse Berent, Rodolphe Jenatton, Efi Kokiopoulou

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07091 2023-10-20 cs.CL cs.AI 62%

Jaeger: A Concatenation-Based Multi-Transformer VQA Model

Jieting Long, Zewei Shi, Penghao Jiang, Yidong Gan

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This paper is the technical research paper of CIKM 2023 DocIU challenges. The authors received the CIKM 2023 DocIU Winner Award, sponsored by Google, Microsoft, and the Centre for data-driven geoscience

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05473 2023-10-10 cs.CV cs.AI 62%

Sentence-level Prompts Benefit Composed Image Retrieval

Yang Bai, Xinxing Xu, Yong Liu, Salman Khan, Fahad Khan, Wangmeng Zuo, Rick Siow Mong Goh, Chun-Mei Feng

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18274 2023-10-10 cs.CV cs.AI q-bio.NC 62%

Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors

Paul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin, Alex Nguyen, Ethan Cohen, Aidan J. Dempster, Nathalie Verlinde, Elad Yundler, David Weisberg, Kenneth A. Norman, Tanishq Mathew Abraham

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Project Page at https://medarc.ai/mindeye. Code at https://github.com/MedARC-AI/fMRI-reconstruction-NSD/. Published as a conference paper at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08839 2023-09-19 cs.SD cs.MM eess.AS 62%

Contrastive Latent Space Reconstruction Learning for Audio-Text Retrieval

Kaiyi Luo, Xulong Zhang, Jianzong Wang, Huaxiong Li, Ning Cheng, Jing Xiao

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted by The 35th IEEE International Conference on Tools with Artificial Intelligence. (ICTAI 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00917 2023-09-15 cs.CL cs.AI 62%

Knowledge Graph Embeddings for Multi-Lingual Structured Representations of Radiology Reports

Tom van Sonsbeek, Xiantong Zhen, Marcel Worring

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03100 2023-09-07 cs.CV cs.MM 62%

FArMARe: a Furniture-Aware Multi-task methodology for Recommending Apartments based on the user interests

Ali Abdari, Alex Falcon, Giuseppe Serra

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments accepted for presentation at the ICCV2023 CV4Metaverse workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01740 2023-09-07 eess.IV cs.CL cs.CV cs.LG 62%

An Empirical Analysis for Zero-Shot Multi-Label Classification on COVID-19 CT Scans and Uncurated Reports

Ethan Dack, Lorenzo Brigato, Matthew McMurray, Matthias Fontanellaz, Thomas Frauenfelder, Hanno Hoppe, Aristomenis Exadaktylos, Thomas Geiser, Manuela Funke-Chambour, Andreas Christe, Lukas Ebner, Stavroula Mougiakakou

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00775 2023-09-06 cs.CV cs.AI cs.LG 62%

Contrastive Feature Masking Open-Vocabulary Vision Transformer

Dahun Kim, Anelia Angelova, Weicheng Kuo

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00976 2023-08-28 cs.CV cs.CL 62%

TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion Synthesis

Mathis Petrovich, Michael J. Black, Gül Varol

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments ICCV 2023 Camera Ready, project page: https://mathis.petrovich.fr/tmr/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03881 2023-08-23 cs.IR cs.CL cs.CV 62%

Fairness in Image Search: A Study of Occupational Stereotyping in Image Retrieval and its Debiasing

Swagatika Dash

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments 20 Pages, Work uses Proprietary Search Systems from the year 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02898 2023-08-15 cs.CV cs.MM 62%

Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search Benchmark

Shuyu Yang, Yinan Zhou, Yaxiong Wang, Yujiao Wu, Li Zhu, Zhedong Zheng

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05144 2023-08-10 cs.CV cs.AI 62%

Adapt and Align to Improve Zero-Shot Sketch-Based Image Retrieval

Shiyin Dong, Mingrui Zhu, Nannan Wang, Xinbo Gao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16226 2023-08-01 cs.CV cs.MM 62%

ScribbleVC: Scribble-supervised Medical Image Segmentation with Vision-Class Embedding

Zihan Li, Yuan Zheng, Xiangde Luo, Dandan Shan, Qingqi Hong

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2023, project page: https://github.com/HUANGLIZI/ScribbleVC

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08541 2023-07-28 cs.CV cs.AI 62%

Fine-Tuned but Zero-Shot 3D Shape Sketch View Similarity and Retrieval

Gianluca Berardi, Yulia Gryaditskaya

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05610 2023-07-13 cs.LG cs.AI cs.CV 62%

Substance or Style: What Does Your Image Embedding Know?

Cyrus Rashtchian, Charles Herrmann, Chun-Sung Ferng, Ayan Chakrabarti, Dilip Krishnan, Deqing Sun, Da-Cheng Juan, Andrew Tomkins

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

Comments 27 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06870 2023-06-13 cs.CV cs.AI 62%

Sticker820K: Empowering Interactive Retrieval with Stickers

Sijie Zhao, Yixiao Ge, Zhongang Qi, Lin Song, Xiaohan Ding, Zehua Xie, Ying Shan

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17652 2023-05-30 cs.CV cs.CL 62%

ConaCLIP: Exploring Distillation of Fully-Connected Knowledge Interaction Graph for Lightweight Text-Image Retrieval

Jiapeng Wang, Chengyu Wang, Xiaodan Wang, Jun Huang, Lianwen Jin

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments ACL 2023 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11327 2023-05-22 cs.CV cs.LG cs.MM 62%

MALM: Mask Augmentation based Local Matching for Food-Recipe Retrieval

Bhanu Prakash Voutharoja, Peng Wang, Lei Wang, Vivienne Guan

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.MM

Comments Under review. Link to the dataset repo - https://github.com/torralba-lab/im2recipe-Pytorch#recipe1m-dataset

详情

展开后加载摘要…

URL PDF HTML 收藏