arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2312.12723 2023-12-21 cs.CV 57%

Multi-Clue Reasoning with Memory Augmentation for Knowledge-based Visual Question Answering

Chengxiang Yin, Zhengping Che, Kun Wu, Zhiyuan Xu, Jian Tang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12659 2023-12-21 cs.CV 57%

Expediting Contrastive Language-Image Pretraining via Self-distilled Encoders

Bumsoo Kim, Jinhyung Kim, Yeonsik Jo, Seung Hwan Kim

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11376 2023-12-20 cs.CV 57%

CLIM: Contrastive Language-Image Mosaic for Region Representation

Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Wentao Liu, Chen Change Loy

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09681 2023-12-18 cs.LG cs.CV cs.DB 57%

Urban Region Embedding via Multi-View Contrastive Prediction

Zechen Li, Weiming Huang, Kai Zhao, Min Yang, Yongshun Gong, Meng Chen

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09511 2023-12-18 cs.IR cs.AI 57%

MONET: Modality-Embracing Graph Convolutional Network and Target-Aware Attention for Multimedia Recommendation

Yungi Kim, Taeri Kim, Won-Yong Shin, Sang-Wook Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Accepted by WSDM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06141 2023-12-14 cs.AI 57%

Survey on Memory-Augmented Neural Networks: Cognitive Insights to AI Applications

Savya Khosla, Zhen Zhu, Yifei He

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06055 2023-12-12 cs.SD eess.AS 57%

Speaker-Text Retrieval via Contrastive Learning

Xuechen Liu, Xin Wang, Erica Cooper, Xiaoxiao Miao, Junichi Yamagishi

专题命中 跨模态检索 :cross-modal(abstract);分类 eess.AS

Comments Submitted to IEEE Signal Processing Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16676 2023-12-12 cs.CL 57%

SSLCL: An Efficient Model-Agnostic Supervised Contrastive Learning Framework for Emotion Recognition in Conversations

Tao Shi, Xiao Liang, Yaoyuan Liang, Xinyi Tong, Shao-Lun Huang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11435 2023-12-12 cs.CV cs.LG 57%

Bidirectional Contrastive Split Learning for Visual Question Answering

Yuwei Sun, Hideya Ochiai

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted for AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02780 2023-12-06 cs.LG cs.CL cs.CR 57%

Scaling Laws for Adversarial Attacks on Language Model Activations

Stanislav Fort

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17949 2023-12-01 cs.CV 57%

Zero-shot Retrieval: Augmenting Pre-trained Models with Search Engines

Hamed Damirchi, Cristian Rodríguez-Opazo, Ehsan Abbasnejad, Damien Teney, Javen Qinfeng Shi, Stephen Gould, Anton van den Hengel

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00349 2023-11-28 cs.CV cs.LG 57%

CALICO: Self-Supervised Camera-LiDAR Contrastive Pre-training for BEV Perception

Jiachen Sun, Haizhong Zheng, Qingzhao Zhang, Atul Prakash, Z. Morley Mao, Chaowei Xiao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09084 2023-11-16 cs.CV 57%

Contrastive Transformer Learning with Proximity Data Generation for Text-Based Person Search

Hefeng Wu, Weifeng Chen, Zhibin Liu, Tianshui Chen, Zhiguang Chen, Liang Lin

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE T-CSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07861 2023-11-16 cs.IR cs.AI 57%

Overview of the TREC 2023 Product Product Search Track

Daniel Campos, Surya Kallumadi, Corby Rosset, Cheng Xiang Zhai, Alessandro Magnani

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments 14 pages, 4 figures, 11 tables - TREC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16604 2023-11-07 cs.CV cs.IR cs.LG 57%

Bi-directional Training for Composed Image Retrieval via Text Prompt Learning

Zheyuan Liu, Weixuan Sun, Yicong Hong, Damien Teney, Stephen Gould

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments WACV 2024 accepted. 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19168 2023-10-31 cs.CV 57%

BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and Mapping

Srikumar Sastry, Subash Khanal, Aayush Dhakal, Di Huang, Nathan Jacobs

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08937 2023-10-27 cs.CL cs.IR 57%

DocumentNet: Bridging the Data Gap in Document Pre-Training

Lijun Yu, Jin Miao, Xiaoyu Sun, Jiayi Chen, Alexander G. Hauptmann, Hanjun Dai, Wei Wei

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.07115 2023-10-24 cs.LG cs.CL cs.SE stat.ML 57%

HiGitClass: Keyword-Driven Hierarchical Classification of GitHub Repositories

Yu Zhang, Frank F. Xu, Sha Li, Yu Meng, Xuan Wang, Qi Li, Jiawei Han

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments 10 pages; Accepted to ICDM 2019; Some typos fixed

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06342 2023-10-11 cs.SE cs.AI 57%

Contrastive Prompt Learning-based Code Search based on Interaction Matrix

Yubo Zhang, Yanfang Liu, Xinxin Fan, Yunfeng Lu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01842 2023-10-04 cs.CV 57%

SelfGraphVQA: A Self-Supervised Graph Neural Network for Scene-based Question Answering

Bruno Souza, Marius Aasan, Helio Pedrini, Adín Ramírez Rivera

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments To appear in Vision-and-Language Algorithmic Reasoning Workshop at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01330 2023-10-03 cs.CV 57%

Towards reporting bias in visual-language datasets: bimodal augmentation by decoupling object-attribute association

Qiyu Wu, Mengjie Zhao, Yutong He, Lang Huang, Junya Ono, Hiromi Wakaki, Yuki Mitsufuji

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14673 2023-09-27 cs.LG cs.AI cs.IR cs.SI 57%

ALEX: Towards Effective Graph Transfer Learning with Noisy Labels

Jingyang Yuan, Xiao Luo, Yifang Qin, Zhengyang Mao, Wei Ju, Ming Zhang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments Accepted by the ACM International Conference on Multimedia (MM) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14007 2023-09-26 cs.CV 57%

Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training

Ziyu Guo, Renrui Zhang, Longtian Qiu, Xianzhi Li, Pheng-Ann Heng

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06552 2023-09-21 cs.CL 57%

MT4CrossOIE: Multi-stage Tuning for Cross-lingual Open Information Extraction

Tongliang Li, Zixiang Wang, Linzheng Chai, Jian Yang, Jiaqi Bai, Yuwei Yin, Jiaheng Liu, Hongcheng Guo, Liqun Yang, Hebboul Zine el-abidine, Zhoujun Li

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08928 2023-09-19 cs.CV 57%

In-Style: Bridging Text and Uncurated Videos with Style Transfer for Text-Video Retrieval

Nina Shvetsova, Anna Kukleva, Bernt Schiele, Hilde Kuehne

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Published at ICCV 2023, code: https://github.com/ninatu/in_style

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08760 2023-09-19 cs.CV 57%

Biased Attention: Do Vision Transformers Amplify Gender Bias More than Convolutional Neural Networks?

Abhishek Mandal, Susan Leavy, Suzanne Little

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08743 2023-09-19 cs.CV 57%

Active Learning for Fine-Grained Sketch-Based Image Retrieval

Himanshu Thakur, Soumitri Chattopadhyay

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted at BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05551 2023-09-12 cs.CV 57%

OpenFashionCLIP: Vision-and-Language Contrastive Learning with Open-Source Fashion Data

Giuseppe Cartella, Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini, Rita Cucchiara

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments International Conference on Image Analysis and Processing (ICIAP) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05503 2023-09-12 cs.CL 57%

Long-Range Transformer Architectures for Document Understanding

Thibault Douzon, Stefan Duffner, Christophe Garcia, Jérémy Espinas

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments Conference: ICDAR 2023 Workshops on Document Analysis and Recognition

Journal ref Document Analysis and Recognition ICDAR 2023 Workshops pages 47 to 64

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01366 2023-09-06 cs.MM 57%

Target-Guided Composed Image Retrieval

Haokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei, Liqiang Nie

专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM

Journal ref ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏