arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2305.04961 2023-05-10 cs.CV cs.CL cs.LG 62%

Joint Moment Retrieval and Highlight Detection Via Natural Language Queries

Richard Luo, Austin Peng, Heidi Yap, Koby Beard

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12445 2023-05-02 cs.CV cs.AI 62%

MEDIMP: 3D Medical Images with clinical Prompts from limited tabular data for renal transplantation

Leo Milecki, Vicky Kalogeiton, Sylvain Bodard, Dany Anglicheau, Jean-Michel Correas, Marc-Olivier Timsit, Maria Vakalopoulou

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13649 2023-04-27 cs.CV cs.CL cs.IR 62%

A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question Answering

Alireza Salemi, Juan Altmayer Pizzorno, Hamed Zamani

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05725 2023-04-13 cs.CV cs.AI 62%

CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational Alignment

Jiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li, Ge Wang, Jun Xia, Yidong Chen, Stan Z. Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2023 (Highlight paper, 2.5% acceptance rate); Open source

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04979 2023-03-17 cs.CV cs.LG cs.MM 62%

VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Shen Yan, Tao Zhu, Zirui Wang, Yuan Cao, Mi Zhang, Soham Ghosh, Yonghui Wu, Jiahui Yu

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.MM

Comments Tech report. arXiv v3: update text

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08039 2023-03-15 cs.CL cs.CV 62%

TQ-Net: Mixed Contrastive Representation Learning For Heterogeneous Test Questions

He Zhu, Xihua Li, Xuemin Zhao, Yunbo Cao, Shan Yu

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments This paper has been accepted for the AAAI2023 AI4Edu Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05093 2023-02-21 cs.IR cs.AI cs.MM 62%

Unified Vision-Language Representation Modeling for E-Commerce Same-Style Products Retrieval

Ben Chen, Linbo Jin, Xinxin Wang, Dehong Gao, Wen Jiang, Wei Ning

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted in The Web Conference (WWW2023) Industry Track. 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12886 2023-02-21 cs.CV cs.AI cs.IR 62%

You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in Videos

Xin Sun, Xuan Wang, Jialin Gao, Qiong Liu, Xi Zhou

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments in SIGIR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05703 2023-02-14 cs.CL cs.AI 62%

HateProof: Are Hateful Meme Detection Systems really Robust?

Piush Aggarwal, Pranit Chawla, Mithun Das, Punyajoy Saha, Binny Mathew, Torsten Zesch, Animesh Mukherjee

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted at TheWebConf'2023 (WWW'2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04858 2023-02-14 cs.CV cs.MM 62%

LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval

Jinbin Bai, Chunhui Liu, Feiyue Ni, Haofan Wang, Mengying Hu, Xiaofeng Guo, Lele Cheng

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.11471 2022-12-23 cs.CV cs.MM 62%

Multi-queue Momentum Contrast for Microvideo-Product Retrieval

Yali Du, Yinwei Wei, Wei Ji, Fan Liu, Xin Luo, Liqiang Nie

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining (WSDM '23), February 27-March 3, 2023, Singapore, Singapore

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15867 2022-11-21 cs.CV cs.CL 62%

Image Retrieval from Contextual Descriptions

Benno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Ponti, Siva Reddy

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments accepted to ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01036 2022-10-28 cs.LG cs.AI cs.CV 62%

Face-to-Face Contrastive Learning for Social Intelligence Question-Answering

Alex Wilf, Martin Q. Ma, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08554 2022-10-18 cs.CV cs.CL 62%

COFAR: Commonsense and Factual Reasoning in Image Search

Prajwal Gatti, Abhirama Subramanyam Penamakuri, Revant Teotia, Anand Mishra, Shubhashis Sengupta, Roshni Ramnani

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted in AACL-IJCNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02833 2022-10-07 cs.IR cs.CL cs.LG cs.SD eess.AS 62%

Matching Text and Audio Embeddings: Exploring Transfer-learning Strategies for Language-based Audio Retrieval

Benno Weck, Miguel Pérez Fernández, Holger Kirchhoff, Xavier Serra

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments 5 pages, 2 figures. Accepted at Detection and Classification of Acoustic Scenes and Events 2022 (DCASE2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07424 2022-09-16 cs.CL cs.LG cs.SD eess.AS 62%

CMSBERT-CLR: Context-driven Modality Shifting BERT with Contrastive Learning for linguistic, visual, acoustic Representations

Junghun Kim, Jihie Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted by IJCNN 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05773 2022-09-14 cs.CV cs.IR cs.MM 62%

CAIBC: Capturing All-round Information Beyond Color for Text-based Person Retrieval

Zijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan, Chao Liu, Tian Wang, Yifeng Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted on ACM MM '22

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12661 2022-07-27 cs.CV cs.CL 62%

Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training

Haoxuan You, Luowei Zhou, Bin Xiao, Noel Codella, Yu Cheng, Ruochen Xu, Shih-Fu Chang, Lu Yuan

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ECCV 2022, 22 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10830 2022-06-23 cs.CV cs.AI 62%

A Feature Memory Rearrangement Network for Visual Inspection of Textured Surface Defects Toward Edge Intelligent Manufacturing

Haiming Yao, Wenyong Yu, Xue Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Revision to IEEE transactions on automation science and engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07991 2022-06-23 cs.CV cs.CL cs.LG 62%

LiT: Zero-Shot Transfer with Locked-image text Tuning

Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, Lucas Beyer

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL

Comments Xiaohua, Xiao, Basil, Andreas and Lucas contributed equally; CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10435 2022-04-25 cs.CV cs.AI 62%

PreTraM: Self-Supervised Pre-training via Connecting Trajectory and Map

Chenfeng Xu, Tian Li, Chen Tang, Lingfeng Sun, Kurt Keutzer, Masayoshi Tomizuka, Alireza Fathi, Wei Zhan

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments The code is available at https://github.com/chenfengxu714/PreTraM.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01971 2022-04-07 cs.CV cs.AI 62%

Non-Local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimation

Jogendra Nath Kundu, Siddharth Seth, Anirudh Jamkhandi, Pradyumna YM, Varun Jampani, Anirban Chakraborty, R. Venkatesh Babu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2021. Project page: https://sites.google.com/view/sa3dhp

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13489 2022-04-05 cs.CV cs.AI cs.LG cs.RO 62%

SurfEmb: Dense and Continuous Correspondence Distributions for Object Pose Estimation with Learnt Surface Embeddings

Rasmus Laurvig Haugaard, Anders Glent Buch

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10299 2022-03-22 cs.CL cs.AI 62%

Neural Machine Translation with Phrase-Level Universal Visual Representations

Qingkai Fang, Yang Feng

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments ACL 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05299 2022-01-17 cs.CV cs.CL cs.IR 62%

A Thousand Words Are Worth More Than a Picture: Natural Language-Centric Outside-Knowledge Visual Question Answering

Feng Gao, Qing Ping, Govind Thattai, Aishwarya Reganti, Ying Nian Wu, Prem Natarajan

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.03651 2021-11-08 cs.CV cs.CL 62%

The Curious Layperson: Fine-Grained Image Recognition without Expert Labels

Subhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea Vedaldi

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments To appear in BMVC 2021 (Oral). Project page: https://www.robots.ox.ac.uk/~vgg/research/clever/

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11881 2021-10-25 cs.CV cs.CL 62%

Simple Dialogue System with AUDITED

Yusuf Tas, Piotr Koniusz

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by the BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04435 2021-10-12 cs.CV cs.CL 62%

Two-stage Visual Cues Enhancement Network for Referring Image Segmentation

Yang Jiao, Zequn Jie, Weixin Luo, Jingjing Chen, Yu-Gang Jiang, Xiaolin Wei, Lin Ma

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ACM MM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.10016 2021-09-22 cs.MM cs.AI 62%

CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval

Zhijian Hou, Chong-Wah Ngo, Wing Kwong Chan

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI、cs.MM

Comments 10 pages, 4 figures, 2021 MultiMedia, code: https://github.com/houzhijian/CONQUER

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04378 2021-09-10 cs.RO cs.AI cs.CV cs.LG 62%

Dynamic Modeling of Hand-Object Interactions via Tactile Sensing

Qiang Zhang, Yunzhu Li, Yiyue Luo, Wan Shou, Michael Foshey, Junchi Yan, Joshua B. Tenenbaum, Wojciech Matusik, Antonio Torralba

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments IROS 2021. First two authors contributed equally. Project page: http://phystouch.csail.mit.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏