arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2303.14376 2023-03-28 cs.CV 57%

ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding

Hongyu Sun, Yongcai Wang, Xudong Cai, Xuewei Bai, Deying Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 9 pages, 4 figures, 7 tables; accepted by ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14348 2023-03-28 cs.CV 57%

Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable Style

Fengyin Lin, Mingkang Li, Da Li, Timothy Hospedales, Yi-Zhe Song, Yonggang Qi

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments CVPR 2023 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10555 2023-03-23 cs.CV cs.LG 57%

LargeKernel3D: Scaling up Kernels in 3D Sparse CNNs

Yukang Chen, Jianhui Liu, Xiangyu Zhang, Xiaojuan Qi, Jiaya Jia

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments In CVPR 2023. Code is at https://github.com/dvlab-research/LargeKernel3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13508 2023-03-21 cs.CV 57%

Motion Transformer with Global Intention Localization and Local Movement Refinement

Shaoshuai Shi, Li Jiang, Dengxin Dai, Bernt Schiele

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2022 as Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09831 2023-03-20 cs.CV 57%

MODIFY: Model-driven Face Stylization without Style Images

Yuhe Ding, Jian Liang, Jie Cao, Aihua Zheng, Ran He

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07775 2023-03-15 cs.CV 57%

Data-Free Sketch-Based Image Retrieval

Abhra Chaudhuri, Ayan Kumar Bhunia, Yi-Zhe Song, Anjan Dutta

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Computer Vision and Pattern Recognition (CVPR) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14785 2023-03-01 cs.CL 57%

Joint Representations of Text and Knowledge Graphs for Retrieval and Evaluation

Teven Le Scao, Claire Gardent

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13173 2023-02-28 cs.CL 57%

MetaAID 2.0: An Extensible Framework for Developing Metaverse Applications via Human-controllable Pre-trained Models

Hongyin Zhu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10178 2023-02-28 cs.CL 57%

Visually-Augmented Language Modeling

Weizhi Wang, Li Dong, Hao Cheng, Haoyu Song, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, Furu Wei

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12552 2023-02-27 cs.CV 57%

Deep Learning for Video-Text Retrieval: a Review

Cunjuan Zhu, Qi Jia, Wei Chen, Yanming Guo, Yu Liu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments International Journal of Multimedia Information Retrieval (IJMIR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.11705 2023-02-24 cs.CV 57%

ACE: Zero-Shot Image to Image Translation via Pretrained Auto-Contrastive-Encoder

Sihan Xu, Zelong Jiang, Ruisi Liu, Kaikai Yang, Zhijie Huang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.11052 2023-02-23 cs.IR cs.AI cs.LG 57%

Que2Engage: Embedding-based Retrieval for Relevant and Engaging Products at Facebook Marketplace

Yunzhong He, Yuxin Tian, Mengjiao Wang, Feier Chen, Licheng Yu, Maolong Tang, Congcong Chen, Ning Zhang, Bin Kuang, Arul Prakash

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Accepted by WWW'2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03293 2023-02-14 cs.SE cs.AI cs.LG 57%

CoCoSoDa: Effective Contrastive Learning for Code Search

Ensheng Shi, Yanlin Wang, Wenchao Gu, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, Hongbin Sun

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Accepted by ICSE 2023 (The 45th International Conference on Software Engineering)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.03990 2023-01-25 cs.CV 57%

TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network

Zhengyi Liu, Yuan Wang, Zhengzheng Tu, Yun Xiao, Bin Tang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08100 2023-01-24 cs.SE cs.AI 57%

CommitBART: A Large Pre-trained Model for GitHub Commits

Shangqing Liu, Yanzhou Li, Xiaofei Xie, Yang Liu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06685 2023-01-18 cs.CV 57%

Distribution Aligned Feature Clustering for Zero-Shot Sketch-Based Image Retrieval

Yuchen Wu, Kun Song, Fangzheng Zhao, Jiansheng Chen, Huimin Ma

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13419 2022-12-29 cs.CV 57%

Position-Aware Contrastive Alignment for Referring Image Segmentation

Bo Chen, Zhiwei Hu, Zhilong Ji, Jinfeng Bai, Wangmeng Zuo

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13558 2022-12-15 cs.CV 57%

FreeSeg: Free Mask from Interpretable Contrastive Language-Image Pretraining for Semantic Segmentation

Yi Li, Huifeng Yao, Hualiang Wang, Xiaomeng Li

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments This paper contains some immature results

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04114 2022-12-09 cs.CV 57%

Group Generalized Mean Pooling for Vision Transformer

Byungsoo Ko, Han-Gyu Kim, Byeongho Heo, Sangdoo Yun, Sanghyuk Chun, Geonmo Gu, Wonjae Kim

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01825 2022-12-06 eess.IV cs.CV 57%

MouseGAN++: Unsupervised Disentanglement and Contrastive Representation for Multiple MRI Modalities Synthesis and Structural Segmentation of Mouse Brain

Ziqi Yu, Xiaoyang Han, Shengjie Zhang, Jianfeng Feng, Tingying Peng, Xiao-Yong Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments IEEE Transactions on Medical Imaging (IEEE-TMI) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16564 2022-12-01 cs.CV cs.LG 57%

Testing GLOM's ability to infer wholes from ambiguous parts

Laura Culp, Sara Sabour, Geoffrey E. Hinton

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07046 2022-11-29 cs.CV 57%

Exploring Visual Interpretability for Contrastive Language-Image Pre-training

Yi Li, Hualiang Wang, Yiqun Duan, Hang Xu, Xiaomeng Li

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11225 2022-11-22 cs.SD cs.LG eess.AS 57%

TimbreCLIP: Connecting Timbre to Text and Images

Nicolas Jonason, Bob L. T. Sturm

专题命中 跨模态检索 :cross-modal(abstract);分类 eess.AS

Comments Submitted to AAAI workshop on creative AI across modalities

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08557 2022-11-17 cs.CV 57%

Unsupervised Feature Clustering Improves Contrastive Representation Learning for Medical Image Segmentation

Yejia Zhang, Xinrong Hu, Nishchal Sapkota, Yiyu Shi, Danny Z. Chen

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted to 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM'22) proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03940 2022-11-09 cs.CL 57%

Tell Your Story: Task-Oriented Dialogs for Interactive Content Creation

Satwik Kottur, Seungwhan Moon, Aram H. Markosyan, Hardik Shah, Babak Damavandi, Alborz Geramifard

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments 8 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10638 2022-11-07 cs.IR cs.AI cs.LG 57%

Digital Human Interactive Recommendation Decision-Making Based on Reinforcement Learning

Xiong Junwu, Xiaoyun Feng, YunZhou Shi, James Zhang, Zhongzhou Zhao, Wei Zhou

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 1 figure, 1 table, the paper has been accepted and this is the final camera-ready for NeurIPS 2022 Workshop on Human in the Loop Learning, https://neurips-hill.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13430 2022-11-01 cs.CV cs.LG 57%

UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Janghyeon Lee, Jongsuk Kim, Hyounguk Shon, Bumsoo Kim, Seung Hwan Kim, Honglak Lee, Junmo Kim

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Neural Information Processing Systems (NeurIPS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13848 2022-11-01 cs.AI cs.RO 57%

ProspectNet: Weighted Conditional Attention for Future Interaction Modeling in Behavior Prediction

Yutian Pang, Zehua Guo, Binnan Zhuang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13591 2022-10-28 cs.CV 57%

Learning by Hallucinating: Vision-Language Pre-training with Weak Supervision

Tzu-Jui Julius Wang, Jorma Laaksonen, Tomas Langer, Heikki Arponen, Tom E. Bishop

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to WACV'23. Please find supplementary material at https://drive.google.com/file/d/1SmCBGsUgkYLAhmK83RZqY03bq4j3214p/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14427 2022-10-27 cs.CL 57%

ReSel: N-ary Relation Extraction from Scientific Text and Tables by Learning to Retrieve and Select

Yuchen Zhuang, Yinghao Li, Jerry Junyang Cheung, Yue Yu, Yingjun Mou, Xiang Chen, Le Song, Chao Zhang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

Comments Accepted to EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏