arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2309.01017 2023-09-06 cs.CV 57%

Contrastive Grouping with Transformer for Referring Image Segmentation

Jiajin Tang, Ge Zheng, Cheng Shi, Sibei Yang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00372 2023-09-04 eess.IV cs.CV 57%

On the Localization of Ultrasound Image Slices within Point Distribution Models

Lennart Bastian, Vincent Bürgin, Ha Young Kim, Alexander Baumann, Benjamin Busam, Mahdi Saleh, Nassir Navab

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments ShapeMI Workshop @ MICCAI 2023; 12 pages 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14126 2023-08-29 cs.CV 57%

Synergizing Contrastive Learning and Optimal Transport for 3D Point Cloud Domain Adaptation

Siddharth Katageri, Arkadipta De, Chaitanya Devaguptapu, VSSV Prasad, Charu Sharma, Manohar Kaul

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.11752 2023-08-29 cs.CV 57%

CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose

Xu Zhang, Wen Wang, Zhe Chen, Yufei Xu, Jing Zhang, Dacheng Tao

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11994 2023-08-24 cs.CV 57%

Progressive Feature Mining and External Knowledge-Assisted Text-Pedestrian Image Retrieval

Huafeng Li, Shedan Yang, Yafei Zhang, Dapeng Tao, Zhengtao Yu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08942 2023-08-24 cs.CV 57%

Spherical Space Feature Decomposition for Guided Depth Map Super-Resolution

Zixiang Zhao, Jiangshe Zhang, Xiang Gu, Chengli Tan, Shuang Xu, Yulun Zhang, Radu Timofte, Luc Van Gool

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11485 2023-08-23 cs.CV 57%

Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features

Alberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto del Bimbo

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Accepted in ACM Transactions on Multimedia Computing Communications and Applications (TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06309 2023-08-22 cs.CV 57%

UATVR: Uncertainty-Adaptive Text-Video Retrieval

Bo Fang, Wenhao Wu, Chang Liu, Yu Zhou, Yuxin Song, Weiping Wang, Xiangbo Shu, Xiangyang Ji, Jingdong Wang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments To appear at ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09107 2023-08-21 cs.CV 57%

Hyperbolic Face Anti-Spoofing

Shuangpeng Han, Rizhao Cai, Yawen Cui, Zitong Yu, Yongjian Hu, Alex Kot

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07078 2023-08-15 cs.CV 57%

ICPC: Instance-Conditioned Prompting with Contrastive Learning for Semantic Segmentation

Chaohui Yu, Qiang Zhou, Zhibin Wang, Fan Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15872 2023-08-01 eess.IV cs.CV cs.LG 57%

Cross-dimensional transfer learning in medical image segmentation with deep learning

Hicham Messaoudi, Ahror Belaid, Douraied Ben Salem, Pierre-Henri Conze

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments 30 pages, 12 figures, 6 tables, Accepted for publication in the Journal of Medical Image Analysis

Journal ref In Medical Image Analysis (Vol. 88, p. 102868). Elsevier BV (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.00223 2023-07-19 cs.CV 57%

Unsupervised Deep Cross-modality Spectral Hashing

Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen, Ngai-Man Cheung

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transaction on Image Processing (TIP) Add Acknowledgement

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01577 2023-07-06 cs.AI q-bio.NC 57%

Conceptual Cognitive Maps Formation with Neural Successor Networks and Word Embeddings

Paul Stoewer, Achim Schilling, Andreas Maier, Patrick Krauss

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06691 2023-06-13 cs.CV 57%

Self-Enhancement Improves Text-Image Retrieval in Foundation Visual-Language Models

Yuguang Yang, Yiming Wang, Shupeng Geng, Runqi Wang, Yimi Wang, Sheng Wu, Baochang Zhang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2023 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06382 2023-06-05 cs.CV 57%

Differentiated Relevances Embedding for Group-based Referring Expression Comprehension

Fuhai Chen, Xuri Ge, Xiaoshuai Sun, Yue Gao, Jianzhuang Liu, Fufeng Chen, Wenjie Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19329 2023-06-01 cs.CV cs.IR cs.LG 57%

Mitigating Test-Time Bias for Fair Image Retrieval

Fanjie Kong, Shuai Yuan, Weituo Hao, Ricardo Henao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13500 2023-05-24 cs.CV 57%

Learning Emotion Representations from Verbal and Nonverbal Communication

Sitao Zhang, Yimu Pan, James Z. Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11503 2023-05-22 cs.CL 57%

A Topic-aware Summarization Framework with Different Modal Side Information

Xiuying Chen, Mingzhe Li, Shen Gao, Xin Cheng, Qiang Yang, Qishen Zhang, Xin Gao, Xiangliang Zhang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments SIGIR 2023, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06923 2023-05-12 cs.CV 57%

EAML: Ensemble Self-Attention-based Mutual Learning Network for Document Image Classification

Souhail Bakkali, Ziheng Ming, Mickael Coustaty, Marçal Rusiñol

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted at IJDAR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02265 2023-05-08 cs.CL 57%

A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex Text

Yunxin Li, Baotian Hu, Yuxin Ding, Lin Ma, Min Zhang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CL

Comments Accepted to ACL 2023 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02760 2023-05-05 cs.CV 57%

Multi-Modality Deep Network for JPEG Artifacts Reduction

Xuhao Jiang, Weimin Tan, Qing Lin, Chenxi Ma, Bo Yan, Liquan Shen

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 18 pages, 17 figures, accepted by IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00131 2023-05-02 cs.CV 57%

Regularizing Self-training for Unsupervised Domain Adaptation via Structural Constraints

Rajshekhar Das, Jonathan Francis, Sanket Vaibhav Mehta, Jean Oh, Emma Strubell, Jose Moura

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09520 2023-05-02 cs.CV 57%

Using Language to Extend to Unseen Domains

Lisa Dunlap, Clara Mohri, Devin Guillory, Han Zhang, Trevor Darrell, Joseph E. Gonzalez, Aditi Raghunathan, Anja Rohrbach

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04662 2023-04-24 cs.LG cs.AI 57%

Cognitively Inspired Learning of Incremental Drifting Concepts

Mohammad Rostami, Aram Galstyan

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 2023 International Joint Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07747 2023-04-18 cs.CV 57%

Language Guided Local Infiltration for Interactive Image Retrieval

Fuxiang Huang, Lei Zhang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments 10 pages, 9 figures, 4 tables, IEEE/CVF Conference on Computer Vision and Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07387 2023-04-18 cs.MM 57%

Cross-domain Food Image-to-Recipe Retrieval by Weighted Adversarial Learning

Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Wing-Kwong Chan

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.07738 2023-04-17 cs.CL cs.IR cs.LG 57%

Alloprof: a new French question-answer education dataset and its use in an information retrieval case study

Antoine Lefebvre-Brossard, Stephane Gazaille, Michel C. Desmarais

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03669 2023-04-10 cs.CV 57%

DATE: Domain Adaptive Product Seeker for E-commerce

Haoyuan Li, Hao Jiang, Tao Jin, Mengyan Li, Yan Chen, Zhijie Lin, Yang Zhao, Zhou Zhao

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments This paper was accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10624 2023-04-04 cs.CV 57%

A Unified Model for Video Understanding and Knowledge Embedding with Heterogeneous Knowledge Graph Dataset

Jiaxin Deng, Dong Shen, Haojie Pan, Xiangyu Wu, Ximan Liu, Gaofeng Meng, Fan Yang, Size Li, Ruiji Fu, Zhongyuan Wang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICMR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.01872 2023-04-04 cs.CV 57%

Parts2Words: Learning Joint Embedding of Point Clouds and Texts by Bidirectional Matching between Parts and Words

Chuan Tang, Xi Yang, Bojian Wu, Zhizhong Han, Yi Chang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏