arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

1902.00378 2019-02-04 cs.CV 79%

Self-Supervised Visual Representations for Cross-Modal Retrieval

Yash Patel, Lluis Gomez, Marçal Rusiñol, Dimosthenis Karatzas, C. V. Jawahar

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:1807.02110

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.09854 2019-01-31 cs.CV 79%

Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space

Indrani Bhattacharya, Arkabandhu Chowdhury, Vikas Raykar

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages including reference, 8 figures. First two authors are equal contributors

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.07702 2019-01-24 cs.CV 79%

Exploring Uncertainty in Conditional Multi-Modal Retrieval Systems

Ahmed Taha, Yi-Ting Chen, Xitong Yang, Teruhisa Misu, Larry Davis

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.11013 2018-12-26 cs.CV 79%

Cycle-Consistent Deep Generative Hashing for Cross-Modal Retrieval

Lin Wu, Yang Wang, Ling Shao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments To appeared on IEEE Trans. Image Processing. arXiv admin note: text overlap with arXiv:1703.10593 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.02534 2018-11-15 cs.CL 79%

Using Sparse Semantic Embeddings Learned from Multimodal Text and Image Data to Model Human Conceptual Knowledge

Steven Derby, Paul Miller, Brian Murphy, Barry Devereux

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Proceedings of the 22nd Conference on Computational Natural Language Learning (CoNLL 2018), pages 260-270. Brussels, Belgium, October 31 - November 1, 2018. Association for Computational Linguistics

Journal ref Proceedings of the 22nd Conference on Computational Natural Language Learning (CoNLL 2018), pages 260-270. Brussels, Belgium, October 31 - November 1, 2018. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.09617 2018-10-24 cs.CV 79%

How to Read Paintings: Semantic Art Understanding with Multi-Modal Retrieval

Noa Garcia, George Vogiatzis

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.08266 2018-08-29 cs.CL 79%

A Visual Attention Grounding Neural Model for Multimodal Machine Translation

Mingyang Zhou, Runxiang Cheng, Yong Jae Lee, Zhou Yu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06468 2018-08-21 cs.CY cs.MM 79%

Endogenous and Exogenous Multi-Modal Layers in Context Aware Recommendation Systems for Health

Nitish Nag, Vaibhav Pandey, Ramesh C. Jain

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.MM

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06277 2018-08-21 cs.MM 79%

An Efficient Approach for Geo-Multimedia Cross-Modal Retrieval

Lei Zhu, Jun Long, Chengyuan Zhang, Ruipeng Chen, Xinpan Yuan, Zhan Yang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.MM

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00833 2018-07-27 cs.CV 79%

Learnable PINs: Cross-Modal Embeddings for Person Identity

Arsha Nagrani, Samuel Albanie, Andrew Zisserman

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments To appear in ECCV 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.07998 2018-06-14 cs.CV 79%

Deep Sparse Coding for Invariant Multimodal Halle Berry Neurons

Edward Kim, Darryl Hannan, Garrett Kenyon

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.10819 2018-05-01 cs.CV 79%

Learning Cross-Modal Deep Embeddings for Multi-Object Image Retrieval using Text and Sketch

Sounak Dey, Anjan Dutta, Suman K. Ghosh, Ernest Valveny, Josep Lladós, Umapada Pal

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at ICPR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.01223 2018-04-05 cs.CV 79%

Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval

Chao Li, Cheng Deng, Ning Li, Wei Liu, Xinbo Gao, Dacheng Tao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.02488 2018-02-08 cs.CV 79%

SCH-GAN: Semi-supervised Cross-modal Hashing by Generative Adversarial Network

Jian Zhang, Yuxin Peng, Mingkuan Yuan

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 12 pages, submitted to IEEE Transactions on Cybernetics

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.01943 2018-02-07 cs.CV 79%

Attribute-Guided Network for Cross-Modal Zero-Shot Hashing

Zhong Ji, Yuxin Sun, Yunlong Yu, Yanwei Pang, Jungong Han

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.00358 2017-12-04 cs.CV 79%

Unsupervised Generative Adversarial Cross-modal Hashing

Jian Zhang, Yuxin Peng, Mingkuan Yuan

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 8 pages, accepted by 32th AAAI Conference on Artificial Intelligence (AAAI), 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.09696 2017-09-01 cs.CV cs.LG 79%

Generalized Multi-view Embedding for Visual Recognition and Cross-modal Retrieval

Guanqun Cao, Alexandros Iosifidis, Ke Chen, Moncef Gabbouj

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.04047 2017-07-14 cs.CV 79%

Discrete Multi-modal Hashing with Canonical Views for Robust Mobile Landmark Search

Lei Zhu, Zi Huang, Xiaobai Liu, Xiangnan He, Jingkuan Song, Xiaofang Zhou

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.07026 2017-04-06 cs.LG cs.CV stat.ML 79%

Cross-modal Deep Metric Learning with Multi-task Regularization

Xin Huang, Yuxin Peng

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Revision: Added reference [7] 6 pages, 1 figure, to appear in the proceedings of the IEEE International Conference on Multimedia and Expo (ICME), Jul 10, 2017 - Jul 14, 2017, Hong Kong, Hong Kong

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.00763 2017-04-05 cs.CV 79%

AMC: Attention guided Multi-modal Correlation Learning for Image Search

Kan Chen, Trung Bui, Fang Chen, Zhaowen Wang, Ram Nevatia

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.01101 2017-02-06 cs.CL 79%

Multilingual Multi-modal Embeddings for Natural Language Processing

Iacer Calixto, Qun Liu, Nick Campbell

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

Comments 4 pages (5 including references), no figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.07295 2016-07-26 cs.CV 79%

Learning Aligned Cross-Modal Representations from Weakly Aligned Data

Lluis Castrejon, Yusuf Aytar, Carl Vondrick, Hamed Pirsiavash, Antonio Torralba

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Conference paper at CVPR 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.00715 2016-04-21 cs.CL cs.SI 79%

Multi-Modal Bayesian Embeddings for Learning Social Knowledge Graphs

Zhilin Yang, Jie Tang, William Cohen

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.02705 2016-01-13 cs.RO cs.AI cs.LG 79%

Robobarista: Learning to Manipulate Novel Objects via Deep Multimodal Embedding

Jaeyong Sung, Seok Hyun Jin, Ian Lenz, Ashutosh Saxena

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Journal Version

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.3970 2016-01-01 cs.CV cs.IR cs.LG cs.NE 79%

A Deep and Autoregressive Approach for Topic Modeling of Multimodal Data

Yin Zheng, Yu-Jin Zhang, Hugo Larochelle

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 24 pages, 10 figures. A version has been accepted by TPAMI on Aug 4th, 2015. Add footnote about how to train the model in practice in Section 5.1. arXiv admin note: substantial text overlap with arXiv:1305.5306

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.06746 2015-11-23 cs.CV cs.LG 79%

Images Don't Lie: Transferring Deep Visual Semantic Features to Large-Scale Multimodal Learning to Rank

Corey Lynch, Kamelia Aryafar, Josh Attenberg

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.05670 2015-11-11 cs.CV cs.LG 79%

Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries

Yuke Zhu, Ce Zhang, Christopher Ré, Li Fei-Fei

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
0905.2435 2009-12-01 cs.AI cs.LO 79%

Quantified Multimodal Logics in Simple Type Theory

Christoph Benzmueller, Lawrence C. Paulson

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments ii + 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20318 2026-08-04 cs.CV cs.MM 版本更新 79%

UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

UniCVR:从对齐到重排序的统一零样本复合视觉检索

Haokun Wen, Xuemeng Song, Haoyu Zhang, Weili Guan, Xiangyu Zhao, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学) Southern University of Science and Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室)

专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM

AI总结 UniCVR提出首个统一零样本复合视觉检索框架,结合多模态大语言模型与视觉语言预训练模型,通过两阶段方法实现多任务联合优化,实验表明其在五个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16918 2026-07-15 cs.CV cs.AI 版本更新 79%

Xray-Visual Models: Scaling Vision models on Industry Scale Data

Xray-Visual模型:在产业级数据上扩展视觉模型

Shlok Mishra, Tsung-Yu Lin, Linda Wang, Hongli Xu, Yimin Liu, Michael Hsu, Chaitanya Ahuja, Hao Yuan, Jianpeng Cheng, Hong-You Chen, Haoyuan Xu, Chao Li, Sreya Dutta Roy, Abhijeet Awasthi, Jihye Moon, Don Husa, Michael Ge, Sumedha Singla, Arkabandhu Chowdhury, Phong Dingh, Satya Narayan Shukla, Yonghuan Yang, David Jacobs, Qi Guo, Jun Xiao, Xiangjun Fan, Aashu Singh

机构 * Meta-AI MIT(麻省理工学院) University of Maryland(马里兰大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 Xray-Visual通过三阶段训练流程和LLM2CLIP技术,在产业级数据上实现高效多模态视觉模型,取得最佳性能并提升鲁棒性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏