arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2408.02392 2024-08-06 cs.CV eess.IV 57%

MaFreeI2P: A Matching-Free Image-to-Point Cloud Registration Paradigm with Active Camera Pose Retrieval

Gongxin Yao, Xinyang Li, Yixin Xuan, Yu Pan

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Conference on Multimedia Expo 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20243 2024-07-31 cs.CL cs.LG 57%

Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions

Jinsung Yoon, Raj Sinha, Sercan O Arik, Tomas Pfister

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17671 2024-07-29 cs.CV cs.LG 57%

Unsqueeze [CLS] Bottleneck to Learn Rich Representations

Qing Su, Shihao Ji

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12713 2024-07-23 cs.CV 57%

Dynamic Identity-Guided Attention Network for Visible-Infrared Person Re-identification

Peng Gao, Yujian Lee, Hui Zhang, Xubo Liu, Yiyang Hu, Guquan Jing

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments I need to further debug my code to improve accuracy

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14026 2024-07-22 cs.CV 57%

Semi-supervised reference-based sketch extraction using a contrastive learning framework

Chang Wook Seo, Amirsaman Ashtari, Junyong Noh

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Main paper 1-12 page, Supplementary 13-34 page

Journal ref ACM Transactions on Graphics (TOG) 2023, Volume 42, Issue 4 Article No.: 56, Pages 1 - 12

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13773 2024-07-22 cs.DL cs.AI 57%

OpenDataLab: Empowering General Artificial Intelligence with Open Datasets

Conghui He, Wei Li, Zhenjiang Jin, Chao Xu, Bin Wang, Dahua Lin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08530 2024-07-18 cs.CV 57%

X-Pose: Detecting Any Keypoints

Jie Yang, Ailing Zeng, Ruimao Zhang, Lei Zhang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments ECCV24

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11816 2024-07-16 cs.CV cs.LG 57%

Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning

Jihai Zhang, Xiang Lan, Xiaoye Qu, Yu Cheng, Mengling Feng, Bryan Hooi

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments ECCV 2024 Camera-Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05898 2024-07-09 cs.AI 57%

Contrastive Learning of Preferences with a Contextual InfoNCE Loss

Timo Bertram, Johannes Fürnkranz, Martin Müller

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03615 2024-07-08 cs.CL 57%

Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models

Chang-Sheng Kao, Yun-Nung Chen

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15935 2024-07-04 cs.LG cs.AI 57%

MLEM: Generative and Contrastive Learning as Distinct Modalities for Event Sequences

Viktor Moskvoretskii, Dmitry Osin, Egor Shvetsov, Igor Udovichenko, Maxim Zhelnin, Andrey Dukhovny, Anna Zhimerikina, Evgeny Burnaev

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 11 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01810 2024-07-03 cs.CV 57%

Freeview Sketching: View-Aware Fine-Grained Sketch-Based Image Retrieval

Aneeshan Sain, Pinaki Nath Chowdhury, Subhadeep Koley, Ayan Kumar Bhunia, Yi-Zhe Song

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted in European Conference on Computer Vision (ECCV) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05646 2024-07-03 cs.CV 57%

Masked Attribute Description Embedding for Cloth-Changing Person Re-identification

Chunlei Peng, Boyu Wang, Decheng Liu, Nannan Wang, Ruimin Hu, Xinbo Gao

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01219 2024-07-02 cs.CL 57%

Searching for Best Practices in Retrieval-Augmented Generation

Xiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang, Yixin Wu, Zhibo Xu, Tianyuan Shi, Zhengyuan Wang, Shizheng Li, Qi Qian, Ruicheng Yin, Changze Lv, Xiaoqing Zheng, Xuanjing Huang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18400 2024-06-28 cs.MM 57%

Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning

Hanyao Wang, Yibing Zhan, Liu Liu, Liang Ding, Yan Yang, Jun Yu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM

Comments This work has been submitted to the lEEE for possible publication. Copyright may betransferred without notice, after which this version may no longer be accessible

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12881 2024-06-21 physics.acc-ph cs.CL 57%

Towards Unlocking Insights from Logbooks Using AI

Antonin Sulc, Alex Bien, Annika Eichler, Daniel Ratner, Florian Rehm, Frank Mayet, Gregor Hartmann, Hayden Hoschouer, Henrik Tuennermann, Jan Kaiser, Jason St. John, Jennefer Maldonado, Kyle Hazelwood, Raimund Kammering, Thorsten Hellert, Tim Wilksen, Verena Kain, Wan-Lin Hu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments 5 pages, 1 figure, 15th International Particle Accelerator Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05620 2024-06-11 cs.CV 57%

Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval

Yiwei Ma, Xiaoshuai Sun, Jiayi Ji, Guannan Jiang, Weilin Zhuang, Rongrong Ji

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments ACM MM2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05344 2024-06-11 cs.CL 57%

MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention

Prince Jha, Raghav Jain, Konika Mandal, Aman Chadha, Sriparna Saha, Pushpak Bhattacharyya

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07104 2024-06-07 cs.AI cs.PL 57%

SGLang: Efficient Execution of Structured Language Model Programs

Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05001 2024-06-07 cs.CV cs.LG 57%

COURIER: Contrastive User Intention Reconstruction for Large-Scale Visual Recommendation

Jia-Qi Yang, Chenglei Dai, Dan OU, Dongshuai Li, Ju Huang, De-Chuan Zhan, Xiaoyi Zeng, Yang Yang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05634 2024-06-04 cs.CV 57%

PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-Identification

Quoc-Huy Trinh, Nhat-Tan Bui, Dinh-Hieu Hoang, Phuoc-Thao Vo Thi, Hai-Dang Nguyen, Debesh Jha, Ulas Bagci, Ngan Le, Minh-Triet Tran

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at AVSS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05510 2024-05-29 cs.CV 57%

Mutimodal Ranking Optimization for Heterogeneous Face Re-identification

Hui Hu, Jiawei Zhang, Zhen Han

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments In the methods section, unpaired face samples should be used for training during negative sample training, rather than paired samples. Corresponding errors are also present in the subsequent experimental results

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16328 2024-05-28 cs.CV 57%

A Classifier-Free Incremental Learning Framework for Scalable Medical Image Segmentation

Xiaoyang Chen, Hao Zheng, Yifang Xie, Yuncong Ma, Tengfei Li

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02141 2024-05-20 cs.CV 57%

Zero-shot sketch-based remote sensing image retrieval based on multi-level and attention-guided tokenization

Bo Yang, Chen Wang, Xiaoshuang Ma, Beiping Song, Zhuang Liu, Fangde Sun

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 44 pages, 6 figures

Journal ref Remote Sens. 2024, 16, 1653

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09594 2024-05-17 eess.IV cs.CV cs.LG 57%

Learning Generalized Medical Image Representations through Image-Graph Contrastive Pretraining

Sameer Khanna, Daniel Michael, Marinka Zitnik, Pranav Rajpurkar

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Accepted into Machine Learning for Health (ML4H) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04771 2024-05-09 cs.CV 57%

Exploring Vision Transformers for 3D Human Motion-Language Models with Motion Patches

Qing Yu, Mikihiro Tanaka, Kent Fujiwara

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2024, Project website: https://yu1ut.com/MotionPatches-HP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18836 2024-04-25 cs.CV 57%

ChatPose: Chatting about 3D Human Pose

Yao Feng, Jing Lin, Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Michael J. Black

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Home page: https://yfeng95.github.io/ChatPose/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10795 2024-04-18 cs.SI cs.AI cs.CY cs.IR 57%

Intelligent Message Behavioral Identification System

Yuvaraju Chinnam, Bosubabu Sambana

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07993 2024-04-12 cs.CV 57%

Connecting NeRFs, Images, and Text

Francesco Ballerini, Pierluigi Zama Ramirez, Roberto Mirabella, Samuele Salti, Luigi Di Stefano

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at CVPRW-INRV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15288 2024-04-09 cs.CV stat.ML 57%

Understanding normalization in contrastive representation learning and out-of-distribution detection

Tai Le-Gia, Jaehyun Ahn

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏