arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2512.21221 2025-12-25 cs.CV cs.AI 73%

Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval

利用轻量级实体提取实现可扩展的基于事件的图像检索

Dao Sy Duy Minh, Huynh Trung Kiet, Nguyen Lam Phu Quy, Phu-Hoa Pham, Tran Chi Nguyen

机构 * University of Science - VNUHCM(越南胡志明市科学大学)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种轻量级两阶段检索流程,利用事件中心实体提取结合BM25和BEiT-3模型,实现高效准确的图像检索。

Comments System description paper for EVENTA Grand Challenge Track 2 at ACM Multimedia 2025 (MM '25). Ranked 4th place. 6 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02792 2025-12-16 cs.CV cs.MM 73%

HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval

HUD: 用于组合视频检索的层次不确定性感知网络

Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Haokun Wen, Weili Guan

机构 * School of Software Shandong University Jinan China(软件学院 山东大学 济南 中国) School of Data Science City University of Hong Kong Hong Kong China(数据科学学院 香港城市大学 香港 中国) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 HUD通过层次不确定性感知网络提升组合视频检索的多模态查询理解与特征学习精度。

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11050 2025-11-18 cs.CV cs.AI 73%

RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models

Yujin Wang, Quanfeng Liu, Jiaqi Fan, Jinlong Hong, Hongqing Chu, Mengjian Tian, Bingzhao Gao, Hong Chen

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15877 2025-10-15 cs.CV cs.CL cs.LG 73%

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval

Siting Li, Xiang Gao, Simon Shaolei Du

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments NeurIPS 2025; 27 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21871 2025-09-29 cs.CV cs.AI 73%

Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization

Boyang Liu, Yifan Hu, Senjie Jin, Shihan Dou, Gonglei Shi, Jie Shao, Tao Gui, Xuanjing Huang

机构 * Fudan University(复旦大学) Tsinghua University(清华大学) Bytedance(字节跳动)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20813 2025-09-26 cs.CV cs.AI 73%

Revolutionizing Precise Low Back Pain Diagnosis via Contrastive Learning

Thanh Binh Le, Hoang Nhat Khang Vo, Tan-Ha Mai, Trong Nhan Phan

机构 * Faculty of Computer Science and Engineering(计算机科学与工程学院) Ho Chi Minh City University of Technology(胡志明市技术大学) Vietnam National University (HCMUT-VNU)(越南国家大学(HCMUT-VNU)) National Taiwan University(国立台湾大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18772 2025-08-27 cs.CV cs.CL 73%

Beyond the Textual: Generating Coherent Visual Options for MCQs

Wanqiang Wang, Longzhu He, Wei Zheng

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03091 2025-08-06 cs.AI cs.CR cs.CV 73%

T2UE: Generating Unlearnable Examples from Text Descriptions

Xingjun Ma, Hanxun Huang, Tianwei Song, Ye Sun, Yifeng Gao, Yu-Gang Jiang

机构 * Fudan University(复旦大学) The University of Melbourne(墨尔本大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments To appear in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03173 2025-06-06 cs.CV cs.AI 73%

FOLIAGE: Towards Physical Intelligence World Models Via Unbounded Surface Evolution

Xiaoyi Liu, Hao Tang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09948 2025-05-30 cs.CV cs.AI 73%

RadCLIP: Enhancing Radiologic Image Analysis through Contrastive Language-Image Pre-training

Zhixiu Lu, Hailong Li, Nehal A. Parikh, Jonathan R. Dillman, Lili He

机构 * Imaging Research Center, Department of Radiology, Cincinnati Children’s Hospital Medical Center(影像研究中心、放射科、辛辛那提儿童医院医学中心) Artificial Intelligence Imaging Research Center, Cincinnati Children’s Hospital Medical Center(人工智能影像研究中心、辛辛那提儿童医院医学中心) Neurodevelopmental Disorders Prevention Center, Perinatal Institute, Cincinnati Children’s Hospital Medical Center(神经发育障碍预防中心、产科研究所、辛辛那提儿童医院医学中心) Department of Radiology, University of Cincinnati College of Medicine(放射科、辛辛那提大学医学院) Department of Pediatrics, University of Cincinnati College of Medicine(儿科、辛辛那提大学医学院) Department of Computer Science, University of Cincinnati(计算机科学系、辛辛那提大学) Department of Biomedical Engineering, University of Cincinnati(生物医学工程系、辛辛那提大学) Department of Biomedical Informatics, University of Cincinnati College of Medicine(生物医学信息学系、辛辛那提大学医学院)

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10995 2025-04-16 cs.CV cs.AI 73%

TMCIR: Token Merge Benefits Composed Image Retrieval

Chaoyang Wang, Zeyu Zhang, Long Teng, Zijun Li, Shichao Kan

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: text overlap with arXiv:2310.05473 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05606 2025-03-11 cs.CV cs.MM 73%

CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization

Nan Chen, Mengqi Huang, Zhuowei Chen, Yang Zheng, Lei Zhang, Zhendong Mao

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11694 2025-03-05 cs.AI cs.CL cs.LG 73%

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

Shixin Jiang, Jiafeng Liang, Jiyuan Wang, Xuan Dong, Heng Chang, Weijiang Yu, Jinhua Du, Ming Liu, Bing Qin

专题命中 跨模态检索 :multi-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI

Comments 35 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12799 2025-02-19 cs.CL cs.CV cs.IR 73%

Towards Text-Image Interleaved Retrieval

Xin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu, Wenjie Li, Min Zhang

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments 16 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04371 2025-02-10 cs.AI cs.CL cs.LG 73%

PerPO: Perceptual Preference Optimization via Discriminative Rewarding

Zining Zhu, Liang Zhao, Kangheng Lin, Jinze Yang, En Yu, Chenglong Liu, Haoran Wei, Jianjian Sun, Zheng Ge, Xiangyu Zhang

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08347 2025-01-16 cs.CV cs.AI 73%

SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval

Bhavin Jawade, Joao V. B. Soares, Kapil Thadani, Deen Dayal Mohan, Amir Erfan Eshratifar, Benjamin Culpepper, Paloma de Juan, Srirangaraj Setlur, Venu Govindaraju

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Paper accepted at WACV 2025 in round 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02202 2024-09-10 cs.CV cs.AI 73%

No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs

Cristian Sbrolli, Matteo Matteucci

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments to be published in BMVC 2024 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05503 2024-08-13 cs.CV cs.AI 73%

Disentangled Noisy Correspondence Learning

Zhuohang Dang, Minnan Luo, Jihong Wang, Chengyou Jia, Haochen Han, Herun Wan, Guang Dai, Xiaojun Chang, Jingdong Wang

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19415 2024-07-30 cs.MM cs.AI 73%

Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval

Zeyu Chen, Pengfei Zhang, Kai Ye, Wei Dong, Xin Feng, Yana Zhang

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19149 2024-05-31 cs.CV cs.AI cs.IR 73%

CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval

Xintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu, Bingwen Hu, Xueming Qian

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments To appear at SIGIR 2024. arXiv admin note: text overlap with arXiv:2309.02169

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02169 2024-02-01 cs.CV cs.AI 73%

Dual Relation Alignment for Composed Image Retrieval

Xintong Jiang, Yaxiong Wang, Yujiao Wu, Meng Wang, Xueming Qian

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments The architecture of our model changes, hence methodolgy and experiments changes a lot, We have significantly revised the original manuscript of the paper, so a withdraw of our original script is needed

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07510 2024-01-23 cs.CL cs.AI 73%

Developing ChatGPT for Biology and Medicine: A Complete Review of Biomedical Question Answering

Qing Li, Lei Li, Yu Li

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments 50 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02240 2023-12-06 cs.CV cs.AI 73%

Contrastive Learning-Based Spectral Knowledge Distillation for Multi-Modality and Missing Modality Scenarios in Semantic Segmentation

Aniruddh Sikdar, Jayant Teotia, Suresh Sundaram

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11331 2023-08-23 cs.CV cs.AI 73%

GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training

Xinchi Deng, Han Shi, Runhui Huang, Changlin Li, Hang Xu, Jianhua Han, James Kwok, Shen Zhao, Wei Zhang, Xiaodan Liang

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06275 2023-04-14 cs.CV cs.MM 73%

Noisy Correspondence Learning with Meta Similarity Correction

Haochen Han, Kaiyao Miao, Qinghua Zheng, Minnan Luo

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01961 2023-04-05 cs.IR cs.CL cs.CV 73%

AToMiC: An Image/Text Retrieval Test Collection to Support Multimedia Content Creation

Jheng-Hong Yang, Carlos Lassance, Rafael Sampaio de Rezende, Krishna Srinivasan, Miriam Redi, Stéphane Clinchant, Jimmy Lin

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10163 2022-10-20 cs.CV cs.CL 73%

MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng Sun

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02531 2022-06-07 cs.CV cs.AI 73%

3D-Augmented Contrastive Knowledge Distillation for Image-based Object Pose Estimation

Zhidan Liu, Zhen Xing, Xiangdong Zhou, Yijiang Chen, Guichun Zhou

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted for presentation at International Conference on Multimedia Retrieval (ICMR '22)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14463 2022-04-18 cs.CV cs.CL 73%

Large-scale Bilingual Language-Image Contrastive Learning

Byungsoo Ko, Geonmo Gu

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICLRW2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03809 2022-03-09 cs.CV cs.CL 73%

Image Search with Text Feedback by Additive Attention Compositional Learning

Yuxin Tian, Shawn Newsam, Kofi Boakye

专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏