arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2510.12323 2025-10-15 cs.AI 70%

RAG-Anything: All-in-One RAG Framework

Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, Chao Huang

机构 * The University of Hong Kong(香港大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21524 2025-10-14 cs.CV cs.LG stat.ML 70%

Learning Shared Representations from Unpaired Data

Amitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri Shaham

机构 * Department of Computer Science Bar-Ilan University(巴伊兰大学计算机科学系) Electrical and Computer Engineering Technion(技术学院电子与计算机工程系)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02790 2025-10-06 cs.CV cs.AI cs.CL cs.MM 70%

MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding

Jingyuan Deng, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted to emnlp2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26330 2025-10-01 cs.CV cs.IR 70%

SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval

Ren-Di Wu, Yu-Yen Lin, Huei-Fang Yang

机构 * National Sun Yat-sen University(国立中山大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 20 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13146 2025-09-23 cs.CV cs.LG 70%

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) University of Michigan(密歇根大学) UIUC(伊利诺伊大学香槟分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Published at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14746 2025-09-19 cs.CV cs.IR 70%

Chain-of-Thought Re-ranking for Image Retrieval Tasks

Shangrong Wu, Yanghong Zhou, Yang Chen, Feng Zhang, P. Y. Mok

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12994 2025-09-17 cs.CL 70%

SitLLM: Large Language Models for Sitting Posture Health Understanding via Pressure Sensor Data

Jian Gao, Fufangchen Zhao, Yiyang Zhang, Danfeng Yan

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09118 2025-09-12 cs.CV 70%

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Tianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng, Kaicheng Yang, Qichuan Ding

机构 * Northeastern University(东北大学) South China University of Technology(南方科技大学) DeepGlint

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by EMNLP2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04376 2025-09-08 cs.CV 70%

AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search

Hao Ju, Hu Zhang, Zhedong Zheng

机构 * Faculty of Science and Technology and Institute of Collaborative Innovation, University of Macau(科技学院和协同创新研究所,澳门大学) CSIRO Data61(CSIRO数据61)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21539 2025-09-01 cs.CV 70%

HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones

Hao Ruan, Jinliang Lin, Yingxin Lai, Zhiming Luo, Shaozi Li

机构 * Department of Artificial Intelligence, Xiamen University(人工智能学院,厦门大学)

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06104 2025-08-11 cs.CV 70%

MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment

Gui Zou, Chaofan Gan, Chern Hong Lim, Supavadee Aramvith, Weiyao Lin

机构 * Shanghai Jiao Tong University, China(上海交通大学) Monash University, Malaysia(墨尔本大学) Chulalongkorn University, Thailand(朱拉隆功大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICMEW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09459 2025-07-15 cs.CV cs.RO 70%

SegVec3D: A Method for Vector Embedding of 3D Objects Oriented Towards Robot manipulation

Zhihan Kang, Boyu Wang

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Undergraduate Theis; 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07902 2025-07-11 cs.CV 70%

MIRA: A Novel Framework for Fusing Modalities in Medical RAG

Jinhong Wang, Tajamul Ashraf, Zongyan Han, Jorma Laaksonen, Rao Mohammad Anwer

机构 * Department of Computer Vision, MBZUAI(视觉计算系,MBZUAI) Department of Computer Science, Aalto University(计算机科学系,阿alto大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04735 2025-07-08 cs.CV 70%

An analysis of vision-language models for fabric retrieval

Francesco Giuliari, Asif Khan Pattan, Mohamed Lamine Mekhalfi, Fabio Poiesi

机构 * Fondazione Bruno Kessler(布雷诺-科塞勒基金会)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at Ital-IA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13066 2025-06-17 cs.CL 70%

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design

Kai Lan, Jiayong Zhu, Jiangtong Li, Dawei Cheng, Guang Chen, Changjun Jiang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments 26 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11082 2025-06-10 cs.CV cs.AI cs.CL cs.MM 70%

Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning

Chen Jiang, Hong Liu, Xuzheng Yu, Qing Wang, Yuan Cheng, Jia Xu, Zhongyi Liu, Qingpei Guo, Wei Chu, Ming Yang, Yuan Qi

机构 * Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院) Ant Group(蚂蚁集团)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06602 2025-06-10 cs.CV 70%

Zero Shot Composed Image Retrieval

Santhosh Kakarla, Gautama Shastry Bulusu Venkata

机构 * George Mason University(乔治·玛森大学)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05443 2025-06-09 cs.LG cs.AI q-bio.GN 70%

UniPTMs: The First Unified Multi-type PTM Site Prediction Model via Master-Slave Architecture-Based Multi-Stage Fusion Strategy and Hierarchical Contrastive Loss

Yiyu Lin, Yan Wang, You Zhou, Xinye Ni, Jiahui Wu, Sen Yang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17709 2025-06-06 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Contrastive Visual Data Augmentation

Yu Zhou, Bingxuan Li, Mohan Tang, Xiaomeng Jin, Te-Lin Wu, Kuan-Hao Huang, Heng Ji, Kai-Wei Chang, Nanyun Peng

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03148 2025-06-04 cs.CV 70%

Self-Supervised Spatial Correspondence Across Modalities

Ayush Shrivastava, Andrew Owens

机构 * University of Michigan(密歇根大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments CVPR 2025. Project link: https://www.ayshrv.com/cmrw . Code: https://github.com/ayshrv/cmrw

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02291 2025-06-04 cs.CV cs.IR 70%

Entity Image and Mixed-Modal Image Retrieval Datasets

Cristian-Ioan Blaga, Paul Suganthan, Sahil Dua, Krishna Srinivasan, Enrique Alfonseca, Peter Dornbach, Tom Duerig, Imed Zitouni, Zhe Dong

机构 * Google Switzerland(谷歌瑞士分公司) Microsoft AI(微软人工智能)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24441 2025-06-02 cs.CV 70%

SORCE: Small Object Retrieval in Complex Environments

Chunxu Liu, Chi Xie, Xiaxu Chen, Wei Li, Feng Zhu, Rui Zhao, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Sensetime Research(商汤科技研究院) Tongji University(同济大学) Beijing Institute of Technology(北京理工大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Project Page: https://github.com/MCG-NJU/SORCE

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02131 2025-05-08 cs.LG cs.CL 70%

Boosting Masked ECG-Text Auto-Encoders as Discriminative Learners

Hung Manh Pham, Aaqib Saeed, Dong Ma

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16352 2025-04-24 cs.IR cs.AI 70%

Disentangling and Generating Modalities for Recommendation in Missing Modality Scenarios

Jiwan Kim, Hongseok Kang, Sein Kim, Kibum Kim, Chanyoung Park

机构 * KAIST(韩国科学技术院)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

Comments SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09518 2025-04-15 cs.CV 70%

3D CoCa: Contrastive Learners are 3D Captioners

Ting Huang, Zeyu Zhang, Yemin Wang, Hao Tang

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08342 2025-03-13 cs.CV 70%

Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs

Chongjun Tu, Peng Ye, Dongzhan Zhou, Lei Bai, Gang Yu, Tao Chen, Wanli Ouyang

专题命中 跨模态检索 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18733 2025-02-27 cs.LG cs.AI 70%

Cross-Modality Investigation on WESAD Stress Classification

Eric Oliver, Sagnik Dakshit

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08438 2025-02-13 cs.CV cs.AI cs.CL cs.IR cs.MM 70%

Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions

Prajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul, Manish Gupta, Anand Mishra

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at AAAI 2024, 9 pages. Project Website: https://vl2g.github.io/projects/cstbir

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10501 2025-02-10 cs.CV 70%

Enhancing medical vision-language contrastive learning via inter-matching relation modelling

Mingjian Li, Mingyuan Meng, Michael Fulham, David Dagan Feng, Lei Bi, Jinman Kim

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Published at IEEE Transactions on Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07331 2025-02-04 cs.CV cs.LG 70%

Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering

Weixi Weng, Jieming Zhu, Xiaojun Meng, Hao Zhang, Rui Zhang, Chun Yuan

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏