arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2508.00217 2025-08-04 cs.CL cs.DB cs.LG 57%

Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges

Xiaofeng Wu, Alan Ritter, Wei Xu

机构 * College of Computing, Georgia Institute of Technology(计算学院、佐治亚理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23217 2025-08-01 cs.LG cs.AI 57%

Zero-Shot Document Understanding using Pseudo Table of Contents-Guided Retrieval-Augmented Generation

Hyeon Seong Jeong, Sangwoo Jo, Byeong Hyun Yoon, Yoonseok Heo, Haedong Jeong, Taehoon Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21015 2025-07-29 cs.CV 57%

Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions

Licai Sun, Xingxun Jiang, Haoyu Chen, Yante Li, Zheng Lian, Biu Liu, Yuan Zong, Wenming Zheng, Jukka M. Leppänen, Guoying Zhao

机构 * University of Oulu(奥卢大学) Southeast University(东南大学) University of Turku(图尔库大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18881 2025-07-28 cs.CV cs.RO 57%

Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?

Bolei Chen, Jiaxu Kang, Haonan Yang, Ping Zhong, Jianxin Wang

机构 * School of Computer Science Engineering, Central South University Changsha Hunan China Engineering, Central South University

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12591 2025-07-28 cs.CL 57%

LLMs are Also Effective Embedding Models: An In-depth Overview

Chongyang Tao, Tao Shen, Shen Gao, Junshuo Zhang, Zhen Li, Kai Hua, Wenpeng Hu, Zhengwei Tao, Shuai Ma

机构 * Beihang University(北航大学) SKLSDE Lab, Beihang University(北航SKLSDE实验室) University of Technology Sydney(悉尼大学) University of Electronic Science and Technology of China(电子科技大学) Peking University(北京大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

Comments 38 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20292 2025-07-25 cs.CV cs.LG 57%

Visual Adaptive Prompting for Compositional Zero-Shot Learning

Kyle Stein, Arash Mahyari, Guillermo Francia, Eman El-Sheikh

机构 * University of West Florida(乌斯托尔大学) Florida Institute for Human and Machine Cognition (IHMC)(佛罗里达人类与机器认知研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17183 2025-07-21 eess.IV cs.CV 57%

Large-Vocabulary Segmentation for Medical Images with Text Prompts

Ziheng Zhao, Yao Zhang, Chaoyi Wu, Xiaoman Zhang, Xiao Zhou, Ya Zhang, Yanfeng Wang, Weidi Xie

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 74 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01048 2025-07-17 cs.CV cs.CR cs.LG 57%

How does Watermarking Affect Visual Language Models in Document Understanding?

Chunxue Xu, Yiwei Wang, Bryan Hooi, Yujun Cai, Songze Li

机构 * Southeast University, China(东南大学) University of California, Merced, USA(加州大学默塞德分校) National University of Singapore, Singapore(新加坡国立大学) The University of Queensland, Australia(昆士兰大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11479 2025-07-16 cs.AI cs.GR cs.HC 57%

Perspective-Aware AI in Extended Reality

Daniel Platnick, Matti Gruener, Marjan Alirezaie, Kent Larson, Dava J. Newman, Hossein Rahnama

机构 * Flybits Labs(Flybits实验室) Creative Ai Hub(创意人工智能中心) Toronto Metropolitan University(多伦多 Metropolitan 大学) MIT Media Lab(MIT媒体实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Accepted to the International Conference on eXtended Reality (2025), 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06979 2025-07-10 cs.LG cs.CV 57%

A Principled Framework for Multi-View Contrastive Learning

Panagiotis Koromilas, Efthymios Georgiou, Giorgos Bouritsas, Theodoros Giannakopoulos, Mihalis A. Nicolaou, Yannis Panagakis

机构 * Department of Informatics and Telecommunications, National and Kapodistrian University of Athens(信息与通信技术系,雅典国家与卡波迪斯特里亚大学) Archimedes AI/Athena Research Center(阿基米德AI/阿塔纳研究中心) ILSP/Athena Research Center(ILSP/阿塔纳研究中心) The Cyprus Institute(塞浦路斯研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00160 2025-07-08 cs.CL 57%

Emergency Department Decision Support using Clinical Pseudo-notes

Simon A. Lee, Sujay Jain, Alex Chen, Kyoka Ono, Jennifer Fang, Akos Rudas, Jeffrey N. Chiang

机构 * Department of Computational Medicine, University of California, Los Angeles, CA 90095 USA(计算医学系,加州大学洛杉矶分校) Department of Neurosurgery, University of California, Los Angeles, CA 90095 USA(神经外科系,加州大学洛杉矶分校) Department of Electrical and Computer Engineering, University of California at Los Angeles, Los Angeles, CA 90095 USA(电气与计算机工程系,加州大学洛杉矶分校) Harbor-UCLA Medical Center, Department of Emergency Medicine, Torrance, CA(Harbor-UCLA医疗中心,急诊医学部) University of California, Los Angeles, Department of Emergency Medicine, Los Angeles, California(加州大学洛杉矶分校,急诊医学部) Department of Statistics and Data Science University of California, Los Angeles(统计与数据科学系,加州大学洛杉矶分校) Department of Natural Sciences International Christian University, Mitaka, Tokyo, Japan(自然科学系,国际基督教大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Journal ref npj Digital Medicine 8 (1), 394, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03250 2025-07-08 cs.CV cs.LG 57%

Subject Invariant Contrastive Learning for Human Activity Recognition

Yavuz Yarici, Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib

机构 * Georgia Institute of Technology(佐治亚理工学院) Center for Signal and Information Processing CSIP(信号与信息处理中心)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11281 2025-07-02 cs.CV q-bio.QM 57%

DynaCLR: Contrastive Learning of Cellular Dynamics with Temporal Regularization

Eduardo Hirata-Miyasaki, Soorya Pradeep, Ziwen Liu, Alishba Imran, Taylla Milena Theodoro, Ivan E. Ivanov, Sudip Khadka, See-Chi Lee, Michelle Grunberg, Hunter Woosley, Madhura Bhave, Carolina Arias, Shalin B. Mehta

机构 * Chan Zuckerberg Biohub(查纳·泽伯生物枢纽) University of California Berkeley(加州大学伯克利分校)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 30 pages, 6 figures, 13 appendix figures, 5 videos (ancillary files)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23827 2025-07-01 cs.CV 57%

Spatially Gene Expression Prediction using Dual-Scale Contrastive Learning

Mingcheng Qu, Yuncong Wu, Donglin Di, Yue Gao, Tonghua Su, Yang Song, Lei Fan

机构 * Faculty of Computing, Harbin Institute of Technology, Harbin, China(哈尔滨工业大学计算机学院) School of Astronautics, Harbin Institute of Technology, Harbin, China(哈尔滨工业大学航天学院) School of Software, Tsinghua University, Beijing, China(清华大学软件学院) School of Computer Science and Engineering, UNSW, Sydney, Australia(新南威尔士大学计算机科学与工程学院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Our paper has been accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12494 2025-07-01 cs.CL cs.IR 57%

FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation

Zhuocheng Zhang, Yang Feng, Min Zhang

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(中国科学院智能信息处理重点实验室) University of Chinese Academy of Sciences, China(中国科学院大学) Key Laboratory of AI Safety, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室) Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen), China(哈尔滨工业大学(深圳)计算机与智能信息研究院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Accepted by ACL 2025 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11134 2025-07-01 cs.CV 57%

Visual Re-Ranking with Non-Visual Side Information

Gustav Hanning, Gabrielle Flood, Viktor Larsson

机构 * Lund University(吕勒欧大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted at Scandinavian Conference on Image Analysis (SCIA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21934 2025-06-30 cs.IR cs.CV 57%

CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design

Najmeh Forouzandehmehr, Reza Yousefi Maragheh, Sriram Kollipara, Kai Zhao, Topojoy Biswas, Evren Korpeoglu, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13897 2025-06-27 cs.CV 57%

DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity Understanding

Thomas Kreutz, Max Mühlhäuser, Alejandro Sanchez Guinea

机构 * Telekooperation Lab, Technical University Darmstadt(德累斯顿技术大学电信协作实验室) NTT DATA, Luxembourg(NTT DATA卢森堡)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20612 2025-06-27 cs.LG cs.CV 57%

Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning

Vicente Balmaseda, Bokun Wang, Ching-Long Lin, Tianbao Yang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16065 2025-06-26 cs.IR cs.CL 57%

Aug2Search: Enhancing Facebook Marketplace Search with LLM-Generated Synthetic Data Augmentation

Ruijie Xi, He Ba, Hao Yuan, Rishu Agrawal, Yuxin Tian, Ruoyan Kong, Arul Prakash

机构 * North Carolina State University(北卡罗来纳州立大学) Meta

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17886 2025-06-25 cs.SD eess.AS 57%

GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

Julien Guinot, Elio Quinton, György Fazekas

专题命中 跨模态检索 :multimodal(abstract);分类 eess.AS

Comments Accepted to ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18856 2025-06-24 cs.CV 57%

RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base

Kuanning Wang, Yuqian Fu, Tianyu Wang, Yanwei Fu, Longfei Liang, Yu-Gang Jiang, Xiangyang Xue

机构 * Fudan University(复旦大学) NeuhHelium Co.,Ltd.(NeuhHelium公司)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15802 2025-06-24 cs.CV 57%

Visual Prompt Engineering for Vision Language Models in Radiology

Stefan Denner, Markus Bujotzek, Dimitrios Bounias, David Zimmerer, Raphael Stock, Klaus Maier-Hein

机构 * Division of Medical Image Computing, German Cancer Research Center, Heidelberg, Germany(德国癌症研究中心医学图像计算部) Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany(海德堡大学数学与计算机科学学院) Medical Faculty Heidelberg, University of Heidelberg, Heidelberg, Germany(海德堡大学医学学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at ECCV 2024 Workshop on Emergent Visual Abilities and Limits of Foundation Models & Medical Imaging with Deep Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16683 2025-06-23 cs.IR cs.AI 57%

A Simple Contrastive Framework Of Item Tokenization For Generative Recommendation

Penglong Zhai, Yifang Yuan, Fanyi Di, Jie Li, Yue Liu, Chen Li, Jie Huang, Sicong Wang, Yao Xu, Xin Li

机构 * AMAP, Alibaba Group(阿里集团AMAP)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments 12 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13496 2025-06-23 cs.CV cs.IR cs.LG 57%

Hierarchical Multi-Positive Contrastive Learning for Patent Image Retrieval

Kshitij Kavimandan, Angelos Nalmpantis, Emma Beauxis-Aussalet, Robert-Jan Sips

机构 * Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) TKH AI(TKH人工智能)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 3 figures, Accepted as a short paper at the 6th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech 2025), co-located with SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14070 2025-06-18 cs.AI 57%

Into the Unknown: Applying Inductive Spatial-Semantic Location Embeddings for Predicting Individuals' Mobility Beyond Visited Places

Xinglei Wang, Tao Cheng, Stephen Law, Zichao Zeng, Ilya Ilyankou, Junyuan Liu, Lu Yin, Weiming Huang, Natchapon Jongwiriyanurak

机构 * SpaceTimeLab, UCL London UK Department of Geography, UCL London UK 3DIMPact \& SpaceTimeLab, UCL London UK School of Computer Science, University of Surrey Surrey UK Institute for Spatial Data Science, University of Leeds Leeds UK SpaceTimeLab, UCL Department of Geography, UCL 3DIMPact \& SpaceTimeLab, UCL School of Computer Science, University of Surrey Institute for Spatial Data Science, University of Leeds

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13836 2025-06-17 cs.LG cs.AI 57%

Quantifying Memorization and Parametric Response Rates in Retrieval-Augmented Vision-Language Models

Peter Carragher, Abhinand Jha, R Raghav, Kathleen M. Carley

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12099 2025-06-17 cs.CY cs.AI 57%

SocialCredit+

Thabassum Aslam, Anees Aslam

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10550 2025-06-13 cs.CV 57%

ContextRefine-CLIP for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2025

Jing He, Yiqing Wang, Lingling Li, Kexin Zhang, Puhua Chen

机构 * Intelligent Perception and Image Understanding Lab, Xidian University(西安电子科技大学智能感知与图像理解实验室)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10408 2025-06-13 cs.AI cs.IR 57%

Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic Retrieval-Augmented Generation for Industry Challenges

Jintao Liang, Gang Su, Huifeng Lin, You Wu, Rui Zhao, Ziyue Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) University of Georgia(佐治亚大学) South China University of Technology(华南理工大学) Technical University of Munich, University of Cologne(慕尼黑技术大学、科隆大学) SenseTime Research(商汤科技研究院) Qingyuan Research Institute, Shanghai Jiaotong University(青原研究院,上海交通大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏