arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2506.08968 2025-06-11 cs.CV 57%

ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations

Amirreza Rouhi, Solmaz Arezoomandan, Knut Peterson, Joseph T. Woods, David K. Han

机构 * David S. Hippocampus Department of Computer Science Cranberry-Lemon University(David S. 哈皮科ampus 计算机科学系 Cranberry-Lemon 大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09356 2025-06-05 cs.CV 57%

Galileo: Learning Global & Local Features of Many Remote Sensing Modalities

Gabriel Tseng, Anthony Fuller, Marlena Reil, Henry Herzog, Patrick Beukema, Favyen Bastani, James R. Green, Evan Shelhamer, Hannah Kerner, David Rolnick

机构 * Mila -- Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学) Arizona State University(亚利桑那州立大学) Carleton University(卡尔顿大学) Allen Institute for AI (Ai2)(人工智能研究所) University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01466 2025-06-03 cs.CV 57%

Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark

Shuyu Yang, Yilun Wang, Yaxiong Wang, Li Zhu, Zhedong Zheng

机构 * Xi’an Jiaotong University(西安交通大学) Hefei University of Technology(合肥工业大学) University of Macau(澳门大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19958 2025-06-03 cs.CV 57%

ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language Models

Ke Niu, Haiyang Yu, Mengyang Zhao, Teng Fu, Siyang Yi, Wei Lu, Bin Li, Xuelin Qian, Xiangyang Xue

机构 * Fudan University(复旦大学) Northwestern Polytechnical University(西北工业大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24641 2025-06-02 cs.CV 57%

A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning

Chengzhi Wu, Qianliang Huang, Kun Jin, Julius Pfrommer, Jürgen Beyerer

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24134 2025-06-02 stat.ML cs.CV cs.LG 57%

A Mathematical Perspective On Contrastive Learning

Ricardo Baptista, Andrew M. Stuart, Son Tran

机构 * Stores Foundational AI, Amazon, Palo Alto CA 94301 and Pasadena CA 91125(亚马逊公司、帕洛阿尔托加州94301和帕萨迪纳加州91125) Computing and Mathematical Sciences, California Institute of Technology, Pasadena CA 91125(计算与数学科学系,加州理工学院,帕萨迪纳加州91125)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 44 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23763 2025-05-30 cs.CV 57%

Sketch Down the FLOPs: Towards Efficient Networks for Human Sketch

Aneeshan Sain, Subhajit Maity, Pinaki Nath Chowdhury, Subhadeep Koley, Ayan Kumar Bhunia, Yi-Zhe Song

机构 * SketchX, CVSSP, University of Surrey, United Kingdom(SketchX、CVSSP、塞夫顿大学、英国) Department of Computer Science, University of Central Florida(计算机科学系、中央佛罗里达大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2025, Project Page: https://subhajitmaity.me/SketchDownTheFLOPs

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12051 2025-05-28 cs.CL cs.LG 57%

How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines

Ayan Sengupta, Yash Goel, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments 21 pages, 11 tables, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19213 2025-05-27 cs.AI 57%

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

Shaohao Rui, Kaitao Chen, Weijie Ma, Xiaosong Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17643 2025-05-26 cs.CL cs.LG 57%

Bridging Electronic Health Records and Clinical Texts: Contrastive Learning for Enhanced Clinical Tasks

Sara Ketabi, Dhanesh Ramachandram

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12761 2025-05-26 cs.LG cs.AI 57%

Enhancing Channel-Independent Time Series Forecasting via Cross-Variate Patch Embedding

Donghwa Shin, Edwin Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Added link to code implementation in PDF abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15867 2025-05-23 cs.CV cs.LG 57%

SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval

Nikolaos Chaidos, Angeliki Dimitriou, Maria Lymperaiou, Giorgos Stamou

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14603 2025-05-21 cs.AI cs.LG eess.SP 57%

Towards a Foundation Model for Communication Systems

Davide Buffelli, Sowmen Das, Yu-Wei Lin, Sattar Vakili, Chien-Yi Wang, Masoud Attarifar, Pritthijit Nath, Da-shan Shiu

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13434 2025-05-20 cs.CL 57%

SMOTExT: SMOTE meets Large Language Models

Mateusz Bystroński, Mikołaj Hołysz, Grzegorz Piotrowski, Nitesh V. Chawla, Tomasz Kajdanowicz

机构 * Wrocław University of Science and Technology(沃拉夫大学科学与技术学院) University of Notre Dame(诺特大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12253 2025-05-20 cs.CV 57%

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding

Hanyu Zhou, Gim Hee Lee

机构 * School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11921 2025-05-20 cs.CV 57%

DC-Seg: Disentangled Contrastive Learning for Brain Tumor Segmentation with Missing Modalities

Haitao Li, Ziyu Li, Yiheng Mao, Zhengyao Ding, Zhengxing Huang

机构 * Zhejiang University(浙江大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07896 2025-05-14 q-bio.GN cs.AI 57%

Bridging Large Language Models and Single-Cell Transcriptomics in Dissecting Selective Motor Neuron Vulnerability

Douglas Jiang, Zilin Dai, Luxuan Zhang, Qiyi Yu, Haoqi Sun, Feng Tian

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14123 2025-05-12 cs.CV 57%

AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities

Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05681 2025-05-12 cs.CV 57%

Fine-Tuning Video-Text Contrastive Model for Primate Behavior Retrieval from Unlabeled Raw Videos

Giulio Cesare Mastrocinque Santo, Patrícia Izar, Irene Delval, Victor de Napole Gregolin, Nina S. T. Hirata

机构 * Institute of Mathematics and Statistics, University of São Paulo (IME-USP)(数学统计研究所,圣保罗大学) Department of Experimental Psychology, Institute of Psychology, University of São Paulo (IP-USP)(心理学实验部门,心理学研究所,圣保罗大学) Institute of Biosciences, University of São Paulo (IB-USP)(生物科学研究所,圣保罗大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04846 2025-05-09 cs.IR cs.CE cs.CL cs.DC cs.LG 57%

HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights

Ozan Gokdemir, Carlo Siebenschuh, Alexander Brace, Azton Wells, Brian Hsu, Kyle Hippe, Priyanka V. Setty, Aswathy Ajith, J. Gregory Pauloski, Varuni Sastry, Sam Foreman, Huihuo Zheng, Heng Ma, Bharat Kale, Nicholas Chia, Thomas Gibbs, Michael E. Papka, Thomas Brettin, Francis J. Alexander, Anima Anandkumar, Ian Foster, Rick Stevens, Venkatram Vishwanath, Arvind Ramanathan

机构 * Argonne National Laboratory(阿贡国家实验室) The University of Chicago(芝加哥大学) NVIDIA Inc.(NVIDIA公司) University of Illinois Chicago(伊利诺伊大学芝加哥分校) California Institute of Technology(加州理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments This paper has been accepted at the Platform for Advanced Scientific Computing Conference (PASC 25), June 16-18, 2025, Brugg-Windisch, Switzerland

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03492 2025-05-07 cs.HC cs.AI 57%

Augmenting Human Cognition through Everyday AR

Xiaoan Liu

机构 * New York University(纽约大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 3 pages, 4 figures. Position paper accepted to CHI'25 Workshop 'Everyday AR through AI-in-the-Loop'

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01457 2025-05-07 cs.IR cs.CV 57%

A Multi-Granularity Retrieval Framework for Visually-Rich Documents

Mingjun Xu, Zehui Wang, Hengxing Cai, Renxin Zhong

机构 * DP Technology(DP技术) School of Intelligent Systems Engineering, Sun Yat-Sen University, Shenzhen, China(中山大学智能系统工程学院,深圳,中国)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02325 2025-05-06 cs.CV 57%

TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment

Zhichuan Wang, Yang Zhou, Jinhai Xiang, Yulong Wang, Xinwei He

机构 * Huazhong Agricultural University(华中农业大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18562 2025-04-29 cs.LG cs.AI 57%

Deep Learning with Pretrained 'Internal World' Layers: A Gemma 3-Based Modular Architecture for Wildfire Prediction

Ayoub Jadouli, Chaker El Amrani

机构 * Computer Science and Smart Systems, Faculty of Sciences and Technology, Abdelmalek Essaâdi University(计算机科学与智能系统系,科学与技术学院,阿卜杜勒马利克·埃萨迪大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04339 2025-04-29 cs.CV 57%

NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval

Peng Gao, Yujian Lee, Zailong Chen, Hui zhang, Xubo Liu, Yiyang Hu, Guquang Jing

机构 * Hong Kong Baptist University(香港 Baptist 大学) BNU-HKBU United International College(BNU-HKBU 国际学院) University of Wollongong(沃林根大学) University of Surrey(萨里大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Has been accepted by ICASSP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17990 2025-04-28 cs.CV 57%

From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval

Yabing Wang, Zhuotao Tian, Qingpei Guo, Zheng Qin, Sanping Zhou, Ming Yang, Le Wang

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究所,西安交通大学) Harbin Institute of Technology(哈尔滨工业大学) Ant Group(蚂蚁集团)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04918 2025-04-25 cs.CV 57%

Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and Similarity

Ren-Di Wu, Yu-Yen Lin, Huei-Fang Yang

机构 * National Sun Yat-sen University(国立中山大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 14 pages, 6 figures, International Conference on Technologies and Applications of Artificial Intelligence (TAAI) Camera Ready

Journal ref Technologies and Applications of Artificial Intelligence, pp. 77-90, Springer, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08548 2025-04-16 cs.GR cs.CV 57%

COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails

Miguel Espinosa, Valerio Marsocci, Yuru Jia, Elliot J. Crowley, Mikolaj Czerkawski

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2025 Workshop MORSE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10084 2025-04-15 cs.CV 57%

UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval

Yating Liu, Yaowei Li, Xiangyuan Lan, Wenming Yang, Zimo Liu, Qingmin Liao

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments 16 pages, 7 figures, first submited to IEEE TCSVT on 2024 May. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09302 2025-04-15 cs.AI 57%

Application of Contrastive Learning on ECG Data: Evaluating Performance in Japanese and Classification with Around 100 Labels

Junichiro Takahashi, JingChuan Guan, Masataka Sato, Kaito Baba, Kazuto Haruguchi, Daichi Nagashima, Satoshi Kodera, Norihiko Takeda

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 13 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏