arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2506.16157 2025-09-23 cs.CV 57%

Proxy-Embedding as an Adversarial Teacher: An Embedding-Guided Bidirectional Attack for Referring Expression Segmentation Models

Xingbai Chen, Tingchao Fu, Renyang Liu, Wei Zhou, Chao Yi

机构 * National Pilot School of Software, Yunnan University(云南大学软件试点学校) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 20pages, 5figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16212 2025-09-23 cs.DB cs.AI 57%

EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics

Ahmad Maroof Karimi, Woong Shin, Jesse Hines, Tirthankar Ghosal, Naw Safrin Sattar, Feiyi Wang

机构 * National Center for Computational Sciences(国家计算科学中心) Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05427 2025-09-18 cs.MM 57%

Reply with Sticker: New Dataset and Model for Sticker Retrieval

Bin Liang, Bingbing Wang, Zhixin Bai, Qiwei Lang, Mingwei Sun, Kaiheng Hou, Lanjun Zhou, Ruifeng Xu, Kam-Fai Wong

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM

Journal ref Liang B, Wang B, Bai Z, et al. Reply with Sticker: New Dataset and Model for Sticker Retrieval[J]. IEEE Transactions on Audio, Speech and Language Processing, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13175 2025-09-17 cs.CV 57%

More performant and scalable: Rethinking contrastive vision-language pre-training of radiology in the LLM era

Yingtai Li, Haoran Lai, Xiaoqian Zhou, Shuai Ming, Wenxin Ma, Wei Wei, Shaohua Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences Medicine, University of Science Technology of China (USTC), Hefei Anhui, 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advance Research, USTC, Suzhou Jiangsu, 215123, China The First Affiliated Hospital of USTC, Division of Life Sciences Medicine, USTC, Hefei Anhui, 230001, China Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Jiangsu, 215123, China State Key Laboratory of Precision

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21875 2025-09-17 cs.AI 57%

Tiny-BioMoE: a Lightweight Embedding Model for Biosignal Analysis

Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis

机构 * Foundation for Research \& Technology-Hellas Heraklion Greece Foundation for Research \& Technology-Hellas Hellenic Mediterranean University Heraklion Greece Hellenic Mediterranean University

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10697 2025-09-16 cs.CL 57%

A Survey on Retrieval And Structuring Augmented Generation with Large Language Models

Pengcheng Jiang, Siru Ouyang, Yizhu Jiao, Ming Zhong, Runchu Tian, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments KDD'25 survey track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09306 2025-09-12 eess.AS eess.IV 57%

Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction

Wenhao Yang, Jianguo Wei, Wenhuan Lu, Xinyue Song, Xianghu Yue

专题命中 跨模态检索 :multimodal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06566 2025-09-09 cs.CV 57%

Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval

Emil Demić, Luka Čehovin Zajc

机构 * Faculty of Computer and Information Science University of Ljubljana(计算机与信息科学系卢布尔雅那大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to BMVC2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01198 2025-09-03 cs.LG cs.AI 57%

Preserving Vector Space Properties in Dimensionality Reduction: A Relationship Preserving Loss Framework

Eddi Weinwurm, Alexander Kovalenko

机构 * Department of Applied Mathematics, Faculty of Information Technology, Czech Technical University in Prague(应用数学系,信息科技学院,布拉格捷克技术大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03567 2025-09-03 cs.CV 57%

Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images

Zengli Luo, Canlong Zhang, Zhixin Li, Zhiwen Wang, Chunrong Wei

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 9 pages, 5 figures. Accepted by the 18th International Conference on Knowledge Science, Engineering and Management (KSEM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05892 2025-09-03 cs.CR cs.AI 57%

PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization

Ruoxi Cheng, Yizhong Ding, Shuirong Cao, Ranjie Duan, Xiaoshuang Jia, Shaowei Yuan, Simeng Qin, Zhiqiang Wang, Xiaojun Jia

机构 * Alibaba Group(阿里巴巴集团) Beijing Electronic Science and Technology Institute(北京电子科技研究所) Nanjing University(南京大学) Renmin University of China(中国人民大学) Northeastern University(东北大学) BraneMatrix AI Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

Comments Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02592 2025-09-03 cs.CV 57%

OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation

Junyuan Zhang, Qintong Zhang, Bin Wang, Linke Ouyang, Zichen Wen, Ying Li, Ka-Ho Chow, Conghui He, Wentao Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Peking University(北京大学) The University of HongKong(香港大学) Shanghai Jiaotong University(上海交通大学) Beihang University(北京航空航天大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19596 2025-08-28 cs.AI cs.IR 57%

Reference-Aligned Retrieval-Augmented Question Answering over Heterogeneous Proprietary Documents

Nayoung Choi, Grace Byun, Andrew Chung, Ellie S. Paek, Shinsun Lee, Jinho D. Choi

机构 * Department of Computer Science Emory University Atlanta Georgia USA(计算机科学系 埃默里大学 阿拉巴马 州 美国) Emory University(埃默里大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments Accepted to CIKM 2025 Applied Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19097 2025-08-27 cs.AI 57%

Reasoning LLMs in the Medical Domain: A Literature Survey

Armin Berger, Sarthak Khanna, David Berghaus, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫研究所媒体工程部门) University of Bonn - Department of Computer Science(波恩大学计算机科学系) West-AI - Federal Ministry of Education and Research(西德人工智能 - 教育与研究部)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17561 2025-08-26 cs.AI cs.LG 57%

Consciousness as a Functor

Sridhar Mahadevan

机构 * Adobe Research and University of Massachusetts, Amherst(Adobe研究院和马萨诸塞大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07493 2025-08-26 cs.CV 57%

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding

Jian Chen, Ming Li, Jihyung Kil, Chenguang Wang, Tong Yu, Ryan Rossi, Tianyi Zhou, Changyou Chen, Ruiyi Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01280 2025-08-26 cs.IR cs.MM 57%

Demo: Soccer Information Retrieval via Natural Queries using SoccerRAG

Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen

专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM

Comments accepted to CBMI 2024 as a demonstration; https://github.com/simula/soccer-rag. arXiv admin note: text overlap with arXiv:2406.01273

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16927 2025-08-26 cs.CV 57%

LGE-Guided Cross-Modality Contrastive Learning for Gadolinium-Free Cardiomyopathy Screening in Cine CMR

Siqing Yuan, Yulin Wang, Zirui Cao, Yueyan Wang, Zehao Weng, Hui Wang, Lei Xu, Zixian Chen, Lei Chen, Zhong Xue, Dinggang Shen

机构 * School of Biomedical Engineering(生物医学工程学院) State Key Laboratory of Advanced Medical Materials and Devices(先进医疗材料与设备国家重点实验室) ShanghaiTech University(上海科技大学) United Imaging Intelligence(联合影像智能) Shanghai Clinical Research and Trial Center(上海临床研究与试验中心) Capital Medical University(首都医科大学) The First Hospital of Lanzhou University(兰州大学第一医院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted to MLMI 2025 (MICCAI workshop); camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16707 2025-08-26 cs.CL cs.IR cs.LG 57%

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval

Jonghyun Song, Youngjune Lee, Gyu-Hwung Cho, Ilhyeon Song, Saehun Kim, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(首尔国立大学数据科学研究生院) NAVER Corporation(NAVER公司)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments accepted to CIKM 2025 short research paper track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15973 2025-08-25 cs.CV 57%

Contributions to Label-Efficient Learning in Computer Vision and Remote Sensing

Minh-Tan Pham

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Habilitation à Diriger des Recherches (HDR) manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12108 2025-08-19 cs.CV 57%

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine

Ziyang Zhang, Yang Yu, Xulei Yang, Si Yong Yeo

机构 * MedVisAI Lab Department of ECE Northwestern University(MedVisAI实验室 电子工程系 西北大学) Institute for Infocomm Research (I 2 R) A*STAR, Singapore(信息与通信研究所(I 2 R)A*STAR,新加坡) MedVisAI Lab Lee Kong Chian School of Medicine, Nanyang Technological University(MedVisAI实验室 李科田医学院,南洋理工大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20844 2025-08-18 cs.IR cs.CL 57%

The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers

Xingyu Deng, Xi Wang, Mark Stevenson

机构 * University of Sheffield(谢菲尔德大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Accepted for ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14493 2025-08-15 cs.IR cs.AI cs.LG 57%

FinSage: A Multi-aspect RAG System for Financial Filings Question Answering

Xinyu Wang, Jijun Chi, Zhenghan Tai, Tung Sum Thomas Kwok, Muzhi Li, Zhuhong Li, Hailin He, Yuchen Hua, Peng Lu, Suyuchen Wang, Yihong Wu, Jerry Huang, Jingrui Tian, Fengran Mo, Yufei Cui, Ling Zhou

机构 * 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments Accepted at the 34th ACM International Conference on Information and Knowledge Management (CIKM2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04376 2025-08-15 cs.CV 57%

MIDAS: Modeling Ground-Truth Distributions with Dark Knowledge for Domain Generalized Stereo Matching

Peng Xu, Zhiyu Xiang, Jingyun Fu, Tianyu Pu, Hanzhi Zhong, Eryun Liu

机构 * Zhejiang University, China(浙江大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07671 2025-08-12 cs.AI cs.CY cs.HC cs.MA stat.AP 57%

EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration

Mohamed Rayan Barhdadi, Mehmet Tuncel, Erchin Serpedin, Hasan Kurban

机构 * Texas A&M University(德克萨斯A&M大学) Istanbul Technical University(伊斯坦布尔技术大学) Hamad Bin Khalifa University(哈利法大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 19 pages, 3 figures (plus 6 figures in supplementary), 2 tables, 1 algorithm. Submitted to NeurIPS 2025 Creative AI Track: Humanity

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16919 2025-08-12 cs.CV 57%

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

Xuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang, Yufei Liu, Kai Wang, Wanli Ouyang, Zhiwei Xiong, Peng Gao, Qibin Hou, Ming-Ming Cheng

机构 * VCIP, CS, Nankai University(南开大学计算机科学与技术学院) NKIARI, Shenzhen Futian(深圳未来科技研究院) USTC(University of Science and Technology of China) CUHK MMLab(香港中文大学MMLab) VAST(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project page: https://github.com/HVision-NKU/TAR3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14428 2025-08-11 cs.CV cs.LG q-bio.QM 57%

WildSAT: Learning Satellite Image Representations from Wildlife Observations

Rangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn, Subhransu Maji

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) GenBio AI University of Edinburgh(爱丁堡大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23029 2025-08-08 cs.CL 57%

Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space

Si Wu, Sebastian Bruch

机构 * Northeastern University(东北大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments The camera-ready version for ACL 2025 in Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23736 2025-08-07 cs.CV cs.IR 57%

Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval

Haiwen Li, Fei Su, Zhicheng Zhao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23735 2025-08-05 cs.RO cs.AI cs.MA 57%

Distributed AI Agents for Cognitive Underwater Robot Autonomy

Markus Buchholz, Ignacio Carlucho, Michele Grimaldi, Yvan R. Petillot

机构 * School of Engineering & Physical Sciences, Heriot-Watt University(工程与物理科学学院,赫里奥特-瓦特大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏