arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2604.18051 2026-04-21 cs.CV 70%

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

INTENT: 为鲁棒的复合图像检索的不变性与判别意识噪声缓解

Zhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li, Jiale Huang, Qinlei Huang, Yinwei Wei

机构 * School of Software, Shandong University(山东大学软件学院)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出INTENT网络,通过视觉不变组成和双目标判别学习处理复合图像检索中的两种噪声类型,提升检索鲁棒性。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17581 2026-04-21 cs.LG cs.AI q-bio.NC 70%

How Much Data is Enough? The Zeta Law of Discoverability in Biomedical Data, featuring the enigmatic Riemann zeta function

需要多少数据?生物医学数据的发现率ζ定律,涉及神秘的黎曼ζ函数

Paul M. Thompson

机构 * Stevens Institute for Neuroimaging & Informatics(神经影像与信息学史蒂文斯研究所) University of Southern California(南加州大学) Los Angeles, CA, USA(洛杉矶,加利福尼亚州,美国)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出基于数据协方差算子谱结构的跨模态发现率定律框架,揭示性能指标与ζ函数的关系,为数据规模、表示学习和多模态扩展提供理论指导。

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16247 2026-04-20 cs.LG cs.AI 70%

Joint-Centric Dual Contrastive Alignment with Structure-Preserving and Information-Balanced Regularization

基于结构保持和信息平衡正则化的联合中心双对比对齐

Habibeh Naderi, Behrouz Haji Soleimani, Stan Matwin

机构 * Dalhousie University(达尔豪斯大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出HILBERT框架,通过跨模态注意力和自注意力池化学习文档级音频文本表示,采用双对比目标和辅助正则化提升跨模态对齐效果,实现长序列语义表示学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14710 2026-04-17 cs.CV 70%

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval

G-MIXER:基于测地混合的隐式语义扩展和显式语义重新排序用于零样本复合图像检索

Jiyoung Lim, Heejae Yang, Jee-Hyong Lee

机构 * Sungkyunkwan University(顺天大学)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 G-MIXER通过测地混合生成隐式语义特征并重新排序显式语义,提升零样本复合图像检索的多样性和准确性,实现跨多个基准的最优性能。

Comments CVPR 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14147 2026-04-16 cs.CV 70%

ROSE: Retrieval-Oriented Segmentation Enhancement

ROSE:基于检索的分割增强

Song Tang, Guangquan Jie, Henghui Ding, Yu-Gang Jiang

机构 * Fudan University(复旦大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出ROSE框架,通过引入网络检索模块和提示增强模块,提升多模态大语言模型在新兴实体和新实体分割任务中的性能,实验表明其在NEST基准上表现优异。

Comments CVPR 2026 Findings, Project Page: https://henghuiding.com/ROSE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18994 2026-04-14 cs.CV 70%

Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy

细粒度长尾植物分类的双边嵌入

Cheng Yaw Low, Heejoon Koo, Jaewoo Park, Meeyoung Cha

机构 * Changwon National University(昌原国立大学) Max Planck Institute for Security and Privacy (MPI-SP)(马克斯·普朗克安全与隐私研究所) AiV Co.(AiV公司)

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 本文提出TaxoNet框架,通过理论支撑的双边目标优化类别决策边界,在类别不平衡情况下提升细粒度区分能力和稀有类表示几何。在开放世界设置中,利用多种植物数据集验证了其优于多模态基础模型的性能。

Comments 4 figures, 5 tables, and 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15623 2026-04-09 cs.CV 70%

PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence Learning

PCSR:基于伪标签一致性的样本细化用于噪声对应学习

Zhuoyao Liu, Yang Liu, Wentao Feng, Shudong Huang

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出PCSR框架,通过伪标签一致性提升对应可靠性,解决噪声对应问题,提升检索鲁棒性。

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26341 2026-03-30 cs.CV 70%

HINT: Composed Image Retrieval with Dual-path Compositional Contextualized Network

HINT: 基于双路径组合上下文化网络的复合图像检索

Mingyu Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Xiaowei Zhu, Jiajia Nie, Yinwei Wei, Yupeng Hu

机构 * School of Software, Shandong University(山东大学软件学院)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出HINT模型,通过双路径组合上下文化网络增强匹配与非匹配样本的相似度差异,提升复杂场景下复合图像检索性能。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25891 2026-03-30 cs.CV 70%

Few Shots Text to Image Retrieval: New Benchmarking Dataset and Optimization Methods

少样本文本到图像检索:新基准数据集和优化方法

Ofer Idan, Vladi Vexler, Gil Lederman, Dima Sivov, Aviad Cohen Zada, Shir Niego Komforti

机构 * Huawei Tel-Aviv Research Center(华为特拉维夫研究中心)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出FSIR任务及FSIR-BD数据集,通过少样本学习方法提升图像检索性能,实验表明其基准挑战性及优化方法优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14961 2026-03-27 cs.LG cs.CV 70%

Graph Memory: A Structured and Interpretable Framework for Modality-Agnostic Embedding-Based Inference

图记忆:一种结构化且可解释的多模态嵌入推理框架

Artur A. Oliveira, Mateus Espadoto, Roberto M. Cesar, Roberto Hirata

机构 * Institute of Mathematics and Statistics, University of São Paulo(数学统计研究所,圣保罗大学)

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 图记忆通过紧凑的可靠性标注原型区域图表示嵌入空间,结合原型关系编码局部几何和区域模糊性,实现查询证据扩散推理,统一实例检索、原型推理和图扩散,提供更准确和可解释的多模态推理。

Comments This version expands the published conference paper (VISAPP 2026) with additional methodological details, experiments, and analysis that were omitted due to page limits. The final published version is available via DOI: 10.5220/0014578800004084

Journal ref Proc. 21st Int. Conf. Comput. Vision Theory Appl. (VISAPP 2026), Vol. 1, pp. 652-659 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21886 2026-03-24 cs.IR cs.CV 70%

ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval

ADaFuSE: 适应性扩散生成图像与文本融合用于交互式文本到图像检索

Zhuocheng Zhang, Xingwu Zhang, Kangheng Liang, Guanxuan Li, Richard Mccreadie, Zijun Long

机构 * Hunan University(湖南大学) University of Glasgow(格拉斯哥大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 ADaFuSE通过引入双分支融合机制,结合自适应门控和语义感知专家混合,提升交互式文本到图像检索的性能,实现更稳健的多模态融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09894 2026-03-17 cs.AI cs.CY cs.LG 70%

Beyond AlphaEarth: Toward Human-Centered Geospatial Foundation Models via POI-Guided Contrastive Learning

超越AlphaEarth:通过POI引导对比学习实现以人为中心的地理空间基础模型

Junyuan Liu, Quan Qin, Guangsheng Dong, Xinglei Wang, Jiazhuang Feng, Zichao Zeng, Tao Cheng

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出AETHER框架,通过POI引导的多模态对齐,将AlphaEarth与以人为中心的城市分析结合,提升地理空间表示的可解释性与语言可访问性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06704 2026-03-10 cs.CV cs.LG 70%

On the Generalization Capacities of MLLMs for Spatial Intelligence

关于多模态大语言模型在空间智能中的泛化能力

Gongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang, Deli Zhao, Shijian Lu, Ran Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) HuPan Lab(虎朋实验室) Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出Camera-Aware MLLM框架,通过注入相机内参、数据增强和蒸馏几何先验,提升多模态大语言模型在空间任务中的泛化能力与鲁棒性。

Comments ICLR 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04614 2026-03-06 cs.CV 70%

SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D

SGR3模型:三维场景图检索-推理模型

Zirui Wang, Ruiping Liu, Yufan Chen, Junwei Zheng, Weijia Fan, Kunyu Peng, Di Wen, Jiale Wei, Jiaming Zhang, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Shenzhen University(深圳大学) Hunan University(湖南大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 SGR3模型通过多模态大语言模型与检索增强生成技术,实现无需训练的三维场景图生成,提升场景关系推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03617 2026-03-05 cs.CV 70%

RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

RAGTrack: 基于检索增强生成的语言感知RGBT跟踪

Hao Li, Yuhao Wang, Wenning Hao, Pingping Zhang, Dong Wang, Huchuan Lu

机构 * College of Command and Control Engineering, Army Engineering University of PLA(中国人民解放军陆军工程大学指挥控制工程学院) School of Future Technology, Dalian University of Technology(大连理工大学未来技术学院) School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 RAGTrack通过引入语言指导的检索增强生成框架,解决RGBT跟踪中外观变化和模态间隙问题,实现稳健的目标定位。

Comments This work is accepted by CVPR2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01632 2026-03-03 cs.LG cs.AI 70%

DeLo: Dual Decomposed Low-Rank Experts Collaboration for Continual Missing Modality Learning

DeLo:双分解低秩专家协作用于持续缺失模态学习

Xiwei Liu, Yulong Li, Feilong Tang, Imran Razzak

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 DeLo通过双分解低秩专家架构解决持续缺失模态学习中的模态干扰问题,结合跨模态引导路由和任务键记忆实现高效推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04932 2026-03-03 cs.CV 70%

UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features

UniView: 通过统一参考特征从单张图像增强新颖视角合成

Haowang Cui, Rui Chen, Jiaze Wang, Tao Guo, Zheng Qin

机构 * Tianjin University(天津大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 UniView通过统一参考特征提升单张图像新颖视角合成性能,采用检索增强系统和多模态大语言模型选择参考图像,并结合解耦三重注意力机制提高合成效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20530 2026-02-25 cs.LG cs.SD eess.AS 70%

Memory-guided Prototypical Co-occurrence Learning for Mixed Emotion Recognition

基于记忆的原型共现学习用于混合情绪识别

Ming Li, Yong-Jin Liu, Fang Liu, Huankun Sheng, Yeying Fan, Yixiang Wei, Minnan Luo, Weizhan Zhang, Wenping Wang

机构 * MOE-Key Laboratory of Pervasive Computing, Department of Computer Science and Technology, Tsinghua University(MOE-Key pervasive computing laboratory,计算机科学与技术系,清华大学) Key State Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播关键实验室,中国传媒大学) Department of Computer Science and Computer Engineering, Texas A&M University(计算机科学与工程系,德克萨斯大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

AI总结 本文提出基于记忆的原型共现学习框架,通过多模态信号融合和原型关系蒸馏,提升混合情绪识别的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03098 2026-02-24 cs.LG cs.AI 70%

TextME: Bridging Unseen Modalities Through Text Descriptions

TextME: 通过文本描述弥合未见模态

Soyeon Hong, Jinchan Kim, Jaegook You, Seungtaek Choi, Suha Kwak, Hyunsouk Cho

机构 * Department of Artificial Intelligence, Ajou University, Suwon, South Korea Division of Language \& AI, Hankuk University of Foreign Studies, Seoul, Korea Graduate School of AI, POSTECH, Pohang, Korea Department of Software, Ajou University, Suwon, South Korea

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 TextME通过仅使用文本描述实现跨模态扩展,无需配对监督,有效弥合不同模态间的差距。

Comments Code available at https://github.com/SoyeonHH/TextME

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17402 2026-02-20 cs.AI 70%

A Contrastive Variational AutoEncoder for NSCLC Survival Prediction with Missing Modalities

用于NSCLC生存预测的对比变分自编码器,具有缺失模态

Michele Zanitti, Vanja Miskovic, Francesco Trovò, Alessandra Laura Giulia Pedrocchi, Ming Shen, Yan Kyaw Tun, Arsela Prelaj, Sokol Kosta

机构 * Department of Electronic Systems, Aalborg University, Copenhagen, Denmark(电子系统系,奥胡斯大学) Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy(电子、信息与生物工程系,米兰理工学院) Department of Medical Oncology, Istituto Nazionale dei Tumori, Milan, Italy(医学肿瘤学系,国家肿瘤研究所)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种多模态对比变分自编码器,用于NSCLC生存预测,通过整合多种数据模态并处理缺失数据,提高预测的鲁棒性和准确性。

Comments Accepted at The 13th IEEE International Conference on Big Data (IEEE BigData 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21033 2026-02-03 cs.SD cs.AI 70%

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

SupCLAP:通过支持向量正则化控制音频-文本对比学习中的优化轨迹漂移

Jiehui Luo, Yuguo Yin, Yuxin Xie, Jinghan Ru, Xianwei Zhuang, Minghua He, Aofan Liu, Zihan Xiong, Dongchao Yang

机构 * Peking University(北京大学) Central Conservatory of Music(中央音乐学院) The Chinese University of Hong Kong(香港中文大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 SupCLAP通过支持向量正则化有效控制音频-文本对比学习中的优化轨迹漂移,提升多模态学习的稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00131 2026-02-03 cs.CV cs.RO 70%

PovNet+: A Deep Learning Architecture for Socially Assistive Robots to Learn and Assist with Multiple Activities of Daily Living

PovNet+: 一种深度学习架构用于社交辅助机器人学习和协助多种日常活动

Fraser Robinson, Souren Pashangpour, Matthew Lisondra, Goldie Nejat

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 PovNet+是一种多模态深度学习架构,用于社交辅助机器人识别多种日常活动并主动发起辅助行为,提升了ADL分类准确率和人机交互能力。

Comments Submitted to Advanced Robotics (Taylor & Francis)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14259 2026-01-21 cs.CV 70%

ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation

ManipShield: 一种用于图像篡改检测、定位和解释的统一框架

Zitong Xu, Huiyu Duan, Xiaoyu Wang, Zhaolin Cai, Kaiwei Zhang, Qiang Hu, Jing Liu, Xiongkuo Min, Guangtao Zhai

机构 * Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所) University of Electronic and Science Technology of China(电子科技大学) Tianjin University(天津大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 ManipShield基于多模态大语言模型,通过对比学习LoRA微调和任务特定解码器,实现图像篡改的统一检测、定位和解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13072 2025-12-16 cs.CV 70%

Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

锻造动态记忆:基于检索的持续学习用于通用医学基础模型

Zizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu, Zhaoyu Chen, Dingkang Yang, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) Fysics Intelligence Technologies Co., Ltd. (Fysics AI)(Fysics智能技术有限公司(Fysics AI)) School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学)

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出基于检索的持续学习方法,通过动态知识蒸馏和RAG技术,解决多模态医学模型在领域迁移和细粒度特征保留中的核心难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18298 2025-11-25 cs.AI 70%

Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery

跨学科知识检索与综合:一种用于科学发现的复合AI架构

Svitlana Volkova, Peter Bautista, Avinash Hiriyanna, Gabriel Ganberg, Isabel Erickson, Zachary Klinefelter, Nick Abele, Hsien-Te Kao, Grant Engberson

机构 * Aptima, Inc.(Aptima公司)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 BioSage通过整合LLMs与RAG,利用专门代理实现跨学科知识检索与综合,提升科学发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV 70%

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10390 2025-11-18 cs.CV 70%

HIBMatch: Hypergraph Information Bottleneck for Semi-supervised Alzheimer's Progression

Zhongying Deng, Shujun Wang, Angelica I Aviles-Rivero, Zoe Kourtzi, Carola-Bibiane Schönlieb

机构 * Department of Applied Mathematics and Theoretical Physics, University of Cambridge(应用数学与理论物理系,剑桥大学) Department of Biomedical Engineering, The Hong Kong Polytechnic University(生物医学工程系,香港理工大学) Research Institute for Artificial Intelligence of Things, The Hong Kong Polytechnic University(物联网人工智能研究所,香港理工大学) Yau Mathematical Sciences Centre, Tsinghua University(叶德平数学科学中心,清华大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to the IEEE Journal of Biomedical and Health Informatics (To appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05630 2025-11-11 q-bio.NC cs.AI 70%

BrainCSD: A Hierarchical Consistency-Driven MoE Foundation Model for Unified Connectome Synthesis and Multitask Brain Trait Prediction

Xiongri Shen, Jiaqi Wang, Yi Zhong, Zhenxi Song, Leilei Zhao, Liling Li, Yichen Wei, Lingyan Liang, Shuqiang Wang, Baiying Lei, Demao Deng, Zhiguo Zhang

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术系) School of Intelligence Science and Engineering, College of Artificial Intelligence, Harbin Institute of Technology(哈尔滨工业大学智能科学与工程学院) School of Biomedical Engineering, National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, Guangdong Key Laboratory for Biomedical, Measurements and Ultrasound Imaging, Shenzhen University Medical School, Shenzhen University(深圳大学医学院生物医学工程学院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15252 2025-11-10 cs.CR cs.CL cs.IR 70%

Retrieval-Augmented Review Generation for Poisoning Recommender Systems

Shiyi Yang, Xinshu Li, Guanglin Zhou, Chen Wang, Xiwei Xu, Liming Zhu, Lina Yao

机构 * University of New South Wales and CSIRO’s Data61(新南威尔士大学和CSIRO的Data61)

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18623 2025-10-17 cs.CV 70%

Training-Free Personalization via Retrieval and Reasoning on Fingerprints

Deepayan Das, Davide Talon, Yiming Wang, Massimiliano Mancini, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯瑟实验室)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏