arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2511.05020 2025-12-19 cs.GR cs.CV 57%

DAFM: Dynamic Adaptive Fusion for Multi-Model Collaboration in Composed Image Retrieval

DAFM:动态自适应融合用于复合图像检索中的多模型协作

Yawei Cai, Jiapeng Mi, Nan Ji, Haotian Rong, Yawei Zhang, Zhangti Li, Wenbin Guo, Rensong Xie

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 DAFM通过动态自适应融合多模型优势,提升复合图像检索的准确性和鲁棒性。

Comments We discovered an error that affects the main conclusions, so we decided to withdraw the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16674 2025-12-18 cs.CV 57%

FitPro: A Zero-Shot Framework for Interactive Text-based Pedestrian Retrieval in Open World

FitPro: 一个用于开放世界中基于文本的交互行人检索的零样本框架

Zengli Luo, Canlong Zhang, Xiaochun Lu, Zhixin Li

机构 * Key Lab of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University(教育区块链与智能技术重点实验室,教育部,广西师范大学) Guangxi Key Lab of Multi-source Information Mining & Security, Guangxi Normal University(广西多源信息挖掘与安全重点实验室,广西师范大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 FitPro通过引入特征对比解码、增量语义挖掘和查询感知分层检索,提升开放世界中基于文本的交互行人检索的语义理解和跨场景适应能力。

Comments 12pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14878 2025-12-18 cs.CV 57%

Visual-textual Dermatoglyphic Animal Biometrics: A First Case Study on Panthera tigris

视觉-文本皮肤纹路动物生物特征:对豹属(Panthera tigris)的首次案例研究

Wenshuo Li, Majid Mirmehdi, Tilo Burghardt

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本研究通过引入皮肤纹路文本描述符,提出了一种视觉-文本结合的动物Re-ID方法,提升跨模态身份检索的准确性并缓解数据稀缺问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01115 2025-12-17 cs.AI cs.MA econ.TH physics.soc-ph 57%

Exploring Network-Knowledge Graph Duality: A Case Study in Agentic Supply Chain Risk Analysis

探索网络-知识图谱二元性:代理供应链风险分析的案例研究

Evan Heus, Rick Bookstaber, Dhruv Sharma

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出基于LLM的代理框架,利用网络与知识图谱的二元性,通过图遍历和上下文壳技术实现供应链风险分析的实时可解释生成。

Comments Accepted to the 2nd Workshop on LLMs and Generative AI in Finance: International Conference on AI in Finance(ICAIF) 2025;7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12309 2025-12-16 cs.CV 57%

WeDetect: Fast Open-Vocabulary Object Detection as Retrieval

WeDetect:快速开放词汇物体检测作为检索

Shenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu, Xiaohua Xie, Wei-Shi Zheng

机构 * WeChat Vision, Tencent Inc.(腾讯微信视觉部) Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔)) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 WeDetect通过检索框架实现快速开放词汇物体检测,结合双塔架构和LMMs,提升检测效率和通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13102 2025-12-16 cs.CV 57%

CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation

CapeNext: 重新思考和细化用于类别无关姿态估计的动态支持信息

Yu Zhu, Dan Zeng, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 CapeNext通过整合层次交叉模态交互与双流特征细化,提升了类别无关姿态估计中动态支持信息的鲁棒性和鉴别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10098 2025-12-13 cs.LG cs.AI 57%

MedXAI: A Retrieval-Augmented and Self-Verifying Framework for Knowledge-Guided Medical Image Analysis

MedXAI: 一种结合检索增强和自我验证的框架,用于知识引导的医学图像分析

Midhat Urooj, Ayan Banerjee, Farhat Shaikh, Kuntal Thakur, Sandeep Gupta

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 MedXAI通过整合深度学习与临床专家知识,提升医学图像分析的可解释性和泛化能力,尤其在稀有疾病检测中表现优异。

Comments https://cmsworkshops.com/Asilomar2025/Papers/Uploads/FinalPapers/Original/1527/20251130102314_899554_1527.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09319 2025-12-11 cs.CE cs.AI 57%

Efficiency-Aware Computational Intelligence for Resource-Constrained Manufacturing Toward Edge-Ready Deployment

面向资源受限制造的高效计算智能:为边缘化部署而努力

Qianyu Zhou

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 本研究提出高效计算框架,解决资源受限制造中的数据贫乏和物理感知问题,通过生成策略、半监督学习和边缘云协作压缩方案,提升工业部署的可靠性与效率。

Comments 2025, University of Connecticut

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08673 2025-12-10 cs.CV 57%

Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds

双分支中心-周围对比:重新思考3D点云中的对比学习

Shaofeng Zhang, Xuanqi Chen, Xiangdong Zhang, Sitong Wu, Junchi Yan

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Chinese University of Hong Kong(香港中文大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出双分支中心-周围对比框架,通过双分支输入和补丁级对比损失,提升3D点云对比学习性能,达到SOTA效果。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08330 2025-12-10 cs.CV 57%

PointDico: Contrastive 3D Representation Learning Guided by Diffusion Models

PointDico: 通过扩散模型引导的对比3D表示学习

Pengbo Li, Yiding Sun, Haozhe Cheng

机构 * International School Beijing University of Posts(国际学校 北京邮电大学) School of Software Engineering Xi'an Jiaotong University Xi'an, China(软件工程学院 西安交通大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 PointDico通过融合扩散模型和对比学习的方法,实现了3D表示学习的突破,达到了ScanObjectNN和ShapeNetPart上的新高精度。

Comments Accepted by IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07807 2025-12-09 cs.CV cs.GR 57%

Lang3D-XL: Language Embedded 3D Gaussians for Large-scale Scenes

Lang3D-XL: 语言嵌入的3D高斯用于大规模场景

Shai Krakovsky, Gal Fiebelman, Sagie Benaim, Hadar Averbuch-Elor

机构 * Tel Aviv University(特拉维夫大学) The Hebrew University of Jerusalem(耶路撒冷希伯来大学) Cornell University(康奈尔大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 Lang3D-XL通过引入语言嵌入的3D高斯表示,提升大规模场景的语义理解和交互效率,优于现有方法。

Comments Accepted to SIGGRAPH Asia 2025. Project webpage: https://tau-vailab.github.io/Lang3D-XL

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04522 2025-12-05 cs.CV 57%

Identity Clue Refinement and Enhancement for Visible-Infrared Person Re-Identification

身份线索细化与增强用于可见-红外人再识别

Guoqing Zhang, Zhun Wang, Hairui Wang, Zhonglin Ye, Yuhui Zheng

机构 * School of Computer Science, Nanjing University of Information Science and Technology(南京信息工程大学计算机学院) Key Laboratory of Social Computing and Cognitive Intelligence (Dalian University of Technology), Ministry of Education(社会计算与认知智能重点实验室(大连理工大学)) State Key Laboratory of Tibetan Intelligent Information Processing and Application, Qinghai Normal University(藏语智能信息处理与应用国家重点实验室(青海师范大学))

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出ICRE网络,通过细化和增强身份线索来提升可见-红外人再识别的性能。

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01636 2025-12-02 cs.CV 57%

Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval

联合视觉-语言空间中的生成编辑用于零样本复合图像检索

Xin Wang, Haipeng Zhang, Mang Li, Zhaohui Xia, Yueguo Chen, Yu Zhang, Chunyu Wei

机构 * Renmin University of China(中国人民大学) Alibaba Group(阿里巴巴集团)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 Fusion-Diff通过联合视觉-语言空间中的生成编辑策略,实现了高效的零样本复合图像检索,提升了多模态对齐的性能和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22850 2025-12-01 cs.CV 57%

Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding

解决证据稀疏性:面向长文档理解的代理情境工程

Keliang Liu, Zizhi Chen, Mingcheng Li, Jingqun Tang, Dingkang Yang, Lihua Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 SLEUTH通过多代理框架解决长文档理解中的证据稀疏问题,提升多模态文档处理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22470 2025-12-01 cs.CV 57%

Hybrid, Unified and Iterative: A Novel Framework for Text-based Person Anomaly Retrieval

混合、统一和迭代:一种基于文本的人员异常检索新框架

Tien-Huy Nguyen, Huu-Loc Tran, Huu-Phong Phan-Nguyen, Quang-Vinh Dinh

机构 * University of Information Technology Vietnam National University(信息技术大学越南国家大学) AI VIETNAM Lab(AI越南实验室)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

AI总结 本文提出了一种混合、统一和迭代的新型框架,通过结合局部-全局视角和统一图像-文本模型,提升基于文本的人员异常检索性能。

Comments Accepted on World Wide Web 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22253 2025-12-01 cs.IR cs.CV 57%

UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries

UNION: 一种轻量级的目标表示用于高效零样本图像引导检索与可选文本查询

Hoang-Bao Le, Allie Tran, Binh T. Nguyen, Liting Zhou, Cathal Gurrin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 UNION通过融合图像嵌入与空文本提示,实现高效零样本图像引导检索,仅需5,000训练样本即在多个基准测试中超越传统基线。

Comments Accepted at ICDM - MMSR Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20783 2025-11-27 cs.CL 57%

On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey

在预训练语言模型中通用文本嵌入的作用:综述

Meishan Zhang, Xin Zhang, Xinping Zhao, Shouzheng Huang, Baotian Hu, Min Zhang

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 国家信息与自动化研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文综述了预训练语言模型在通用文本嵌入中的作用,探讨了其架构、核心方法及未来研究方向。

Comments 45 pages, 4 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19200 2025-11-26 cs.CV 57%

Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?

现代视觉模型能否理解物体与相似物之间的差异?

Itay Cohen, Ethan Fetaya, Amir Rosenfeld

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文研究了现代视觉模型能否区分真实物体与相似物,通过构建RoLA数据集并改进CLIP模型的嵌入空间方向,提升跨模态检索和描述生成的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11916 2025-11-26 cs.CV 57%

NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition

NeuroGaze-Distill:基于脑科学的蒸馏与抑郁启发的几何先验用于鲁棒面部情绪识别

Zilin Li, Weiwei Xu, Xuanqi Zhao, Yiran Zhu

机构 * School of Information and Intelligent Science(信息与智能科学学院) Department of Computer(计算机系) North China Electric Power University (BaoDing)(华北电力大学(保定))

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 NeuroGaze-Distill通过脑科学先验和抑郁启发的几何先验提升面部情绪识别的鲁棒性,采用跨模态蒸馏框架实现无需生物信号的部署。

Comments Preprint. Vision-only deployment; EEG used to form static prototypes. Includes appendix, 7 figures and 3 tables. Considering submission to ICLR 2026. Revision note: This version corrects inaccuracies in the authors' institutional affiliations. No technical content has been modified

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24466 2025-11-25 cs.CV 57%

SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking

SA-Person: 基于文本的人检索与场景感知重排序

Yingjia Xu, Jinlin Wu, Daming Gao, Zhen Chen, Yang Yang, Min Cao, Mang Ye, Zhen Lei

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人中心) Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统(MAIS)) Wuhan University(武汉大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 SA-Person通过整合个体外观和全局场景上下文,提升基于文本的人检索准确性,提出ScenePerson-13W数据集和两阶段检索框架。

Comments 13 pages, 8 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17255 2025-11-24 cs.CV cs.IR 57%

A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback

略多一些像这个:利用视觉语言模型进行文本到图像检索的相关反馈

Bulat Khaertdinov, Mirela Popa, Nava Tintarev

机构 * Maastricht University(马斯特里赫特大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于相关反馈机制提升视觉语言模型文本到图像检索性能的方法,通过四种反馈策略在Flickr30k和COCO数据集上验证,提升了检索效果并增强了多轮检索的鲁棒性。

Comments Accepted to WACV'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08008 2025-11-20 cs.AI 57%

Combining LLM Semantic Reasoning with GNN Structural Modeling for Multi-View Multi-Label Feature Selection

Zhiqi Chen, Yuzhou Liu, Jiarui Liu, Wanfu Gao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14160 2025-11-20 cs.CL cs.LG 57%

Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * Queen Mary University of London(伦敦大学玛丽女王学院) Institut Jožef Stefan(乔泽夫·斯蒂芬研究所)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03989 2025-11-19 cs.LG cs.AI 57%

Dynamic User-controllable Privacy-preserving Few-shot Sensing Framework

Ajesh Koyatan Chathoth, Shuhao Yu, Stephen Lee

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09754 2025-11-18 cs.LG cs.AI 57%

History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting

Sarthak Khanna, Armin Berger, Muskaan Chopra, David Berghaus, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫研究所-媒体工程部门) University of Bonn - Department of Computer Science(波恩大学-计算机科学系) Lamarr Institute for Machine Learning(拉马尔机器学习研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Accepted in IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09347 2025-11-17 cs.CV 57%

FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection

Jiangyong Yu, Changyong Shu, Sifan Zhou, Zichen Yu, Xing Hu, Yan Chen, Dawei Yang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments I made an operational error. I intended to update the paper with Identifier arXiv:2502.15488, not submit a new paper with a different identifier. Therefore, I would like to withdraw the current submission and resubmit an updated version for Identifier arXiv:2502.15488

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10244 2025-11-14 cs.AI 57%

PepTriX: A Framework for Explainable Peptide Analysis through Protein Language Models

Vincent Schilling, Akshat Dubey, Georges Hattab

机构 * Center for Artificial Intelligence in Public Health Research (ZKI-PH), Robert Koch Institute(人工智能与公共健康研究所以及罗伯特·科赫研究所) Department of Mathematics and Computer Science, Free University of Berlin(数学与计算机科学系,柏林自由大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06259 2025-11-11 cs.LG cs.AI 57%

Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra

Yiwen Zhang, Keyan Ding, Yihang Wu, Xiang Zhuang, Yi Yang, Qiang Zhang, Huajun Chen

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05894 2025-11-11 cs.CV 57%

Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning

Fei Yu, Quan Deng, Shengeng Tang, Yuehua Li, Lechao Cheng

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02073 2025-11-11 cs.AI 57%

Large model retrieval enhancement framework for construction site risk identification

Jiawei Li, Chengye Yang, Yaochen Zhang, Weilin Sun, Lei Meng, Xiangxu Meng

专题命中 跨模态检索 :image-text(abstract);分类 cs.AI

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏