arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2505.23809 2025-06-04 cs.CL cs.AI cs.IR 62%

LLM-Driven E-Commerce Marketing Content Optimization: Balancing Creativity and Conversion

Haowei Yang, Haotian Lyu, Tianle Zhang, Dingzhou Wang, Yushang Zhao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01388 2025-06-03 cs.CV cs.AI 62%

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

Yihao Ding, Soyeon Caren Han, Yan Li, Josiah Poon

机构 * The University of Melbourne(墨尔本大学) The University of Sydney(悉尼大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at IJCAI 2025 Demonstrations Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23045 2025-05-30 cs.CV cs.AI 62%

Multi-Sourced Compositional Generalization in Visual Question Answering

Chuanhao Li, Wenbo Ye, Zhen Li, Yuwei Wu, Yunde Jia

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学,中国) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,深圳MSU-BIT大学,中国)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07688 2025-05-29 cs.CV cs.AI 62%

ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG

Zilun Zhang, Haozhan Shen, Tiancheng Zhao, Zian Guan, Bin Chen, Yuhao Wang, Xu Jia, Yuxiang Cai, Yongheng Shang, Jianwei Yin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究院) School of Software Engineering of Zhejiang University(浙江大学软件工程学院) Polytechnic Institute of Zhejiang University(浙江大学 polytechnic 院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Geoscience and Remote Sensing Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15876 2025-05-29 cs.CV cs.AI cs.LG 62%

End-to-End Breast Cancer Radiotherapy Planning via LMMs with Consistency Embedding

Kwanyoung Kim, Yujin Oh, Sangjoon Park, Hwa Kyung Byun, Joongyo Lee, Jin Sung Kim, Yong Bae Kim, Jong Chul Ye

机构 * Samsung Research(三星研究院) Center for Advanced Medical Computing and Analysis (CAMCA)(先进医学计算与分析中心) Massachusetts General Hospital (MGH) and Harvard Medical School(麻省总医院和哈佛医学院) Yonsei University College of Medicine(延世大学医学院) Yonsei University(延世大学) Institute for Innovation in Digital Healthcare(数字医疗创新研究所) Yongin Severance Hospital(Yongin Severance医院) Gachon University Gil Hospital(高仁大学Gil医院) Oncosoft Inc(Oncosoft公司) Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST)(人工智能研究生院,韩国科学技术院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted for Medical Image Analysis 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17796 2025-05-26 cs.CV cs.AI cs.IR 62%

DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval

Yuxin Yang, Yinan Zhou, Yuxin Chen, Ziqi Zhang, Zongyang Ma, Chunfeng Yuan, Bing Li, Lin Song, Jun Gao, Peng Li, Weiming Hu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17085 2025-05-26 cs.CR cs.AI cs.CL 62%

GSDFuse: Capturing Cognitive Inconsistencies from Multi-Dimensional Weak Signals in Social Media Steganalysis

Kaibo Huang, Zipei Zhang, Yukun Wei, TianXin Zhang, Zhongliang Yang, Linna Zhou

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing IntokenTech Co., Ltd.(北京IntokenTech公司)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17058 2025-05-26 cs.CL cs.AI 62%

DO-RAG: A Domain-Specific QA Framework Using Knowledge Graph-Enhanced Retrieval-Augmented Generation

David Osei Opoku, Ming Sheng, Yong Zhang

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Beijing National Research Center for Information Science and Technology - Tsinghua University(北京信息科学与技术国家研究中心-清华大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 5 figures;

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16193 2025-05-23 cs.CL cs.CV 62%

An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability

Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01222 2025-05-23 cs.CV cs.CL 62%

Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG

Wenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang, Li Shen, Yong Luo, Bo Du, Dacheng Tao

机构 * Wuhan University(武汉大学) Nanyang Technological University(南洋理工大学) The University of Sydney(悉尼大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09243 2025-05-19 cs.RO cs.AI cs.CV 62%

GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation

Ruihai Wu, Ziyu Zhu, Yuran Wang, Yue Chen, Jiarui Wang, Hao Dong

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03242 2025-05-07 cs.CV cs.AI 62%

Seeing the Abstract: Translating the Abstract Language for Vision Language Models

Davide Talon, Federico Girella, Ziyue Liu, Marco Cristani, Yiming Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR25. Project page: https://davidetalon.github.io/fashionact-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12309 2025-04-21 cs.CY cs.AI cs.CL cs.IR 62%

Large Language Model-Based Knowledge Graph System Construction for Sustainable Development Goals: An AI-Based Speculative Design Perspective

Yi-De Lin, Guan-Ze Liao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This is a minor revision: fixed a typo in the abstract (time range) and corrected minor textual errors

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02064 2025-04-18 cs.CV cs.AI 62%

ArtCrafter: Text-Image Aligning Style Transfer via Embedding Reframing

Nisha Huang, Kaer Huang, Yifan Pu, Jiangshan Wang, Jie Guo, Yiqiang Yan, Xiu Li, Tong-Yee Lee

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages, 17 figures, submitted to a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05316 2025-04-09 cs.IR cs.AI cs.CV 62%

Scale Up Composed Image Retrieval Learning via Modification Text Generation

Yinan Zhou, Yaxiong Wang, Haokun Lin, Chen Ma, Li Zhu, Zhedong Zheng

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07766 2025-04-09 cs.CL cs.AI 62%

Large Language Models for Knowledge Graph Embedding: A Survey

Bingchen Liu, Yuanyuan Fang, Naixing Xu, Shihao Hou, Xin Li, Qian Li

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21309 2025-03-28 cs.CV cs.AI 62%

FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval

Zixu Li, Zhiheng Fu, Yupeng Hu, Zhiwei Chen, Haokun Wen, Liqiang Nie

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19296 2025-03-26 cs.CV cs.MM 62%

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

Haoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu, Yupeng Hu, Liqiang Nie

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09387 2025-03-25 cs.CV cs.AI cs.LG 62%

RankCLIP: Ranking-Consistent Language-Image Pretraining

Yiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng, Zenghui Ding, Yining Sun

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Code and model checkpoints are available at https://github.com/Jam1ezhang/RankCLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17871 2025-03-25 cs.CV cs.AI 62%

good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval

Pranavi Kolouju, Eric Xing, Robert Pless, Nathan Jacobs, Abby Stylianou

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14733 2025-03-24 cs.LG cs.AI cs.CL 62%

Knowledge Graph Embeddings: A Comprehensive Survey on Capturing Relation Properties

Guanglin Niu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 8 figures, 3 tables, this paper is a modified English version of our article already published in Computer Science journal (in Chinese), released to facilitate communication among international researchers in the relevant fields

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08722 2025-03-18 cs.CV cs.AI cs.LG 62%

A Recipe for Improving Remote Sensing VLM Zero Shot Generalization

Aviad Barzilai, Yotam Gigi, Amr Helmy, Vered Silverman, Yehonathan Refael, Bolous Jaber, Tomer Shekel, George Leifman, Genady Beryozkin

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04287 2025-03-07 cs.CV cs.AI 62%

MARS: Paying more attention to visual attributes for text-based person search

Alex Ergasti, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00793 2025-03-04 cs.CV cs.AI cs.RO 62%

Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning

Ukcheol Shin, Kyunghyun Lee, Jean Oh

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICRA 2025, Github link: https://github.com/UkcheolShin/BridgeMultiSpectralDepth

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00361 2025-03-04 cs.CV cs.AI 62%

Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

Wei Suo, Lijun Zhang, Mengyang Sun, Lin Yuanbo Wu, Peng Wang, Yanning Zhang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16641 2025-02-25 cs.CV cs.CL cs.IR 62%

Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines

Xinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang, Biqing Qi, Bowen Zhou

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments AAAI-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11301 2025-02-24 cs.CL cs.AI 62%

Question-to-Question Retrieval for Hallucination-Free Knowledge Access: An Approach for Wikipedia and Wikidata Question Answering

Santhosh Thottingal

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14113 2025-02-21 cs.CV cs.AI 62%

Object-centric Binding in Contrastive Language-Image Pretraining

Rim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal, Adriana Romero-Soriano

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10735 2025-02-11 cs.AI cs.CL 62%

Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning

Kuofeng Gao, Huanqia Cai, Qingyao Shuai, Dihong Gong, Zhifeng Li

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06224 2025-02-07 cs.CV cs.AI 62%

Detection, Retrieval, and Explanation Unified: A Violence Detection System Based on Knowledge Graphs and GAT

Wen-Dong Jiang, Chih-Yung Chang, Diptendu Sinha Roy

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏