arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2511.05363 2025-11-10 cs.CY cs.AI 57%

AI Literacy for Community Colleges: Instructors' Perspectives on Scenario-Based and Interactive Approaches to Teaching AI

Aparna Maya Warrier, Arav Agarwal, Jaromir Savelka, Christopher A Bogart, Heather Burte

机构 * Computer Science(计算机科学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17899 2025-11-07 cs.CV 57%

What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images

Dongheng Lin, Han Hu, Jianbo Jiao

机构 * University of Birmingham(伯明翰大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02996 2025-11-06 cs.CV 57%

SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics

Ailar Mahdizadeh, Puria Azadi Moghadam, Xiangteng He, Shahriar Mirabbasi, Panos Nasiopoulos, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02770 2025-11-05 cs.CL cs.IR 57%

Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval

Hung-Ting Chen, Xiang Liu, Shauli Ravfogel, Eunsol Choi

机构 * Department of Computer Science, New York University(纽约大学计算机科学系) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21738 2025-11-04 cs.CY cs.CV 57%

From Drone Imagery to Livability Mapping: AI-powered Environment Perception in Rural China

Weihuan Deng, Yaofu Huang, Luan Chen, Xun Li, Yu Gu, Yao Yao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00925 2025-11-04 cs.CV 57%

Dynamic Multi-level Weighted Alignment Network for Zero-shot Sketch-based Image Retrieval

Hanwen Su, Ge Song, Jiyan Wang, Yuanbo Zhu

机构 * School of Computer and Electronic Information(计算机与电子信息学院) Nanjing Normal University(南京师范大学) Nanjing(南京) CHN(中国)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22998 2025-10-28 cs.AI 57%

ProfileXAI: User-Adaptive Explainable AI

Gilber A. Corrales, Carlos Andrés Ferro Sánchez, Reinel Tabares-Soto, Jesús Alfonso López Sotelo, Gonzalo A. Ruz, Johan Sebastian Piña Durán

机构 * Facultad de Ingeniería y Ciencias Básicas, Universidad Autónoma de Occidente(工程与基础科学学院,自治大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments pages, 1 figure, 3 tables. Preprint. Evaluated on UCI Heart Disease (1989) and UCI Differentiated Thyroid Cancer Recurrence (2023). Uses IEEEtran

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22937 2025-10-28 cs.CV cs.LG 57%

Bi-Encoder Contrastive Learning for Fingerprint and Iris Biometrics

Matthew So, Judah Goldfeder, Mark Lis, Hod Lipson

机构 * Department of Computer Science(计算机科学系) Columbia University(哥伦比亚大学) College of Medicine(医学院) SUNY Downstate Health Sciences University(SUNY 下州健康科学大学) Department of Mechanical Engineering(机械工程系)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21717 2025-10-28 cs.HC cs.AI cs.SE 57%

AI-Enhanced Operator Assistance for UNICOS Applications

Bernard Tam, Jean-Charles Tournier, Fernando Varela Rodriguez

机构 * The University of Sydney(悉尼大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments Prepared as part of the CERN openlab programme 2025. Also available on Zenodo, a repository operated by CERN and co-funded by the European Union

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11293 2025-10-27 cs.CV 57%

Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining

Raghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K, Mingyi Su, Ping Nie, Semih Yavuz, Yingbo Zhou, Wenhu Chen, Bhuwan Dhingra

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19398 2025-10-23 cs.CL 57%

SONAR-SLT: Multilingual Sign Language Translation via Language-Agnostic Sentence Embedding Supervision

Yasser Hamidullah, Shakib Yazdani, Cennet Oguz, Josef van Genabith, Cristina España-Bonet

机构 * German Research Center for Artificial Intelligence (DFKI GmbH)(德国人工智能研究中心(DFKI GmbH)) Saarland Informatics Campus(萨尔兰信息技术校区) Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心(BSC-CNS))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Journal ref published at WMT2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04039 2025-10-22 cs.CR cs.AI 57%

BlockScan: Detecting Anomalies in Blockchain Transactions

Jiahao Yu, Xian Wu, Hao Liu, Wenbo Guo, Xinyu Xing

机构 * UC Santa Barbara(加州大学圣芭芭拉分校) Meta AI New York University(纽约大学) sec3 Northwestern University(西北大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20322 2025-10-21 cs.CV cs.LG 57%

Fine-Grained Classification: Connecting Metadata via Cross-Contrastive Pre-Training

Sumit Mamtani, Yash Thesia

机构 * New York University(纽约大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 4 figures. Accepted at IEEE ISCMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14965 2025-10-17 cs.CV 57%

ChangingGrounding: 3D Visual Grounding in Changing Scenes

Miao Hu, Zhiwei Huang, Tai Wang, Jiangmiao Pang, Dahua Lin, Nanning Zheng, Runsen Xu

机构 * Xi’an Jiaotong University(西安交通大学) Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12474 2025-10-15 cs.CL cs.LG 57%

SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression

Biao Zhang, Lixin Chen, Tong Liu, Bo Zheng

机构 * Taobao & Tmall Group of Alibaba(淘宝与天猫集团)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11204 2025-10-14 cs.CV 57%

Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos

Rohit Gupta, Anirban Roy, Claire Christensen, Sujeong Kim, Sarah Gerard, Madeline Cincebeaux, Ajay Divakaran, Todd Grindal, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) SRI International(SRI国际)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Published at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10787 2025-10-14 cs.CL 57%

Review of Inference-Time Scaling Strategies: Reasoning, Search and RAG

Zhichao Wang, Cheng Wan, Dong Nie

机构 * Inflection AI Georgia Institute of Technology(佐治亚理工学院) ChatAlpha AI

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05970 2025-10-14 cs.CV 57%

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

Haiwen Li, Delong Liu, Zhaohui Hou, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments This paper was originally submitted to ACM MM 2025 on April 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09586 2025-10-13 cs.CV 57%

Vision Language Models: A Survey of 26K Papers

Fengming Lin

机构 * School of Computer Science, The University of Manchester, Manchester, UK(曼彻斯特大学计算机科学学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments VLM/LLM Learning Notes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08578 2025-10-13 cs.MA cs.AI cs.HC 57%

AgenticAD: A Specialized Multiagent System Framework for Holistic Alzheimer Disease Management

Adib Bazgir, Amir Habibdoust, Xing Song, Yuwen Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20444 2025-10-09 cs.LG cs.CV 57%

HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models

Haoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu, Ruiwen Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Xiaohongshu Inc.(小红书公司)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04396 2025-10-07 cs.MM cs.IR 57%

Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections

Bastian Jäckl, Jiří Kruchina, Lucas Joos, Daniel A. Keim, Ladislav Peška, Jakub Lokoč

专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM

Comments 28 Pages, 17 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03614 2025-10-07 cs.LG cs.AI stat.ML 57%

Neural Bayesian Filtering

Christopher Solinas, Radovan Haluska, David Sychrovsky, Finbarr Timbers, Nolan Bard, Michael Buro, Martin Schmid, Nathan R. Sturtevant, Michael Bowling

机构 * University of Alberta(阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所) Charles University(查尔斯大学) EquiLibre Technologies, Inc.(EquiLibre技术公司) Allen Institute for AI(艾伦人工智能研究所) Sony AI(索尼人工智能)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02605 2025-10-07 cs.CV 57%

ReMoMask: Retrieval-Augmented Masked Motion Generation

Zhengdao Li, Siheng Wang, Zeyu Zhang, Hao Tang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03200 2025-10-06 cs.CV 57%

MonSTeR: a Unified Model for Motion, Scene, Text Retrieval

Luca Collorone, Matteo Gioia, Massimiliano Pappa, Paolo Leoni, Giovanni Ficarra, Or Litany, Indro Spinelli, Fabio Galasso

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Technion, NVIDIA(技术学院与NVIDIA) WSense

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26012 2025-10-01 cs.CV 57%

SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval

Yuqi Xiao, Yingying Zhu

机构 * Yuqi Xiao, Yingying Zhu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24370 2025-09-30 cs.CV 57%

DINOReg: Strong Point Cloud Registration with Vision Foundation Model

Congjia Chen, Yufu Qu

机构 * School of Instrumentation and Optoelectronic Engineering(仪器与光电工程学院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23242 2025-09-30 cs.CV 57%

TATTOO: Training-free AesTheTic-aware Outfit recOmmendation

Yuntian Wu, Xiaonan Hu, Ziqi Zhou, Hao Lu

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21336 2025-09-29 cs.IR cs.CL 57%

HetaRAG: Hybrid Deep Retrieval-Augmented Generation across Heterogeneous Data Stores

Guohang Yan, Yue Zhang, Pinlong Cai, Ding Wang, Song Mao, Hongwei Zhang, Yaoze Zhang, Hairong Zhang, Xinyu Cai, Botian Shi

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21259 2025-09-26 cs.NI cs.AI 57%

Semantic Edge-Cloud Communication for Real-Time Urban Traffic Surveillance with ViT and LLMs over Mobile Networks

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏