arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3460 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3460 篇

2511.18598 2025-11-25 q-bio.OT 50%

Assessing Gaze and Pointing: Human Cue Interpretation by Indian Free-Ranging Dogs in a Food Retrieval Task

评估目光与指认:印度自由放养狗在食物获取任务中的人类提示解读

Srijaya Nandi, Dipanjan Roy, Aesha Lahiri, Anamitra Roy, Anindita Bhadra

专题命中 跨模态检索 :multimodal(abstract)

AI总结 研究发现印度自由放养狗在结合指认和目光提示时能准确找到隐藏食物,但单一或冲突提示下表现无显著差异,且狗的气质影响其参与意愿和接近延迟,但不影响选择准确性。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10584 2025-11-25 cs.IR 50%

DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System

DAS: 基于双对齐语义ID的工业推荐系统

Wencai Ye, Mingjie Sun, Shaoyun Shi, Peng Wang, Wenjin Wu, Peng Jiang

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 DAS通过双对齐语义ID方法,提升推荐系统中多模态内容整合与协同信号对齐效率,有效解决信息损失与灵活性问题。

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16950 2025-11-24 physics.optics 50%

Single-Axis Ptychographic Coherent Diffractive Imaging for Spectroscopic and Wavefront Retrieval

单轴投影衍射成像用于光谱和波前恢复

Qijun You, Lingshuo Meng, Fangrui Quan, Wei Cao

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 单轴投影衍射成像技术通过单轴扫描提高通量,实现光谱和波前的同步成像,用于高效多模式实时成像。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11912 2025-11-18 cs.LG cs.CR 50%

A Systematic Study of Model Extraction Attacks on Graph Foundation Models

Haoyan Xu, Ruizhi Qian, Jiate Li, Yushun Dong, Minghao Lin, Hanson Yan, Zhengtao Yao, Qinghua Liu, Junhao Dong, Ruopeng Huang, Yue Zhao, Mengyuan Li

机构 * University of Southern California(南加州大学) Florida State University(佛罗里达州立大学) The Ohio State University(俄亥俄州立大学) Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14772 2025-11-12 cs.HC 50%

UMind: A Unified Multitask Network for Zero-Shot M/EEG Visual Decoding

Chengjian Xu, Yonghao Song, Zelin Liao, Haochuan Zhang, Qiong Wang, Qingqing Zheng

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25718 2025-10-30 cs.IR cs.DL 50%

Retrieval-Augmented Search for Large-Scale Map Collections with ColPali

Jamie Mahowald, Benjamin Charles Germain Lee

专题命中 跨模态检索 :multimodal(abstract)

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22613 2025-10-28 cs.SE 50%

DynaCausal: Dynamic Causality-Aware Root Cause Analysis for Distributed Microservices

Songhan Zhang, Aoyang Fang, Yifan Yang, Ruiyi Cheng, Xiaoying Tang, Pinjia He

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07759 2025-10-28 cs.IR 50%

A Survey of Long-Document Retrieval in the PLM and LLM Era

Minghan Li, Miyang Luo, Tianrui Lv, Yishuai Zhang, Siqi Zhao, Ercong Nie, Guodong Zhou

专题命中 跨模态检索 :multimodal(abstract)

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16803 2025-10-21 cs.IR 50%

An Efficient Framework for Whole-Page Reranking via Single-Modal Supervision

Zishuai Zhang, Sihao Yu, Wenyi Xie, Ying Nie, Junfeng Wang, Zhiming Zheng, Dawei Yin, Hainan Zhang

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13406 2025-10-16 cs.LG 50%

When Embedding Models Meet: Procrustes Bounds and Applications

Lucas Maystre, Alvaro Ortega Gonzalez, Charles Park, Rares Dolga, Tudor Berariu, Yu Zhao, Kamil Ciosek

机构 * UiPath Spotify

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12213 2025-10-15 astro-ph.IM 50%

Encapsulating Textual Contents into a MOC data Structure for Advanced Applications

Giuseppe Greco, Thomas Boch, Pierre Fernique, Manon Marchand, Mark Allen, Francois Xavier Pineau, Matthieu Baumann, Marco Molinaro, Roberto De Pietri, Marica Branchesi, Steven Schramm, Gergely Dalya, Elahe Khalouei, Barbara Patricelli, Giulia Stratta

专题命中 跨模态检索 :multimodal(abstract)

Comments Published in Astronomy and Computing; 11 pages, 4 figures

Journal ref Astron. Comput. 54 (2026) 101014

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01853 2025-10-07 cs.LG cs.LO 50%

Learning Representations Through Contrastive Neural Model Checking

Vladimir Krsmanovic, Matthias Cosler, Mohamed Ghanem, Bernd Finkbeiner

机构 * CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全研究中心)

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23543 2025-09-30 q-bio.GN cs.NE q-bio.MN 50%

Contrastive Learning Enhances Language Model Based Cell Embeddings for Low-Sample Single Cell Transcriptomics

Luxuan Zhang, Douglas Jiang, Qinglong Wang, Haoqi Sun, Feng Tian

专题命中 跨模态检索 :multimodal(abstract)

Comments 14 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21507 2025-09-29 cs.CE 50%

QuantMind: A Context-Engineering Based Knowledge Framework for Quantitative Finance

Haoxue Wang, Keli Wen, Yuante Li, Qiancheng Qu, Xiangxu Mu, Xinjie Shen, Jiaqi Gao, Chenyang Chang, Chuhan Xie, San Yu Cheung, Zhuoyuan Hu, Xinyu Wang, Sirui Bi, Bi'an Du

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14031 2025-09-22 cs.NE cs.LG 50%

Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms

Shreya Saha, Ishaan Chadha, Meenakshi Khosla

机构 * Electrical and Computer Engineering University of California, San Diego(电气与计算机工程大学加州大学圣地亚哥分校) Halıcıoğlu Data Science Institute University of California, San Diego(Halıcıoğlu数据科学研究所大学加州大学圣地亚哥分校) Department of Cognitive Science, Department of Computer Science and Engineering University of California, San Diego(认知科学系计算机科学与工程系大学加州大学圣地亚哥分校)

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13326 2025-09-18 cs.HC cs.LG 50%

LLM Chatbot-Creation Approaches

Hemil Mehta, Tanvi Raut, Kohav Yadav, Edward F. Gehringer

专题命中 跨模态检索 :multimodal(abstract)

Comments Forthcoming in Frontiers in Education (FIE 2025), Nashville, Tennessee, USA, Nov 2-5, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12824 2025-09-18 cs.IR 50%

DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval

Zechao Liu, Zheng Zhou, Xiangkun Chen, Tao Liang, Dapeng Lang

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04836 2025-09-08 cs.RO 50%

COMMET: A System for Human-Induced Conflicts in Mobile Manipulation of Everyday Tasks

Dongping Li, Shaoting Peng, John Pohovey, Katherine Rose Driggs-Campbell

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Department of Electrical and Computer Engineering(电气与计算机工程系) ZJU-UIUC Institute(浙大-伊利诺伊大学联合学院)

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09647 2025-09-08 cs.RO 50%

Object Instance Retrieval in Assistive Robotics: Leveraging Fine-Tuned SimSiam with Multi-View Images Based on 3D Semantic Map

Taichi Sakaguchi, Akira Taniguchi, Yoshinobu Hagiwara, Lotfi El Hafi, Shoichi Hasegawa, Tadahiro Taniguchi

机构 * Ritsumeikan University(立命馆大学) Soka University(早稻田大学) Kyoto University(京都大学)

专题命中 跨模态检索 :multimodal(abstract)

Comments See website at https://emergentsystemlabstudent.github.io/MultiViewRetrieve/. Accepted to IROS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01184 2025-09-03 cs.IR 50%

MARS: Modality-Aligned Retrieval for Sequence Augmented CTR Prediction

Yutian Xiao, Shukuan Wang, Binhao Wang, Zhao Zhang, Yanze Zhang, Shanqi Liu, Chao Feng, Xiang Li, Fuzhen Zhuang

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09128 2025-09-03 stat.ML cs.LG 50%

A Generalization Theory for Zero-Shot Prediction

Ronak Mehta, Zaid Harchaoui

机构 * University of Washington(华盛顿大学)

专题命中 跨模态检索 :multimodal(abstract)

Comments Published at ICML '25 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11535 2025-08-29 stat.ML cs.LG cs.SY eess.SY stat.CO 50%

Canonical Bayesian Linear System Identification

Andrey Bryutkin, Matthew E. Levine, Iñigo Urteaga, Youssef Marzouk

机构 * Massachusetts Institute of Technology(麻省理工学院) Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) Basis Research Institute(Basis研究机构) Eric and Wendy Schmidt Center(埃里克和温迪·施密特中心) BCAM (Basque Center for Applied Mathematics)(BCAM(巴斯克应用数学中心)) Ikerbasque (Basque Foundation for Science)(Ikerbasque(巴斯克科学基金会))

专题命中 跨模态检索 :multi-modal(abstract)

Comments 46 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19942 2025-08-28 cs.HC 50%

Socially Interactive Agents for Preserving and Transferring Tacit Knowledge in Organizations

Martin Benderoth, Patrick Gebhard, Christian Keller, C. Benjamin Nakhosteen, Stefan Schaffer, Tanja Schneeberger

专题命中 跨模态检索 :multimodal(abstract)

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14940 2025-08-27 cs.LG 50%

Cohort-Aware Agents for Individualized Lung Cancer Risk Prediction Using a Retrieval-Augmented Model Selection Framework

Chongyu Qu, Allen J. Luna, Thomas Z. Li, Junchao Zhu, Junlin Guo, Juming Xiong, Kim L. Sandler, Bennett A. Landman, Yuankai Huo

机构 * Vanderbilt University(范德堡大学) Vanderbilt University Medical Center(范德堡大学医学中心)

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11246 2025-08-18 cs.ET cs.IR 50%

RAG for Geoscience: What We Expect, Gaps and Opportunities

Runlong Yu, Shiyuan Luo, Rahul Ghosh, Lingyao Li, Yiqun Xie, Xiaowei Jia

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11210 2025-08-18 cs.LG stat.ML 50%

Borrowing From the Future: Enhancing Early Risk Assessment through Contrastive Learning

Minghui Sun, Matthew M. Engelhard, Benjamin A. Goldstein

机构 * Department of Biostatistics & Bioinformatics, Duke University(生物统计学与生物信息学系,杜克大学)

专题命中 跨模态检索 :multi-modal(abstract)

Comments accepted by Machine Learning for Healthcare 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06691 2025-08-12 cond-mat.mtrl-sci cs.LG 50%

Role of Large Language Models and Retrieval-Augmented Generation for Accelerating Crystalline Material Discovery: A Systematic Review

Agada Joseph Oche, Arpan Biswas

机构 * Bredesen Center for Interdisciplinary Research, University of Tennessee, Knoxville, USA(跨学科研究中心,田纳西大学,诺克斯维尔,美国) University of Tennessee-Oak Ridge Innovation Institute, University of Tennessee, Knoxville, USA(田纳西大学橡树岭创新研究所,田纳西大学,诺克斯维尔,美国)

专题命中 跨模态检索 :multi-modal(abstract)

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02523 2025-08-05 cs.CR 50%

Transportation Cyber Incident Awareness through Generative AI-Based Incident Analysis and Retrieval-Augmented Question-Answering Systems

Ostonya Thomas, Muhaimin Bin Munir, Jean-Michel Tine, Mizanur Rahman, Yuchen Cai, Khandakar Ashrafi Akbar, Md Nahiyan Uddin, Latifur Khan, Trayce Hockstad, Mashrur Chowdhury

专题命中 跨模态检索 :multimodal(abstract)

Comments This paper has been submitted to the Transportation Research Board (TRB) for consideration for presentation at the 2026 Annual Meeting

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00513 2025-08-04 cs.LG 50%

Text-Attributed Graph Anomaly Detection via Multi-Scale Cross- and Uni-Modal Contrastive Learning

Yiming Xu, Xu Hua, Zhen Peng, Bin Shi, Jiarun Chen, Xingbo Fu, Song Wang, Bo Dong

机构 * School of Computer Science and Technology, Xi'an Jiaotong University(西安交通大学计算机科学与技术学院) Shaanxi Provincial Key Laboratory of Big Data Knowledge Engineering, Xi’an Jiaotong University(陕西省大数据知识工程重点实验室) School of Distance Education, Xi’an Jiaotong University(西安交通大学继续教育学院) University of Virginia(弗吉尼亚大学)

专题命中 跨模态检索 :cross-modal(abstract)

Comments Accepted by ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16133 2025-07-31 cs.GR 50%

Multi-Prompt Style Interpolation for Fine-Grained Artistic Control

Lei Chen, Hao Li, Yuxin Zhang, Chao Li, Kai Wen

专题命中 跨模态检索 :cross-modal(abstract)

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏