arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2510.05016 2025-10-08 astro-ph.IM cs.AI cs.CL 62%

Large Language Models Achieve Gold Medal Performance at the International Olympiad on Astronomy & Astrophysics (IOAA)

Lucas Carrit Delgado Pinheiro, Ziru Chen, Bruno Caixeta Piazza, Ness Shroff, Yingbin Liang, Yuan-Sen Ting, Huan Sun

机构 * The Ohio State University(俄亥俄州立大学) Universidade de São Paulo(圣保罗大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 18 pages, 6 figures, to be submitted, comments are welcome. Reproducibility details can be found at: https://github.com/OSU-NLP-Group/LLM-IOAA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02328 2025-10-06 cs.CL cs.AI cs.MA 62%

AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering

Ziqing Wang, Chengsheng Mao, Xiaole Wen, Yuan Luo, Kaize Ding

机构 * Northwestern University(西北大学) Microsoft(微软公司)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24823 2025-09-30 cs.CR cs.AI cs.CV cs.LG 62%

Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated images with arbitrary size

Benedetta Tondi, Andrea Costanzo, Mauro Barni

机构 * University of Siena(锡耶纳大学)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21356 2025-09-29 cs.CV cs.AI 62%

Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports

Razi Mahmood, Diego Machado-Reyes, Joy Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep Kalra, Ge Wang, Pingkun Yan, Tanveer Syeda-Mahmood

机构 * Rensselaer Polytechnic Institute, NY, USA(罗文学院) IBM Research, Almaden, CA, USA(IBM研究院) Stanford University, CA, USA(斯坦福大学) Massachusetts General Hospital (MGH), Boston, USA(麻省总医院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments In proceedings MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06290 2025-09-25 cs.LG cs.AI cs.CV 62%

CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning

Mingyu Lu, Ethan Weinberger, Chanwoo Kim, Su-In Lee

机构 * Paul G. Allen School of Computer Science & Engineering University of Washington(保罗·G·艾伦计算机科学与工程学院华盛顿大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05930 2025-09-24 cs.CL cs.CV cs.LG 62%

WebLINX: Real-World Website Navigation with Multi-Turn Dialogue

Xing Han Lù, Zdeněk Kasner, Siva Reddy

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17802 2025-09-23 cs.CV cs.AI 62%

TS-P$^2$CL: Plug-and-Play Dual Contrastive Learning for Vision-Guided Medical Time Series Classification

Qi'ao Xu, Pengfei Wang, Bo Zhong, Tianwen Qian, Xiaoling Wang, Ye Wang, Hong Yu

机构 * East China Normal University(东华大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11587 2025-09-16 cs.CV cs.AI 62%

Hierarchical Identity Learning for Unsupervised Visible-Infrared Person Re-Identification

Haonan Shi, Yubin Wang, De Cheng, Lingfeng He, Nannan Wang, Xinbo Gao

机构 * IEEE Publication Technology Department(IEEE出版技术部门) State Key Laboratory of Integrated Services Networks, School of Telecommunications Engineering, Xidian University(信息服务网络国家重点实验室,电信工程学院,西安电子科技大学) Department of Computer Science and Technology, Tongji University(计算机科学与技术系,同济大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08624 2025-09-11 cs.CV cs.AI 62%

UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation

Zhihao Zhao, Yinzheng Zhao, Junjie Yang, Xiangtong Yao, Quanmin Liang, Daniel Zapp, Kai Huang, Nassir Navab, M. Ali Nasseri

机构 * Technical University of Munich(慕尼黑技术大学) Sun Yat-Sen University(中山大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments BIBM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10559 2025-09-10 cs.CV cs.AI cs.ET 62%

From Images to Insights: Explainable Biodiversity Monitoring with Plain Language Habitat Explanations

Yutong Zhou, Masahiro Ryo

机构 * Leibniz Centre for Agricultural Landscape Research (ZALF)(莱比锡农业景观研究中心(ZALF)) Brandenburg University of Technology Cottbus–Senftenberg(勃兰登堡技术大学科特博斯-森芬根堡)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments AISE workshop camera-ready version @ ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06020 2025-09-08 cs.AI cs.CV 62%

ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding

Shuai Wang, Ivona Najdenkoska, Hongyi Zhu, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学) College of Business(商学院) Economics, University of Johannesburg(经济系,约翰内斯堡大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11452 2025-09-03 cs.AI cs.CL cs.HC 62%

Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps

Kangyu Wang, Hongliang He, Lin Liu, Ruiqi Liang, Zhenzhong Lan, Jianguo Li

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Westlake University(西湖大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Our platform is publicly accessible at https://www.tbox.cn/about/model-ranking

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09040 2025-08-26 cs.RO cs.AI cs.CV cs.LG 62%

RT-Cache: Training-Free Retrieval for Real-Time Manipulation

Owen Kwon, Abraham George, Alison Bartsch, Amir Barati Farimani

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures. 2025 IEEE-RAS 24th International Conference on Humanoid Robots

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09545 2025-08-26 cs.CL cs.AI cs.CY 62%

Does GPT-4 surpass human performance in linguistic pragmatics?

Ljubisa Bojic, Predrag Kovacevic, Milan Cabarkapa

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16636 2025-08-18 cs.CL cs.CV 62%

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Yin Wu, Quanyu Long, Jing Li, Jianfei Yu, Wenya Wang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 21 pages, 6 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00589 2025-08-13 cs.CV cs.CL cs.IR cs.RO 62%

Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving

Stefan Englmeier, Max A. Büttner, Katharina Winter, Fabian B. Flohr

机构 * Munich University of Applied Sciences(慕尼黑应用科学大学) Intelligent Vehicles Lab (IVL)(智能车辆实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Project page: https://iv.ee.hm.edu/contextmotionclip/; This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03644 2025-08-06 cs.CL cs.CV cs.IR 62%

Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?

Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen, Junjie Yang, Yao Wan, Weiwei Lin

机构 * South China University of Technology(华南理工大学) Huazhong University of Science and Technology(华中科技大学) University of Maryland(马里兰大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments In submission. Project website: https://double-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03262 2025-08-06 cs.CL cs.AI 62%

Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?

Junhyuk Choi, Hyeonchu Park, Haemin Lee, Hyebeen Shin, Hyun Joung Jin, Bugeun Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20110 2025-07-29 cs.CV cs.AI cs.LG 62%

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

Shiyu Liu, Lianlei Shan

机构 * School of Electrical and Electronic Engineering(电气与电子工程学院) Nanyang Technological University(南洋理工大学) School of Computer Science and Technology(计算机科学与技术学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments **14 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15639 2025-07-22 cs.CL cs.AI cs.LG 62%

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

Anirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina

专题命中 跨模态检索 :MLLM(abstract);分类 cs.CL、cs.AI

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12425 2025-07-17 cs.CL cs.AI cs.CE cs.IR 62%

Advancing Retrieval-Augmented Generation for Structured Enterprise and Internal Data

Chandana Cheerla

机构 * IIT Roorkee(罗尔基大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09482 2025-07-15 cs.CL cs.AI cs.HC 62%

ViSP: A PPO-Driven Framework for Sarcasm Generation with Contrastive Learning

Changli Wang, Rui Wu, Fang Yin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07748 2025-07-11 cs.CL cs.AI 62%

When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance

Peizhang Shao, Linrui Xu, Jinxi Wang, Wei Zhou, Xingyu Wu

机构 * School of Law, China University of Political Science and Law(中国政法大学法学院) Zhejiang University of Finance and Economics Dongfang College(浙江财经大学东洋学院) School of Information Management for Law, China University of Political Science and Law(中国政法大学信息管理法学院) Department of Artificial Intelligence, Chung-Ang University(成均馆大学人工智能系) Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University(香港理工大学数据科学与人工智能系)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15379 2025-07-11 cs.IR cs.AI cs.CV 62%

Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval

Zijun Long, Kangheng Liang, Gerardo Aragon-Camarasa, Richard Mccreadie, Paul Henderson

机构 * Hunan University(湖南大学) University of Glasgow(格拉斯哥大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05513 2025-07-09 cs.CV cs.AI 62%

Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model

Mengyao Xu, Gabriel Moreira, Ronay Ak, Radek Osmulski, Yauhen Babakhin, Zhiding Yu, Benedikt Schifferer, Even Oldridge

机构 * Proceedings of Conference XXX(会议论文集)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20964 2025-07-09 cs.CV cs.AI 62%

Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering

Zhengxuan Zhang, Yin Wu, Yuyu Luo, Nan Tang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19834 2025-06-18 cs.LG cs.CV cs.MM 62%

Knowledge Bridger: Towards Training-free Missing Modality Completion

Guanzhou Ke, Shengfeng He, Xiao Li Wang, Bo Wang, Guoqing Chao, Yuanyang Zhang, Yi Xie, HeXing Su

机构 * Beijing Jiaotong University(北京交通大学) Singapore Management University(新加坡国立大学) Southeast University(东南大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Harbin Institute of Technology(哈尔滨工业大学) Nanjing University of Science and Technology(南京理工大学) South China University of Technology(华南理工大学) Xiamen Institute of Technology(厦门理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07785 2025-06-10 cs.CV cs.AI cs.LG 62%

Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger

Qi Yang, Chenghao Zhang, Lubin Fan, Kun Ding, Jieping Ye, Shiming Xiang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences, China(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences, China(中国科学院自动化研究所MAIS部) Alibaba Cloud Computing, China(阿里巴巴云计算)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICML 2025 Spotlight. 22 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16809 2025-06-10 cs.CV cs.MM 62%

Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities

Junze Wang, Lei Fan, Weipeng Jing, Donglin Di, Yang Song, Sidong Liu, Cong Cong

机构 * College of Computer and Control Engineering, Northeast Forestry University(东北林业大学计算机与控制工程学院) The Centre for Healthy Brain Ageing (CHeBA), UNSW(健康脑年龄中心(CHeBA),UNSW) School of Computer Science and Engineering, UNSW(计算机科学与工程学院,UNSW) School of Software, Tsinghua University(清华大学软件学院) Centre for Health Informatics, Macquarie University(健康信息学中心,麦考瑞大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments MICCAI 2025 Early Accept. The code is available at https://github.com/reeive/ReHyDIL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06208 2025-06-09 cs.CL cs.AI 62%

Building Models of Neurological Language

Henry Watkins

机构 * UCL Queen Square Institute of Neurology(伦敦大学学院女王广场神经病学研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏