arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4644 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2505.24837 2025-06-02 cs.CV 79%

Zero-Shot Chinese Character Recognition with Hierarchical Multi-Granularity Image-Text Aligning

Yinglian Zhu, Haiyang Yu, Qizao Wang, Wei Lu, Xiangyang Xue, Bin Li

机构 * Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理关键实验室) School of Computer Science, Fudan University(复旦大学计算机学院)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments The first three authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22869 2025-05-30 cs.CV 79%

SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories

Alexey Gavryushin, Alexandros Delitzas, Luc Van Gool, Marc Pollefeys, Kaichun Mo, Xi Wang

机构 * ETHZ(苏黎世联邦理工学院) MPI for Informatics(信息研究所) INSAIT(国际人工智能技术研究所) KU Leuven(鲁汶大学) Microsoft(微软公司) NVIDIA(英伟达公司) TUM(慕尼黑工业大学)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13199 2025-05-28 cs.CR cs.AI 79%

Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks

Mohammad Saleh, Azadeh Tabatabaei

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref International Journal of Web Research, vol.8, no.2,pp.11-24, 2025,

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04395 2025-05-27 cs.CV cs.LG 79%

Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting

Siru Zhong, Weilin Ruan, Ming Jin, Huan Li, Qingsong Wen, Yuxuan Liang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19702 2025-05-27 cs.CV 79%

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

Minheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin, Kevin Lin, Wangmeng Zuo, Lijuan Wang

机构 * Hong Kong Polytechnic University(香港理工大学) Harbin Institute of Technology(哈尔滨工业大学) Microsoft(微软公司)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17214 2025-05-26 cs.AI 79%

MEDMKG: Benchmarking Medical Knowledge Exploitation with Multimodal Knowledge Graph

Xiaochen Wang, Yuan Zhong, Lingwei Zhang, Lisong Dai, Ting Wang, Fenglong Ma

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Renmin Hospital of Wuhan University(武汉大学仁民医院) Stony Brook University(石溪大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Submitted to Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15629 2025-05-22 cs.MM 79%

Relationship Analysis of Image-Text Pair in SNS Posts

Takuto Nabeoka, Yijun Duan, Qiang Ma

专题命中 图文多模态 :image-text(title,abstract);分类 cs.MM

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15217 2025-05-22 cs.CV 79%

Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection

Haotian Qin, Dongliang Chang, Yueying Gao, Bingyao Yu, Lei Chen, Zhanyu Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 24 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14340 2025-05-21 cs.CV cs.LG 79%

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey

Seunghyuk Cho, Zhenyue Qin, Yang Liu, Youngbin Choi, Seungbeom Lee, Dongwoo Kim

机构 * Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系) Australian National University(澳大利亚国立大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08851 2025-05-20 cs.LG cs.AI 79%

Mimic In-Context Learning for Multimodal Tasks

Yuchu Jiang, Jiale Fu, Chenduo Hao, Xinting Hu, Yingzhe Peng, Xin Geng, Xu Yang

机构 * Southeast University(东南大学) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, (Southeast University), Ministry of Education(新一代人工智能技术及其跨学科应用重点实验室) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 7 figures,CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18589 2025-05-14 cs.CV 79%

Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency

Zhikai Wang, Jiashuo Sun, Wenqi Zhang, Zhiqiang Hu, Xin Li, Fan Wang, Deli Zhao

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎扑实验室) Zhejiang University(浙江大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Home page: https://alibaba-damo-academy.github.io/VCBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02278 2025-05-06 cs.CV 79%

Compositional Image-Text Matching and Retrieval by Grounding Entities

Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká

机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted at CVPR-W

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11786 2025-04-17 cs.CV 79%

DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generation

Sang-Jun Park, Keun-Soo Heo, Dong-Hee Shin, Young-Han Son, Ji-Hye Oh, Tae-Eui Kam

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09130 2025-04-15 cs.CL 79%

VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search

Yikun Wang, Siyin Wang, Qinyuan Cheng, Zhaoye Fei, Liang Ding, Qipeng Guo, Dacheng Tao, Xipeng Qiu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05644 2025-04-09 cs.CV 79%

iEBAKER: Improved Remote Sensing Image-Text Retrieval Framework via Eliminate Before Align and Keyword Explicit Reasoning

Yan Zhang, Zhong Ji, Changxu Meng, Yanwei Pang, Jungong Han

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05575 2025-04-09 cs.CV cs.LG 79%

A Lightweight Large Vision-language Model for Multimodal Medical Images

Belal Alsinglawi, Chris McCarthy, Sara Webb, Christopher Fluke, Navid Toosy Saidy

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01700 2025-04-03 cs.HC cs.AI cs.RO 79%

Reasoning LLMs for User-Aware Multimodal Conversational Agents

Hamed Rahimi, Jeanne Cattoni, Meriem Beghili, Mouad Abrini, Mahdi Khoramshahi, Maribel Pino, Mohamed Chetouani

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24182 2025-04-01 cs.CV 79%

CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization

Yingrui Ji, Xi Xiao, Gaofei Chen, Hao Xu, Chenrui Ma, Lijing Zhu, Aokun Liang, Jiansheng Chen

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23667 2025-04-01 cs.CV 79%

Context-Independent OCR with Multimodal LLMs: Effects of Image Resolution and Visual Complexity

Kotaro Inoue

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23503 2025-04-01 cs.CL 79%

Evolutionary Prompt Optimization Discovers Emergent Multimodal Reasoning Strategies in Vision-Language Models

Sid Bharthulwar, John Rho, Katrina Brown

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Published at ICLR 2025 Workshop on Reasoning and Planning for LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01167 2025-04-01 cs.CV 79%

Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data

Haoxin Li, Boyang Li

专题命中 图文多模态 :multimodal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16707 2025-03-31 cs.CV 79%

Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding

Jinlong Li, Cristiano Saltori, Fabio Poiesi, Nicu Sebe

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20188 2025-03-27 cs.CV 79%

Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector

Xiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu, Xiaoming Liu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8 figures; 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16843 2025-03-24 cs.CV 79%

LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models

Jian Liang, Wenke Huang, Guancheng Wan, Qu Yang, Mang Ye

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16413 2025-03-21 cs.CV cs.RO 79%

M3: 3D-Spatial MultiModal Memory

Xueyan Zou, Yuchen Song, Ri-Zhao Qiu, Xuanbin Peng, Jianglong Ye, Sifei Liu, Xiaolong Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICLR2025 homepage: https://m3-spatial-memory.github.io code: https://github.com/MaureenZOU/m3-spatial

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15940 2025-03-21 cs.CV 79%

UniCrossAdapter: Multimodal Adaptation of CLIP for Radiology Report Generation

Yaxiong Chen, Chuang Du, Chunlei Li, Jingliang Hu, Yilei Shi, Shengwu Xiong, Xiao Xiang Zhu, Lichao Mou

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CV

Comments MICCAI 2024 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23996 2025-03-18 cs.LG cs.AI cs.IT math.IT 79%

An Information Criterion for Controlled Disentanglement of Multimodal Data

Chenyu Wang, Sharut Gupta, Xinyi Zhang, Sana Tonekaboni, Stefanie Jegelka, Tommi Jaakkola, Caroline Uhler

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00153 2025-03-12 cs.CV cs.LG 79%

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

Kunyang Han, Yibo Hu, Mengxue Qu, Hailin Shi, Yao Zhao, Yunchao Wei

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07090 2025-03-11 stat.ML cs.AI cs.LG 79%

Generative Distribution Prediction: A Unified Approach to Multimodal Learning

Xinyu Tian, Xiaotong Shen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments 31 pages 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20548 2025-03-11 cs.RO cs.AI cs.HC 79%

Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant

Anxing Xiao, Nuwan Janaka, Tianrun Hu, Anshul Gupta, Kaixin Li, Cunjun Yu, David Hsu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏