arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2508.06220 2025-08-14 cs.CL cs.AI 62%

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

Keummin Ka, Junhyeong Park, Jaehyun Jeon, Youngjae Yu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06685 2025-08-14 cs.MM cs.CV 62%

Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding

Dawei Huang, Qing Li, Chuan Yan, Zebang Cheng, Zihao Han, Yurong Huang, Xiang Li, Bin Li, Xiaohui Wang, Zheng Lian, Zhi-Qi Cheng, Xiaojiang Peng

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06795 2025-08-13 cs.CL cs.CV 62%

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Yuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang, Zhengwei Fang, Jingyuan Zhang, Jiawei Chen, Zinan Liu, Yu Tian

机构 * University of Chinese Academy of Sciences(中国科学院大学) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学) Gaoling School of Artificial Intelligence, Renmin University of China(人工智能学院,中国人民大学) Kuaishou Technology Inc.(快手科技有限公司) Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University(多信息处理重点实验室,华东师范大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07432 2025-08-12 cs.CV cs.AI 62%

Freeze and Reveal: Exposing Modality Bias in Vision-Language Models

Vivek Hruday Kavuri, Vysishtya Karanam, Venkata Jahnavi Venkamsetty, Kriti Madumadukala, Lakshmipathi Balaji Darur, Ponnurangam Kumaraguru

机构 * IIIT Hyderabad

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04469 2025-08-07 cs.CV cs.CL 62%

FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding

Emmanuelle Bourigault, Pauline Bourigault

机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) Department of Electrical Engineering, Imperial College London(伦敦帝国学院电子工程系)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01697 2025-08-07 cs.CV cs.AI 62%

Hulk: A Universal Knowledge Translator for Human-Centric Tasks

Yizhou Wang, Yixuan Wu, Weizhen He, Xun Guo, Feng Zhu, Lei Bai, Rui Zhao, Jian Wu, Tong He, Wanli Ouyang, Shixiang Tang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by TPAMI2025

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, Jul. 2025, pp. 5672-5689, vol. 47

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02890 2025-08-06 cs.CV cs.CL 62%

VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction

Rongxin Jiang, Robert Long, Chenghao Gu, Mingrui Yan

机构 * Heilongjiang University of Science and Technology(黑龙江科技大学) University of Padua(帕多瓦大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10342 2025-08-05 cs.CV cs.AI 62%

UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models

Jun Yin, Jing Zhong, Peilin Li, Ruolin Pan, Pengyu Zeng, Miao Zhang, Shuai Lu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01932 2025-08-05 cs.CV cs.AI 62%

Proactive Disentangled Modeling of Trigger-Object Pairings for Backdoor Defense

Kyle Stein, Andrew A. Mahyari, Guillermo Francia, Eman El-Sheikh

机构 * Department of Intelligent Systems and Robotics, University of West Florida(智能系统与机器人系,西佛罗里达大学) Florida Institute For Human and Machine Cognition (IHMC)(佛罗里达人类与机器认知研究所) Center for Cybersecurity, University of West Florida(网络安全中心,西佛罗里达大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Journal ref Computers, Materials & Continua, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01540 2025-08-05 cs.CV cs.AI 62%

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning

Yi Liu, Xiao Xu, Zeyu Xu, Meng Zhang, Yibo Li, Haoyu Chen, Junkang Zhang, Qiang Wang, Jifa Sun, Siling Lin, Shengxun Cheng, Lingshu Zhang, Kang Wang

机构 * Project Leader(项目负责人) Corresponding Author(通讯作者)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01387 2025-08-05 cs.CV cs.AI 62%

Video-based Vehicle Surveillance in the Wild: License Plate, Make, and Model Recognition with Self Reflective Vision-Language Models

Pouya Parsa, Keya Li, Kara M. Kockelman, Seongjin Choi

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 19 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21246 2025-07-30 cs.CV cs.AI 62%

On Explaining Visual Captioning with Hybrid Markov Logic Networks

Monika Shah, Somdeb Sarkhel, Deepak Venugopal

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18915 2025-07-28 cs.CL cs.CV 62%

Mining Contextualized Visual Associations from Images for Creativity Understanding

Ananya Sahu, Amith Ananthram, Kathleen McKeown

机构 * Columbia University(哥伦比亚大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05211 2025-07-28 cs.CV cs.AI 62%

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang, Rao Muhammad Anwer

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫兹哈德大学人工智能大学) Technical University of Munich(慕尼黑技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15265 2025-07-28 cs.CV cs.AI cs.CR 62%

Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs

Zihao Pan, Yu Tong, Weibin Wu, Jingyi Wang, Lifeng Chen, Zhe Zhao, Jiajia Wei, Yitong Qiao, Zibin Zheng

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments The paper needs major revisions, so it is being withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06788 2025-07-25 cs.CV cs.AI 62%

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Haiwen Diao, Xiaotong Li, Yufeng Cui, Yueze Wang, Haoge Deng, Ting Pan, Wenxuan Wang, Huchuan Lu, Xinlong Wang

机构 * DLUT(大连理工大学) BAAI(北京人工智能研究院) PKU(北京大学) BUPT(北京邮电大学) UCAS(中国科学院大学) CASIA(中国科学院自动化研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 10 figures, Accepted by ICCV2025 (highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17467 2025-07-24 cs.CV cs.AI 62%

Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls

Elena Pitta, Tom Kouwenhoven, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science (LIACS), Leiden University, The Netherlands(莱顿先进计算机科学研究所(LIACS)、莱顿大学、荷兰)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments LUHME: 2nd Workshop on Language Understanding in the Human-Machine Era

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05056 2025-07-23 cs.CV cs.AI 62%

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling

Xin Dong, Shichao Dong, Jin Wang, Jing Huang, Li Zhou, Zenghui Sun, Lihua Jing, Jingsong Lan, Xiaoyong Zhu, Bo Zheng

机构 * University of Chinese Academy of Sciences(中国科学院大学) The University of Hong Kong(香港大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05195 2025-07-23 cs.LG cs.CL cs.CV 62%

Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder

Siting Li, Pang Wei Koh, Simon Shaolei Du

机构 * University of Washington(华盛顿大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments ACL 2025; 19 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09487 2025-07-22 cs.CV cs.AI 62%

HMID-Net: An Exploration of Masked Image Modeling and Knowledge Distillation in Hyperbolic Space

Changli Wang, Fang Yin, Jiafeng Liu, Rui Wu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Modified the abstract and reformatted it using latex

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13363 2025-07-21 cs.CV cs.AI 62%

Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop

Atharv Goel, Mehar Khurana

机构 * Indraprastha Institute of Information Technology(印度理工学院信息技术研究所)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04106 2025-07-18 cs.CV cs.AI 62%

MRGen: Segmentation Data Engine for Underrepresented MRI Modalities

Haoning Wu, Ziheng Zhao, Ya Zhang, Yanfeng Wang, Weidi Xie

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025; Project Page: https://haoningwu3639.github.io/MRGen/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06210 2025-07-17 cs.CV cs.CL 62%

CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions

Yuchen Huang, Zhiyuan Fan, Zhitao He, Sandeep Polisetty, Wenyan Li, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港理工大学) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) University of Copenhagen(哥本哈根大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 25 pages, COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08021 2025-07-14 cs.CL cs.AI 62%

Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis

Li Li, Yongliang Wu, Jingze Zhu, Jiawei Peng, Jianfei Cai, Xu Yang

机构 * School of Computer Science & Engineering, Key Lab of New Generation Artificial Intelligence Technology & Its Interdisciplinary Applications (Ministry of Education), Southeast University, China(计算机科学与工程学院,新一代人工智能技术及跨学科应用重点实验室,东南大学,中国) Faculty of IT, Monash University, Australia(信息科技学院,墨尔本大学,澳大利亚) School of Computer Science & Engineering, Southeast University, China(计算机科学与工程学院,东南大学,中国)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 16 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04509 2025-07-08 cs.CV cs.AI 62%

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization

Zhendong Xiao, Wu Wei, Shujie Ji, Shan Yang, Changhao Chen

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能方向,香港科学与技术大学(广州))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments PRCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04141 2025-07-08 cs.CV cs.AI cs.ET cs.LG cs.RO 62%

Pedestrian Intention Prediction via Vision-Language Foundation Models

Mohsen Azarmi, Mahdi Rezaei, He Wang

机构 * Institute for Transport Studies, Faculty of Environment, Computer Vision and Machine Learning Group, University of Leeds(运输研究学院、环境学院、计算机视觉与机器学习小组、莱斯特大学) AI Centre, Department of Computer Science, University College London(人工智能中心、计算机科学系、伦敦大学学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03458 2025-07-08 cs.CV cs.AI 62%

Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach

Leyan Xue, Zongbo Han, Guangyu Wang, Qinghua Hu, Mingyue Cheng, Changqing Zhang

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00898 2025-07-02 cs.CV cs.CL 62%

ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models

Zifu Wan, Ce Zhang, Silong Yong, Martin Q. Ma, Simon Stepputtis, Louis-Philippe Morency, Deva Ramanan, Katia Sycara, Yaqi Xie

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICCV 2025. Project page: https://zifuwan.github.io/ONLY/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03704 2025-07-02 cs.CV cs.CL cs.LG 62%

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Xiyao Wang, Zhengyuan Yang, Linjie Li, Hongjin Lu, Yuancheng Xu, Chung-Ching Lin, Kevin Lin, Furong Huang, Lijuan Wang

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Microsoft(微软公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13949 2025-07-02 cs.CV cs.AI 62%

SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning

Ziqi Wang, Chang Che, Qi Wang, Yangyang Li, Zenglin Shi, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) Tsinghua University(清华大学) Academy of Cyber(网络学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏