arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45832 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4634 篇

2503.19311 2025-10-30 cs.CV cs.AI 84%

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Weizhi Chen, Yupeng Deng, Jin Wei, Jingbo Chen, Jiansheng Chen, Yuman Feng, Zhihao Xi, Diyou Liu, Kai Li, Yu Meng

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院 aerospace information research institute) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Information Network Security, People’s Public Security University of China(中国人民公安大学信息网络安全学院)

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15510 2025-10-28 cs.CV cs.CL 84%

Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought

Zihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang, Weiyun Wang, Hao Fei, Yidong Wang, Alex Jinpeng Wang, Zhi Chen, Wanxiang Che, Libo Qin

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(哈尔滨工业大学社会计算与交互机器人研究中心) Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院计算与智能研究所) Text Computing and Cognitive Intelligence Ministry of Education Engineering Research Center, Guizhou University(贵州大学文字计算与认知智能教育部工程研究中心) Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室) National University of Singapore(新加坡国立大学) Peking University(北京大学) ByteDance Seed (China)(字节跳动种子(中国))

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025;

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21346 2025-10-27 cs.CV cs.AI 84%

CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments

Lemin Liu, Fangchao Hu, Honghua Jiang, Yaru Chen, Limin Liu, Yongliang Qiao

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09966 2025-10-21 eess.IV cs.AI cs.CV cs.LG 84%

Multimodal Fusion at Three Tiers: Physics-Driven Data Generation and Vision-Language Guidance for Brain Tumor Segmentation

Mingda Zhang

机构 * Software School, Yunnan University, Kunming 650504, Yunnan, China(云南大学软件学院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 31 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10013 2025-10-16 cs.CV cs.CL 84%

Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect

Tom Kouwenhoven, Kiana Shahrasbi, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science(莱顿先进计算机科学研究所) Leiden University(莱顿大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09815 2025-10-14 cs.CV cs.AI 84%

Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning

Yufei Wang, Adriana Kovashka, Loretta Fernández, Marc N. Coutanche, Seth Wiener

机构 * University of Pittsburgh(匹兹堡大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to International Conference on Development and Learning (ICDL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14715 2025-10-07 cs.CV cs.AI 84%

Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models

Young Kyun Jang, Ser-nam Lim

机构 * Google DeepMind(谷歌DeepMind) University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20961 2025-09-26 cs.CV cs.AI 84%

Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos

Sarmistha Das, R E Zera Marveen Lyngkhoi, Sriparna Saha, Alka Maurya

机构 * Indian Institute of Technology Patna(印度帕纳杰大学) CRISIL LTD(CRISIL公司)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05244 2025-09-22 cs.CV cs.AI 84%

RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding

Tianchen Fang, Guiru Liu

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Upon further review, we identified that our dataset requires optimization to ensure research reliability and accuracy. Additionally, considering the target journal's latest submission policies, we believe comprehensive manuscript revisions are necessary

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15688 2025-09-16 cs.CL cs.AI cs.LG 84%

Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts

Haodi Ma, Dzmitry Kasinets, Daisy Zhe Wang

机构 * Department of Computer and Information Science and Engineering, University of Florida(计算机与信息科学与工程系,佛罗里达大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06569 2025-08-21 cs.CL cs.AI 84%

Social Debiasing for Fair Multi-modal LLMs

Harry Cheng, Yangyang Guo, Qingpei Guo, Ming Yang, Tian Gan, Weili Guan, Liqiang Nie

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments Project page: https://github.com/xaCheng1996/Social_Debiasing_For_Fair_MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19925 2025-08-14 cs.CL cs.CV cs.LG 84%

Improving Multimodal Large Language Models Using Continual Learning

Shikhar Srivastava, Md Yousuf Harun, Robik Shrestha, Christopher Kanan

机构 * University of Rochester(罗切斯特大学) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments CoLLAs 2025 and Scalable Continual Learning for Lifelong Foundation Models, NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19015 2025-08-11 cs.CV cs.MM 84%

Can Multimodal Large Language Models Understand Spatial Relations?

Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou, Yinan Zou, Weiyan Zhang, Haiyun Jiang, Tong Ruan

机构 * School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China(信息科学与工程学院,东华大学,上海,中国) School of Computer Science, Fudan University, Shanghai, China(计算机科学学院,复旦大学,上海,中国)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 7 figures, published to ACL 2025

Journal ref In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 620-632, 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12591 2025-08-06 cs.CV cs.CL 84%

CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base

Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu, Shuai Zhao, Thong Nguyen, Anh Tuan Luu

机构 * Nanyang Technological University, Singapore(南洋理工大学) National University of Singapore, Singapore(国立新加坡大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18659 2025-08-01 cs.CV cs.AI 84%

DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Yudong Zhang, Ruobing Xie, Xingwu Sun, Yiqing Huang, Jiansheng Chen, Zhanhui Kang, Di Wang, Yu Wang

机构 * Tsinghua University, Tencent(清华大学、腾讯) Tencent(腾讯) Tencent, University of Macau(腾讯、澳门大学) University of Science and Technology Beijing(北京科技大学) Tsinghua University(清华大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20913 2025-07-29 cs.CV cs.AI 84%

HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection

Jialei Cui, Jianwei Du, Yanzhe Li, Lei Gao, Hui Jiang, Chenfu Bao

机构 * Baidu Inc.(百度公司) Southeast University(东南大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20156 2025-07-29 cs.CV cs.AI 84%

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

Daulet Toibazar, Kesen Wang, Sherif Mohamed, Abdulaziz Al-Badawi, Abdulrahman Alfulayt, Pedro J. Moreno

机构 * Humain Riyadh, KSA(利雅得人类,沙特阿拉伯)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17080 2025-07-24 cs.IR cs.AI cs.CV 84%

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings

Ramin Giahi, Kehui Yao, Sriram Kollipara, Kai Zhao, Vahid Mirjalili, Jianpeng Xu, Topojoy Biswas, Evren Korpeoglu, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at RecSys 2025; DOI:https://doi.org/10.1145/3705328.3748064

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10202 2025-07-15 cs.CV cs.AI 84%

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

Jaeseong Lee, Yeeun Choi, Heechan Choi, Hanjung Kim, Seonjoo Kim

机构 * Yonsei University(延世大学)

专题命中 图文多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at CVPR 2025 Workshop on Emergent Visual Abilities and Limits of Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15251 2025-07-08 cs.CL cs.AI 84%

AgentPS: Agentic Process Supervision for Content Moderation with Multimodal LLMs

Mingchao Liu, Yu Sun, Ruixiao Sun, Xin Dong, Xiang Shen, Hongyu Xiong

机构 * TikTok, Inc.(字节跳动公司)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16760 2025-06-23 cs.CL cs.CV 84%

Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models

Lei Jiang, Zixun Zhang, Zizhou Wang, Xiaobing Sun, Zhen Li, Liangli Zhen, Xiaohua Xu

机构 * University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Institute of High Performance Computing, A*STAR, Singapore(新加坡A*STAR高性能计算研究所)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15645 2025-06-19 cs.CV cs.AI 84%

Demystifying the Visual Quality Paradox in Multimodal Large Language Models

Shuo Xing, Lanqing Guo, Hongyuan Hua, Seoyoung Lee, Peiran Li, Yufei Wang, Zhangyang Wang, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Toronto(多伦多大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07399 2025-06-10 cs.CV cs.AI 84%

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems

Peiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai, Huili Wang, Yufei Sun, Xintian Li, Shangguang Wang, Yongfeng Huang, Tao Qi

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07227 2025-06-10 cs.CV cs.CL 84%

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

Tianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun, Jiayi Song, Junlin Han, Zichen Liu, Conghui He, Wentao Zhang, Binhang Yuan

机构 * HKUST(香港科技大学) Shanghai AI Lab(上海人工智能实验室) Peking University(北京大学) HKUST(GZ)(香港科技大学(广州)) Oxford University(牛津大学) Imperial College London(伦敦帝国理工学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15239 2025-06-10 cs.CV cs.AI cs.IR 84%

Benchmark Granularity and Model Robustness for Image-Text Retrieval

Mariya Hendriksen, Shuo Zhang, Ridho Reinanda, Mohamed Yahya, Edgar Meij, Maarten de Rijke

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments accepted at SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04277 2025-06-06 cs.CV cs.AI 84%

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Yi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li, Licheng Tang, Yangguang Ji, Chong Wu, Jay Wu, Wenbo Zhu

专题命中 图文多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11784 2025-06-05 cs.AI cs.CV cs.LG 84%

Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development

Daoyuan Chen, Haibin Wang, Yilun Huang, Ce Ge, Yaliang Li, Bolin Ding, Jingren Zhou

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICML 2025 (Spotlight). 33 pages, 16 tables, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14035 2025-05-21 cs.MM cs.CL 84%

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs

Shiyao Cui, Qinglin Zhang, Xuan Ouyang, Renmiao Chen, Zhexin Zhang, Yida Lu, Hongning Wang, Han Qiu, Minlie Huang

机构 * The Conversational AI (CoAI) group, DCST, Tsinghua University China(清华大学人工智能对话组,国防科技大学,清华大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00958 2025-05-14 cs.CV cs.CL cs.LG 84%

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Wenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun, Yongliang Shen, Weiming Lu, Deli Zhao, Yueting Zhuang, Lidong Bing

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06152 2025-05-12 cs.CV cs.AI 84%

MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks

Wenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan, Bo Yan

机构 * Shanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University(上海智能信息处理关键实验室,计算机科学学院,复旦大学)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏