arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1565 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1565 篇

2605.10012 2026-05-12 cs.HC cs.CR 67%

Sketch-based Access Control: A Multimodal Interface for Translating User Preferences into Intent-Aligned Policies

基于草图的访问控制:一种多模态接口,用于将用户偏好转化为意图对齐的策略

Kyzyl Monteiro, Sauvik Das

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract_cn)

AI总结 本文提出SBAC系统,结合草图和多模态大语言模型,帮助用户逐步完善访问策略,发现潜在问题并验证策略行为。

Comments 27 pages including appendix; 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02200 2026-04-22 cs.NE cs.AI cs.CV cs.LG 67%

Learning Evolution via Optimization Knowledge Adaptation

通过优化知识适应学习进化

Chao Wang, Lingling Li, Licheng Jiao, Jiaxuan Zhao, Fang Liu, Shuyuan Yang

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education(教育部智能感知与图像理解重点实验室) International Research Center for Intelligent Perception and Computation(智能感知与计算国际研究中心) Xidian University(西安电子科技大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出OKAEM模型,通过注意力机制参数化进化算子,实现优化知识的预训练和自适应优化,提升进化算法的知识转移与在线适应能力。

Comments This work has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07833 2026-04-13 cs.CV cs.AI cs.LG 67%

Relational Visual Similarity

关系视觉相似性

Thao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang, Jing Shi, Nicholas Kolkin, Eli Shechtman, Yong Jae Lee, Yuheng Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of California, Los Angeles(加利福尼亚大学洛杉矶分校) Adobe Research(Adobe研究院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出通过关系相似性而非视觉属性来衡量图像相似性,构建了首个基于关系逻辑的图像表示空间,揭示了现有视觉模型在捕捉关系相似性方面的不足。

Comments CVPR 2026 camera-ready; Project page, data, and code: https://thaoshibe.github.io/relsim

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26107 2026-03-30 cs.HC 67%

One Is Not Enough: How People Use Multiple AI Models in Everyday Life

一个不够:人们如何在日常生活中使用多个AI模型

Seunghwa Pyo, Donggun Lee, Jungwoo Rhee, Soobin Park, Youn-kyung Lim

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

AI总结 研究探讨了人们如何在日常生活中协调多个多模态大语言模型,揭示用户构建模型层级和切换策略以优化任务效率与输出可信度。

Comments Accepted as a poster at CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16786 2026-03-06 cs.LG cs.AI cs.CV 67%

Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach

重新审视多模态KV缓存压缩:一种基于频域的异常KV感知方法

Yaoxin Yang, Peng Ye, Xudong Tan, Chongjun Tu, Maosen Zhao, Jia Hao, Tao Chen

机构 * College of Future Information Technology, Fudan University(未来信息科技学院,复旦大学) Shanghai Innovation Institute(上海创新研究院) The Chinese University of Hong Kong(香港中文大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhangjiang Laboratory(张江实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出FlashCache,一种基于频域的异常KV感知KV缓存压缩框架,通过保留关键KV对提升解码效率并降低内存使用。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07666 2026-01-01 cs.LG cs.AI cs.CL cs.CV 67%

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

在大语言模型、多模态大语言模型及更广泛的领域中进行模型融合:方法、理论、应用与机遇

Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, Dacheng Tao

机构 * Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区) Northeastern University China(东北大学) Shenzhen Campus of Sun Yat-sen University China(中山大学深圳校区) Nanyang Technological University Singapore(南洋理工大学) Northeastern University(东北大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Nanyang Technological University(南洋理工大学) Institute for Clarity in Documentation Dublin Ohio USA(文档清晰研究所) Inria Paris-Rocquencourt Rocquencourt France(巴黎-罗quentourt研究所) Rajiv Gandhi University Doimukh Arunachal Pradesh India(拉贾·甘地大学) Tsinghua University Haidian Qu Beijing Shi China(清华大学) Palmer Research Laboratories San Antonio Texas USA(帕勒研究中心) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗quentourt研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究中心)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文综述了模型融合的方法、理论、应用及未来方向,提出新的分类方法并探讨其在多个机器学习领域的应用及挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20781 2025-12-25 cs.IR 67%

Soft Filtering: Guiding Zero-shot Composed Image Retrieval with Prescriptive and Proscriptive Constraints

软过滤:通过规劝性与禁止性约束指导零样本复合图像检索

Youjin Jung, Seongwoo Cho, Hyun-seok Min, Sungchul Choi

专题命中 其他VLM :vision-language model(abstract);multimodal large language model(abstract)

AI总结 本文提出SoFT方法,通过规劝性和禁止性约束提升零样本复合图像检索的准确性与鲁棒性。

Comments Accepted to AAAI 2026 Workshop on New Frontiers in Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19933 2025-12-24 cs.CL 67%

PRISM: A Personality-Driven Multi-Agent Framework for Social Media Simulation

PRISM: 一种基于个性的多智能体框架用于社交媒体模拟

Zhixiang Lu, Xueyuan Deng, Yiran Liu, Yulong Li, Qiang Yan, Imran Razzak, Jionglong Su

机构 * University of Liverpool(利物浦大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University College London(伦敦大学学院) Xi'an Jiaotong-Liverpool University(西安交通大学-利物浦大学) Chinese Academy of Sciences(中国科学院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

AI总结 PRISM通过结合连续情绪演变与基于个性的决策过程,提供了一种更准确模拟社交媒体中个性驱动意见极化的框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26536 2025-11-26 cs.CL cs.AI cs.CV cs.LG cs.RO 67%

OceanGym: A Benchmark Environment for Underwater Embodied Agents

OceanGym: 一个用于水下具身智能体的基准环境

Yida Xue, Mingjun Mao, Xiangyuan Ru, Yuqi Zhu, Baochang Ren, Shuofei Qiao, Mengru Wang, Shumin Deng, Xinyu An, Ningyu Zhang, Ying Chen, Huajun Chen

专题命中 其他VLM :MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 OceanGym通过多模态大语言模型构建首个水下具身智能体基准,揭示水下环境感知与规划的挑战,推动水下自主系统发展。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19518 2025-11-26 cs.CV cs.AI cs.IT cs.LG math.IT 67%

Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning

迈向高效的VLMs:通过自适应结构压缩的信息论驱动压缩

Zhaoqi Xu, Yingying Zhang, Jian Li, Jianwei Guo, Qiannan Zhu, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Zhongtai Securities Institute for Financial Studies, Shandong University(山东大学中泰证券金融研究学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出InfoPrune框架,通过信息论驱动的自适应结构压缩方法,在保持性能的同时显著提升视觉语言模型的效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13883 2025-10-17 q-bio.NC cs.MA 67%

Large Language Model Agents Enable Autonomous Design and Image Analysis of Microwell Microfluidics

Dinh-Nguyen Nguyen, Sadia Shakil, Raymond Kai-Yu Tong, Ngoc-Duy Dinh

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21956 2025-09-30 cs.CV cs.AI cs.CL cs.LG 67%

Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation

Mengdan Zhu, Senhao Cheng, Guangji Bai, Yifei Zhang, Liang Zhao

机构 * Emory University(埃默里大学) University of Michigan, Ann Arbor(密歇根大学安娜堡分校)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16995 2025-09-23 cs.DC 67%

MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

Zheming Yang, Qi Guo, Yunqing Hu, Chang Zhao, Chang Zhang, Jian Zhao, Wen Ji

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06850 2025-09-16 cs.NE 67%

Visual Evolutionary Optimization on Graph-Structured Combinatorial Problems with MLLMs: A Case Study of Influence Maximization

Jie Zhao, Kang Hao Cheong

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13180 2025-07-25 cs.CV cs.AI cs.LG 67%

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras, Tushar Nagarajan, Muhammad Maaz, Yale Song, Tengyu Ma, Shuming Hu, Suyog Jain, Miguel Martin, Huiyu Wang, Hanoona Rasheed, Peize Sun, Po-Yao Huang, Daniel Bolya, Nikhila Ravi, Shashank Jain, Tammy Stark, Shane Moon, Babak Damavandi, Vivian Lee, Andrew Westbury, Salman Khan, Philipp Krähenbühl, Piotr Dollár, Lorenzo Torresani, Kristen Grauman, Christoph Feichtenhofer

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11798 2025-07-22 cs.CR 67%

BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model

Weilin Lin, Nanjun Zhou, Yanyun Wang, Jianze Li, Hui Xiong, Li Liu

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12378 2025-07-17 cs.IR cs.CL 67%

Developing Visual Augmented Q&A System using Scalable Vision Embedding Retrieval & Late Interaction Re-ranker

Rachna Saxena, Abhijeet Kumar, Suresh Shanmugam

专题命中 其他VLM :visual language model(abstract);MLLM(abstract)

Comments Presented at NLP@IR workshop at SIGIR conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08260 2025-06-25 cs.SE 67%

FixDrive: Automatically Repairing Autonomous Vehicle Driving Behaviour for $0.08 per Violation

Yang Sun, Christopher M. Poskitt, Kun Wang, Jun Sun

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

Comments Accepted by the 47th IEEE/ACM International Conference on Software Engineering (ICSE 2025)

Journal ref Proc. ICSE'25, pages 1921-1933. IEEE, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18095 2025-06-24 cs.CV cs.AI cs.LG 67%

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Junying Chen, Zhenyang Cai, Pengcheng Chen, Shunian Chen, Ke Ji, Xidong Wang, Yunjin Yang, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08023 2025-06-11 q-bio.BM cs.AI cs.CE cs.CV cs.LG 67%

Aligning Proteins and Language: A Foundation Model for Protein Retrieval

Qifeng Wu, Zhengzhe Liu, Han Zhu, Yizhou Zhao, Daisuke Kihara, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Purdue University(普渡大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 4 pages for body, 3 pages for appendix, 11 figures. Accepted to CVPR 2025 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities(MMFM-BIOMED)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07484 2025-06-10 cs.CV cs.AI cs.LG 67%

CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization

Dasol Hong, Wooju Lee, Hyun Myung

机构 * Urban Robotics Lab, School of Electrical Engineering, Korea Advanced Institute of Science and Technology, Republic of Korea(乌尔班机器人实验室,电气工程学院,韩国科学技术院,大韩民国)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 8 pages, 5 figures; accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13487 2025-05-23 cs.CL cs.AI cs.CV cs.LG 67%

Transferring Textual Preferences to Vision-Language Understanding through Model Merging

Chen-An Li, Tzu-Han Lin, Yun-Nung Chen, Hung-yi Lee

机构 * National Taiwan University(台湾大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16855 2025-04-02 cs.CL cs.IR 67%

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Xin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, Min Zhang

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

Comments Accepted to CVPR 2025, models at https://huggingface.co/Alibaba-NLP/gme-Qwen2-VL-2B-Instruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23156 2025-03-04 cs.AI cs.CV cs.LG cs.RO 67%

VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning

Yichao Liang, Nishanth Kumar, Hao Tang, Adrian Weller, Joshua B. Tenenbaum, Tom Silver, João F. Henriques, Kevin Ellis

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments ICLR 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04263 2025-02-07 cs.CV cs.AI cs.LG 67%

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

Marco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Andrew D. Bagdanov

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted for publication at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15188 2025-02-06 cs.CL cs.AI cs.CV cs.LG 67%

LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Weijia Shi, Xiaochuang Han, Chunting Zhou, Weixin Liang, Xi Victoria Lin, Luke Zettlemoyer, Lili Yu

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Name change: LlamaFusion to LMFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17171 2025-01-30 cs.CV cs.AI cs.LG eess.IV 67%

Separated Inter/Intra-Modal Fusion Prompts for Compositional Zero-Shot Learning

Sua Jung

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments AIAP 2025

Journal ref Published at AIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12627 2025-01-07 cs.CL 67%

Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation

Andong Chen, Yuchen Song, Kehai Chen, Muyun Yang, Tiejun Zhao, Min Zhang

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract)

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12232 2024-12-18 cs.CV cs.AI cs.LG 67%

You Only Submit One Image to Find the Most Suitable Generative Model

Zhi Zhou, Lan-Zhe Guo, Peng-Xiao Song, Yu-Feng Li

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted by NeurIPS 2023 Workshop on Diffusion Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10063 2024-11-18 cs.AI cs.CV cs.LG 67%

Federated Domain Generalization via Prompt Learning and Aggregation

Shuai Gong, Chaoran Cui, Chunyun Zhang, Wenna Wang, Xiushan Nie, Lei Zhu

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏