arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1566 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1566 篇

2510.09078 2025-10-13 cs.GR cs.LG 57%

MCMC: Bridging Rendering, Optimization and Generative AI

Gurprit Singh, Wenzel Jakob

机构 * Max Planck Institute for Informatics(马克斯·普朗克研究所信息学研究所) EPFL(瑞士联邦理工学院)

专题命中 其他VLM :vision language model(abstract);分类 cs.LG

Comments SIGGRAPH Asia 2024 Courses. arXiv admin note: text overlap with arXiv:2208.11970 by other authors

Journal ref SIGGRAPH Asia 2024 Courses, Article No.: 8, Pages 1 - 27

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17132 2025-10-08 cs.AI 57%

Applications of Large Models in Medicine

YunHe Su, Zhengyang Lu, Junhui Liu, Ke Pang, Haoran Dai, Sa Liu, Yuxin Jia, Lujia Ge, Jing-min Yang

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04257 2025-10-07 cs.CR cs.AI 57%

AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents

Yanjie Li, Yiming Cao, Dong Wang, Bin Xiao

机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02815 2025-10-06 cs.CV 57%

Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis

Feng Yuan, Yifan Gao, Yuehua Ye, Haoyue Li, Xin Gao

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所) Chinese Academy of Sciences(中国科学院) The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICLR2026 under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02787 2025-10-06 cs.CV 57%

OTR: Synthesizing Overlay Text Dataset for Text Removal

Jan Zdenek, Wataru Shimoda, Kota Yamaguchi

机构 * CyberAgent Tokyo Japan(CyberAgent东京日本)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01247 2025-10-03 cs.CL cs.AI 57%

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

Punit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, Jose G Moreno

机构 * Indian Institute of Technology Patna(印度理工学院帕纳布分校) Sardar Patel Institute of Technology(萨达尔·帕特尔技术学院) Universitas Gadjah Mada(加查马大学) King Mongkut’s Institute of Technology Ladkrabang(拉差班国王技术学院) Shenzhen Technology University(深圳技术大学) Université de Toulouse(图卢兹大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments 52 pages, 56 figures; appearing at EMNLP'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01185 2025-10-02 cs.LG 57%

Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs

Leyla Mirvakhabova, Babak Ehteshami Bejnordi, Gaurav Kumar, Hanxue Liang, Wanru Zhao, Paul Whatmough

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25297 2025-10-02 cs.SE cs.AI 57%

Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development

Yuxuan Wan, Tingshuo Liang, Jiakai Xu, Jingyu Xiao, Yintong Huo, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学) Columbia University in the City of New York(哥伦比亚大学) Singapore Management University(新加坡管理学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00806 2025-10-02 cs.CV 57%

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心、自动化研究所、中国科学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国) School of Artificial Intelligence, University of Chinese Academy of Science, Beijing, China(人工智能学院、中国科学院大学、北京中国) Wuhan AI Research, Wuhan, China(武汉人工智能研究、武汉中国) MAPLE Lab, Westlake University(MAPLE实验室、西湖大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26555 2025-10-01 cs.CV 57%

Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation

Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, Varun Jampani

机构 * Stability AI Arizona State University(亚利桑那州立大学) Google DeepMind(谷歌DeepMind)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments NeurIPS 2025. Project Page : https://stable-cinemetrics.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25863 2025-10-01 cs.CV 57%

MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image Classification

Junjie Zhou, Wei Shao, Yagao Yue, Wei Mu, Peng Wan, Qi Zhu, Daoqiang Zhang

机构 * The College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25817 2025-10-01 cs.CL cs.CV 57%

Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer

Jaeyoung Kim, Jongho Lee, Hongjun Choi, Sion Jang

机构 * Teamreboott Inc.(Teamreboott公司) MIRI D.I.H Inc.(MIRI D.I.H公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23242 2025-09-30 cs.CV 57%

TATTOO: Training-free AesTheTic-aware Outfit recOmmendation

Yuntian Wu, Xiaonan Hu, Ziqi Zhou, Hao Lu

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21980 2025-09-29 cs.CV 57%

Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm

Zeyu Wang, Baiyu Chen, Kun Yan, Hongjing Piao, Hao Xue, Flora D. Salim, Yuanchun Shi, Yuntao Wang

机构 * Key Laboratory of Pervasive Computing, Tsinghua University(清华大学普适计算重点实验室) The University of New South Wales(新南威尔士大学) SKLSDE Lab, Beihang University(北航SKLSDE实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15472 2025-09-29 cs.CV 57%

Efficient Multimodal Dataset Distillation via Generative Models

Zhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang, Gaowen Liu, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Central Florida(中央佛罗里达大学) Cisco Research(思科研究)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16663 2025-09-26 cs.CL cs.AI 57%

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

Yujin Han, Hao Chen, Andi Han, Zhiheng Wang, Xinyu Liu, Yingya Zhang, Shiwei Zhang, Difan Zou

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

Comments 31 pages, 16 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20279 2025-09-25 cs.CV q-bio.QM 57%

A co-evolving agentic AI system for medical imaging analysis

Songhao Li, Jonathan Xu, Tiancheng Bao, Yuxuan Liu, Yuchen Liu, Yihang Liu, Lilin Wang, Wenhui Lei, Sheng Wang, Yinuo Xu, Yan Cui, Jialu Yao, Shunsuke Koga, Zhi Huang

机构 * Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学) Department of Electrical and System Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学) The Wharton School, University of Pennsylvania(沃顿商学院,宾夕法尼亚大学) Department of Bioengineering, University of Pennsylvania(生物工程系,宾夕法尼亚大学) Department of Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学) Department of Biostatistics, Epidemiology & Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16560 2025-09-23 cs.CV 57%

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

Ji Soo Lee, Byungoh Ko, Jaewon Cho, Howoong Lee, Jaewoon Byun, Hyunwoo J. Kim

机构 * Korea University(韩国大学) Hanwha Vision(翰威英航) KAIST(韩国科学技术院)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13939 2025-09-18 cs.CV 57%

Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation

Gia Khanh Nguyen, Yifeng Huang, Minh Hoai

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Stony Brook University(石溪大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11829 2025-09-15 cs.CL cs.AI 57%

Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation

Julia Kreutzer, Eleftheria Briakou, Sweta Agrawal, Marzieh Fadaee, Kocmi Tom

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09541 2025-09-12 cs.AI 57%

Compositional Concept Generalization with Variational Quantum Circuits

Hala Hawashin, Mina Abbaszadeh, Nicholas Joseph, Beth Pearson, Martha Lewis, Mehrnoosh sadrzadeh

机构 * School of Computer Science Engineering University of New South Wales Sydney, Australia Stanford University California, USA Computer Science University College London London, UK School of Eng. Maths. \& Tech University of Bristol Bristol, UK Inst. Logic Language \& Computation University of Amsterdam Amsterdam, NL

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments Accepted to: 2025 IEEE International Conference on Quantum Artificial Intelligence (QAI), Naples, Italy, Nov 2-5, 2025. This is the authors' accepted manuscript (AAM). An IEEE copyright notice appears on page 1. The final published version will appear in IEEE Xplore; DOI to be added when available

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11538 2025-09-12 cs.CL cs.AI eess.AS 57%

MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

Muhammad Huzaifah, Geyu Lin, Tianchi Liu, Hardik B. Sailor, Kye Min Tan, Tarun K. Vangani, Qiongqiong Wang, Jeremy H. M. Wong, Jinyang Wu, Nancy F. Chen, Ai Ti Aw

机构 * MERaLiON Team(MERaLiON团队) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10090 2025-09-10 cs.CV 57%

InteractPro: A Unified Framework for Motion-Aware Image Composition

Weijing Tao, Xiaofeng Yang, Miaomiao Cui, Guosheng Lin

机构 * College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) DAMO Academy, Alibaba Group(阿里达摩院)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18201 2025-09-05 cs.CL cs.CV cs.HC 57%

Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications

Bushra Asseri, Estabraq Abdelaziz, Maha Al Mogren, Tayef Alhefdhi, Areej Al-Wabil

机构 * College of Engineering & Advanced Computing(工程与高级计算学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01259 2025-09-03 cs.CV 57%

ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization

Thinh-Phuc Nguyen, Thanh-Hai Nguyen, Gia-Huy Dinh, Lam-Huy Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南胡志明市科学大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11452 2025-09-03 cs.AI cs.CL cs.HC 57%

Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps

Kangyu Wang, Hongliang He, Lin Liu, Ruiqi Liang, Zhenzhong Lan, Jianguo Li

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Westlake University(西湖大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments Our platform is publicly accessible at https://www.tbox.cn/about/model-ranking

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13859 2025-09-03 cs.CV 57%

Learning Visual Proxy for Compositional Zero-Shot Learning

Shiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing, Lei Zhou, Wenjun Wang

机构 * Tianjin University(天津大学) Zhejiang University(浙江大学) Zhejiang University of Technology(浙江工业大学) Hainan University(海南大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13602 2025-09-03 cs.CV 57%

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction

Xiaolu Hou, Bing Ma, Jiaxiang Cheng, Xuhua Ren, Kai Yu, Wenyue Li, Tianxiang Zheng, Qinglin Lu

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Project Page: https://personavlog-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02518 2025-09-03 cs.LG 57%

AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs

Yao Lai, Souradip Poddar, Sungyoung Lee, Guojin Chen, Mengkang Hu, Bei Yu, Ping Luo, David Z. Pan

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21660 2025-09-03 cs.LG 57%

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

Xiaojie Xu, Xinli Xu, Sirui Chen, Haoyu Chen, Fan Zhang, Ying-Cong Chen

机构 * The Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.LG

Comments Accepted at EMNLP 2025, Findings

详情

展开后加载摘要…

URL PDF HTML 收藏