Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
眼见无需耗能,言语却要代价:揭示边缘视觉语言模型推理中的真正能量瓶颈
Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
Jinan University(暨南大学)
On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing
遥感领域基于视觉语言模型的联邦学习中适配策略的有效性研究
Simon Lösche, Barış Büyüktaş, Mathis Adler, Angelos Zavras, Ioannis Papoutsis, Begüm Demir
机构
*
BIFOLD - Berlin Institute for the Foundations of Learning and Data(BIFOLD - 柏林学习与数据基础研究所)
;
Technische Universität Berlin(柏林工业大学)
;
National Technical University of Athens(雅典国家技术大学)
;
National Observatory of Athens(雅典国家天文台)
;
Harokopio University of Athens(哈罗科皮奥雅典大学)
MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models
MEDLAYXPLAIN: 医学视觉语言模型中的专家-外行差距基准测试
Han Jang, Junhyeok Lee, Songsoo Kim, Chae Young Lim, Hyeonjin Goh, Heeseong Eum, Kyu Sung Choi
机构
*
Seoul National University(首尔大学)
;
Seoul National University Hospital(首尔大学医院)
;
Seoul National University College of Medicine(首尔大学医学院)
;
Sungkyunkwan University School of Medicine(成均馆大学医学院)
VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models
面向冻结视觉语言模型的VLM感知元光学前端设计
Chanik Kang, Raphaël Pestourie, Haejun Chung
机构
*
Department of Artificial Intelligence, Hanyang University(汉阳大学人工智能系)
;
School of Computational Science and Engineering, Georgia Institute of Technology(佐治亚理工学院计算科学与工程学院)
;
Department of Electronic Engineering, Hanyang University(汉阳大学电子工程系)
FineGen: A VLM-based Multi-Agent Framework for Fine-Grained Image-Text Dataset Construction
FineGen:基于VLM的多智能体框架用于细粒度图像-文本数据集构建
Chang Kong, Yuebing Li, Peng Mo, Haigang Zhang, Qiuming Luo
机构
*
Shenzhen Polytechnic University(深圳职业技术大学)
;
Institute of Applied Artificial Intelligence of the Guangdong-Hong Kong Macao Greater Bay Area(粤港澳大湾区应用人工智能研究所)
;
Shenzhen University(深圳大学)
VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI
VL2Spike:面向具身AI低功耗视觉感知的VLM脉冲驱动蒸馏
Zinan Liu, Eric Zheng, Soumyaratna Debnath, Hao Shi, Ling Xiao, Lin Wang
机构
*
School of EEE, Nanyang Technological University (NTU)(南洋理工大学电气与电子工程学院)
;
Department of Computer Science, University of Toronto(多伦多大学计算机科学系)
;
Advanced Micro Devices, Inc.(超威半导体公司)
;
State Key Laboratory of Extreme Photonics and Instrumentation, Zhejiang University(浙江大学极端光子学与仪器国家重点实验室)
;
Faculty of Information Science and Technology, Hokkaido University(北海道大学信息科学与技术学院)