VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI
VL2Spike:面向具身AI低功耗视觉感知的VLM脉冲驱动蒸馏
Zinan Liu, Eric Zheng, Soumyaratna Debnath, Hao Shi, Ling Xiao, Lin Wang
机构
*
School of EEE, Nanyang Technological University (NTU)(南洋理工大学电气与电子工程学院)
;
Department of Computer Science, University of Toronto(多伦多大学计算机科学系)
;
Advanced Micro Devices, Inc.(超威半导体公司)
;
State Key Laboratory of Extreme Photonics and Instrumentation, Zhejiang University(浙江大学极端光子学与仪器国家重点实验室)
;
Faculty of Information Science and Technology, Hokkaido University(北海道大学信息科学与技术学院)
机构
*
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
University of California, Berkeley(加州大学伯克利分校)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI
The Vision Encoder as a Privacy Boundary: Visual-Token Side Channels in Encoder-Free Vision-Language Models
视觉编码器作为隐私边界:无编码器视觉-语言模型中的视觉令牌侧信道
Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
机构
*
School of Engineering, Institute of Science Tokyo(东京科学大学工学院)
;
College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院)
;
Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电气与计算机工程系)
CPS4: Class Prompt driven Semi-Supervised Spine Segmentation with Class-specific Consistency Constraint
CPS4: 基于类别提示的半监督脊柱分割与类别特定一致性约束
Qingtao Pan, Hongzan Sun, Bing Ji, Shuo Li
机构
*
School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院)
;
Department of Nuclear Medicine, Shengjing Hospital of China Medical University(中国医科大学附属盛京医院核医学科)
;
Department of Computer and Data Science, Case Western Reserve University(凯斯西储大学计算机与数据科学系)
;
Department of Biomedical Engineering, Case Western Reserve University(凯斯西储大学生物医学工程系)
专题命中
VLM训练与架构
:VLM(summary_cn,abstract);vision language model(abstract);分类 cs.CV
机构
*
School of Computer Science, Peking University(北京大学计算机科学学院)
;
School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院)
;
School of Information, Renmin University of China(中国人民大学信息学院)
;
School of Integrated Circuit Science and Engineering, Beihang University(北京航空航天大学集成电路科学与工程学院)
专题命中
VLM训练与架构
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
Tool-IQA: Augmenting Image Quality Assessment with Simple Tools
Tool-IQA: 利用简单工具增强图像质量评估
Guanyi Qin, Junjie Zhang, Chunming He, Yibing Fu, Jie Liang, Tianhe Wu, Lei Zhang
机构
*
National University of Singapore(新加坡国立大学)
;
OPPO Research Institute(OPPO研究院)
;
Nanyang Technical University(南洋理工大学)
;
Duke University(杜克大学)
;
City University of Hong Kong(香港城市大学)
;
The Hong Kong Polytechnic University(香港理工大学)
iTRIALSPACE: Programmable Virtual Lesion Trials for Controlled Evaluation of Lung CT Models
iTRIALSPACE:用于肺CT模型受控评估的可编程虚拟病灶试验
Fakrul Islam Tushar, Umme Hafsa Momy, Joseph Y. Lo, Geoffrey D. Rubin
机构
*
Department of Radiology and Imaging Sciences, University of Arizona(亚利桑那大学放射科和影像科学系)
;
Department of Biomedical Engineering, Florida International University(佛罗里达国际大学生物医学工程系)
;
Center for Virtual Imaging Trials, Department of Radiology, Duke University Medical Center(达特茅斯大学医学中心虚拟成像试验中心,放射科)
Attention, not scale, drives human-AI alignment in multimodal language prediction
注意力,而非规模,驱动多模态语言预测中的人机对齐
Viktor Kewenig, Andrew Lampinen, Samuel A. Nastase, Christopher Edwards, Quitterie Lacome D'Elascombe, Akilles Rechardt, Jeremy I Skipper, Gabriella Vigliocco
机构
*
Psychology and Language Science, Experimental Psychology, University College London, London, UK(心理学与语言科学、实验心理学,伦敦大学学院,伦敦,英国)
;
Google Deepmind, Mountain View, US(谷歌DeepMind,山景城,美国)
;
Princeton Neuroscience Institute, Princeton University, Princeton, NJ, USA(普林斯顿神经科学研究所,普林斯顿大学,普林斯顿,新泽西州,美国)
;
Computer Science Department, Exeter University(计算机科学系,埃克塞特大学)