Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models
通用增强,特定抑制:基于稀疏自编码器引导的医学视觉语言模型
Farhad Nooralahzadeh, Benjamin Gundersen, Nicolas Deperrois, Hidetoshi Matsuom, Mizuho Nishio, Thomas Frauenfelder, Ahmed Allam, Christian Blüthgen, Michael Moor, Michael Krauthammer
机构
*
University of Zurich and University Hospital of Zurich(苏黎世大学及苏黎世大学医院)
;
Kobe University(Kobe大学)
;
ETH AI Center(苏黎世联邦理工学院人工智能中心)
;
ETH Zurich(苏黎世联邦理工学院)
;
Stanford University(斯坦福大学)
;
Zurich University of Applied Sciences(苏黎世应用科学大学)
机构
*
School of Cyberspace Security, Northwestern Polytechnical University(网络安全学院,西北工业大学)
;
School of Computer Science, Northwestern Polytechnical University(计算机学院,西北工业大学)
;
Intellifusion(智融科技)
机构
*
The University of Hong Kong(香港大学)
;
Nanjing University(南京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Fudan University(复旦大学)
Progressive Multimodal Alignment for Continual Instruction Tuning
用于持续指令微调的渐进式多模态对齐
Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心)
;
Kyoto University(京都大学)
;
Migu Culture Technology Co.,Ltd.(咪咕文化科技有限公司)
;
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)
专题命中
VLM训练与架构
:MLLM(summary_cn,abstract);multimodal large language model(abstract,abstract_cn);分类 cs.CV、cs.AI
CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning
CLASP: 基于语言驱动的机器人技能选择与组合,采用任务参数化学习
Markus Knauer, Valentin Gieraths, Tai Mai, Samuel Bustamante, Alin Albu-Schäffer, Freek Stulp, João Silvério
机构
*
German Aerospace Center (DLR), Institute of Robotics and Mechatronics (RMC)(德国航空航天中心(DLR),机器人与机电一体化研究所(RMC))
;
Technical University of Munich (TUM)(慕尼黑工业大学(TUM))
机构
*
Department of Innovation Engineering(创新工程系)
;
University of Salento, Italy(意大利萨伦托大学)
;
Institute of Applied Sciences and Intelligent Systems - CNR(应用科学与智能系统研究所 - CNR)
;
University of the Basque Country UPV/EHU(巴斯克国家大学UPV/EHU)
;
IKERBASQUE, Basque Foundation for Science(伊基塔斯克巴塞克基金会)
;
Sorbonne University Abu Dhabi(索邦大学阿布扎比分校)
专题命中
VLM训练与架构
:VLM(title,abstract);vision language model(title);分类 cs.CV、cs.AI
Jailbreaking Vision-Language Models Through the Visual Modality
通过视觉模态 jailbreak 视觉-语言模型
Aharon Azulay, Jan Dubiński, Zhuoyun Li, Atharv Mittal, Yossi Gandelsman
机构
*
Warsaw University of Technology(华沙理工大学)
;
NASK National Research Institute(波兰国家研究研究院)
;
University of Liverpool(利物浦大学)
;
Indian Institute of Technology Roorkee(印度理工学院拉胡尔分校)
;
Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Helmholtz Munich(亥姆霍兹慕尼黑)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
Munich Data Science Institute(慕尼黑数据科学研究所)
;
University of Tübingen(图宾根大学)
;
University of Copenhagen(哥本哈根大学)
FastVLM: Efficient Vision Encoding for Vision Language Models
Pavan Kumar Anasosalu Vasu, Fartash Faghri, Chun-Liang Li, Cem Koc, Nate True, Albert Antony, Gokul Santhanam, James Gabriel, Peter Grasch, Oncel Tuzel, Hadi Pouransari
机构
*
Apple(苹果公司)
专题命中
VLM训练与架构
:vision language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG