TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment
TEVI: 基于稀疏自编码器的文本条件视觉表示编辑以改进视觉-语言对齐
Sweta Mahajan, Sukrut Rao, Jiahao Xie, Alexander Koller, Bernt Schiele
机构
*
Max Planck Institute for Informatics, Saarland Informatics Campus, Saarbrücken, Germany(马克斯·普朗克研究所信息学院,萨尔兰信息学院,德国萨尔布吕肯)
;
Department of Language Science and Technology, Saarland University, Saarbrücken, Germany(语言科学与技术系,萨尔兰大学,德国萨尔布吕肯)
CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection
CL-CLIP: 基于CLIP的持续学习框架与代价体积类别解耦用于目标检测
Zihan Liu, Yuguang Yang, Shengjie Su, Jianing Pang, Linlin Yang, Chunyu Xie, Nikolai Yu. Zolotykh, Baochang Zhang
机构
*
National College for Excellent Engineers, Beihang University(卓越工程师学院,北京航空航天大学)
;
AI Research, Qihoo 360(360人工智能研究院,奇虎360)
;
School of Electronic Information Engineering, Beihang University(电子信息学院,北京航空航天大学)
;
School of Cyber Science and Technology, Beihang University(网络安全科学与技术学院,北京航空航天大学)
;
School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北京航空航天大学)
;
State Key Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播国家重点实验室,中国传媒大学)
;
Institute of Information Technology, Mathematics and Mechanics, Lobachebsky University(信息技术、数学与力学学院,洛瓦茨基大学)
;
School of Artificial Intelligence, Beihang University(人工智能学院,北京航空航天大学)