TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment
TEVI: 基于稀疏自编码器的文本条件视觉表示编辑以改进视觉-语言对齐
Sweta Mahajan, Sukrut Rao, Jiahao Xie, Alexander Koller, Bernt Schiele
机构
*
Max Planck Institute for Informatics, Saarland Informatics Campus, Saarbrücken, Germany(马克斯·普朗克研究所信息学院,萨尔兰信息学院,德国萨尔布吕肯)
;
Department of Language Science and Technology, Saarland University, Saarbrücken, Germany(语言科学与技术系,萨尔兰大学,德国萨尔布吕肯)
CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection
CL-CLIP: 基于CLIP的持续学习框架与代价体积类别解耦用于目标检测
Zihan Liu, Yuguang Yang, Shengjie Su, Jianing Pang, Linlin Yang, Chunyu Xie, Nikolai Yu. Zolotykh, Baochang Zhang
机构
*
National College for Excellent Engineers, Beihang University(卓越工程师学院,北京航空航天大学)
;
AI Research, Qihoo 360(360人工智能研究院,奇虎360)
;
School of Electronic Information Engineering, Beihang University(电子信息学院,北京航空航天大学)
;
School of Cyber Science and Technology, Beihang University(网络安全科学与技术学院,北京航空航天大学)
;
School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北京航空航天大学)
;
State Key Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播国家重点实验室,中国传媒大学)
;
Institute of Information Technology, Mathematics and Mechanics, Lobachebsky University(信息技术、数学与力学学院,洛瓦茨基大学)
;
School of Artificial Intelligence, Beihang University(人工智能学院,北京航空航天大学)
M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition
M2S-AVSR:面向鲁棒视听语音识别的模态感知多视角自监督表示
Fei Su, Cancan Li, Ming Li, Juan Liu
机构
*
School of Artificial Intelligence and the School of Computer Science, Wuhan University, China(人工智能学院和计算机科学学院,武汉大学,中国)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(人工智能学院,香港中文大学(深圳),中国)
;
School of Artificial Intelligence, Wuhan University, China(人工智能学院,武汉大学,中国)
Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Watch, Remember, Reason: 基于多模态大语言模型的人类视角视频理解
Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
机构
*
School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院)
;
Wuhan University(武汉大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Nanyang Technological University(南洋理工大学)
;
CASIA(中国科学院自动化研究所)
;
University of Tokyo(东京大学)
;
University of Liverpool(利物浦大学)
;
Zhejiang University(浙江大学)
;
National University of Singapore(新加坡国立大学)
;
UC Merced(加州大学默塞德分校)
机构
*
institutetext: MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models Yue Wu Changyuan Wang Zixuan Wang Shilin Ma Yansong Tang(机构文本:MorphoQuant:多模态大语言模型的模态感知量化 Yue Wu 王昌元 王梓轩 马世林 唐彦松)
TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance
TraRA: 面向城市监控视频文本识别的轨迹级识别聚合方法
Duc Tri Tran, Trung Thanh Nguyen, Vijay John, Phi Le Nguyen, Yasutomo Kawanishi
机构
*
RIKEN(日本理化学研究所)
;
Hanoi University of Science and Technology(河内科学技术大学)
;
Nagoya University(名古屋大学)
;
Lawrence Technological University(劳伦斯技术大学)
;
Ritsumeikan University(立命馆大学)
CommentsThis paper has been accepted for publication in the Proceedings of the IEEE Radar Conference (RadarConf 2026). The final authenticated version will be available through IEEE Xplore