From Attribution to Action: A Human-Centered Application of Activation Steering
从归因到行动:激活导向的人本应用
Tobias Labarta, Maximilian Dreyer, Katharina Weitz, Wojciech Samek, Sebastian Lapuschkin
机构
*
Fraunhofer Heinrich-Hertz-Institut(弗劳恩霍夫 Heinrich-Hertz 研究所)
;
Technische Universität Berlin(柏林技术大学)
;
BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
面向视觉-语言导航的不确定性感知高斯地图
Jianzhe Gao, Rui Liu, Yuxuan Xu, Tongtong Cao, Yingxue Zhang, Zhanguang Zhang, Sida Peng, Yi Yang, Wenguan Wang
机构
*
The State Key Lab of Brain-Machine Intelligence(脑机智能国家重点实验室)
;
Department of Foundation model, 2012 Labs, Huawei(基础模型部门,2012实验室,华为)
;
Noah’s Ark Lab, 2012 Labs, Huawei(诺亚方舟实验室,2012实验室,华为)
;
School of Software Technology, Zhejiang University(浙江大学软件学院)
Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
揭示视觉-语言模型的脆弱性:通过纹理约束扰动和跨模态优化的多模态对抗协同
Xiang Fang, Wanlong Fang, Changshuo Wang
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
Comments12 pages, 3 figures, accepted at ICMHI 2026, 10th International Conference on Medical and Health Informatics, Kyoto, Japan. To appear in ACM Conference Proceedings
机构
*
School of Intelligence Science and Technology(智能科学与技术学院)
;
State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室)
;
Nanjing University(南京大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.AI、cs.LG
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
看见 vs. 相信:评估开源多模态大模型在反直觉场景中的语言偏见
Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
机构
*
Zhejiang University(浙江大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.CV、cs.AI
CommentsThis paper has been accepted by the International Journal of Computer Vision (IJCV), 2026. The first two authors contributed equally to this work. 28 pages
机构
*
McGill University(麦吉尔大学)
;
Mila - Quebec AI Institute(魁北克人工智能研究所)
;
University of Cambridge(剑桥大学)
;
MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩苏尔·本·扎耶德人工智能大学)
;
University of Toronto(多伦多大学)
;
Salesforce
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
一种用于工业检测中自动缺陷推理与报告生成的混合视觉-语言架构
Malikussaid, Imad Gohar
机构
*
School of Computing, Telkom University(Telkom大学计算机学院)
;
Faculty of Engineering and Technology, School of Computing and Artificial Intelligence(工程与技术学院,计算与人工智能学院)
机构
*
MT Lab, Meitu Inc., Beijing 100083, China(美图实验室,美图公司,北京100083,中国)
;
Department of Computer Science and Technology, BNRist, IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing 100084, China(计算机科学与技术系,BNRist,IDG/麦戈文脑研究学院,清华大学,北京100084,中国)
;
Beijing University of Posts(北京邮电大学)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
FTibSuite:面向藏语视觉语言建模的综合资源套件
Guixian Xu, Yide Liang, Zeli Su, Xuexian Song, Ziyin Zhang, Yushuang Dong, Ting Zhang, Xu Han
机构
*
Hainan International College, Minzu University of China(民族大学海南国际学院)
;
School of Information Engineering, Minzu University of China(民族大学信息工程学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
弥合视觉令牌剪枝中的语义-动作鸿沟以实现高效VLA推理
Ziyan Liu, Yeqiu Chen, Hongyi Cai, Tao Lin, Shuo Yang, Zheng Liu, Bo Zhao
机构
*
School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院)
;
University of Science and Technology of China(中国科学技术大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
BAAI(北京人工智能研究院)
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm
超越文本提示:视觉到视觉生成作为统一范式
Yaofang Liu, Kangning Cui, Meng Chu, Zhaoqing Li, Suiyun Zhang, Jean-Michel Morel, Xiaodong Cun, Haoxuan Che, Rui Liu, Raymond H. Chan
机构
*
City University of Hong Kong(香港城市大学)
;
City University of Hong Kong (Dongguan)(香港城市大学(东莞))
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Celia Research HK(Celia研究香港)
;
Great Bay University(大湾大学)
;
Lingnan University(岭南大学)
BioFact-MoE: Biologically Factorized Mixture of Experts for Vision-Language Prognostic Modeling in Hepatocellular Carcinoma
BioFact-MoE:基于生物学因子分解的混合专家模型用于肝细胞癌的视觉-语言预后建模
Junlin Yang, Tian Yu, Nicha C. Dvornek, Yuexi Du, Peiyu Duan, Annabella Shewarega, Lawrence H. Staib, James S. Duncan, Julius Chapiro
机构
*
Department of Radiology \& Biomedical Imaging, Department of Biomedical Engineering, Department of Electrical Engineering, Department of Statistics \& Data Science Yale University, New Haven, CT, 06510, USA
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
并非所有标记都同等重要:基于关键标记监督的动态上下文向量蒸馏用于长医学报告生成
Ning Wu, Rui Liu, Xinkun Lin, Weixing Chen, Jinxi Xiang, Tao Wei, Lina Yao, Mingjie Li
机构
*
UNSW Sydney(新南威尔士大学悉尼分校)
;
University of Technology Sydney(技术大学悉尼分校)
;
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Stanford University(斯坦福大学)
;
Shanghai Jiao Tong University(上海交通大学)
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
D-OPSD:用于连续调优步蒸馏扩散模型的在线自蒸馏方法
Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du, Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang, Steven Hoi
机构
*
The Hong Kong University of Science and Technology(香港科技大学)
;
Z-Image Team, Alibaba Group(阿里集团Z-Image团队)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
The Chinese University of Hong Kong(香港中文大学)