Attention, not scale, drives human-AI alignment in multimodal language prediction
注意力,而非规模,驱动多模态语言预测中的人机对齐
机构 * Psychology and Language Science, Experimental Psychology, University College London, London, UK(心理学与语言科学、实验心理学,伦敦大学学院,伦敦,英国) ; Google Deepmind, Mountain View, US(谷歌DeepMind,山景城,美国) ; Princeton Neuroscience Institute, Princeton University, Princeton, NJ, USA(普林斯顿神经科学研究所,普林斯顿大学,普林斯顿,新泽西州,美国) ; Computer Science Department, Exeter University(计算机科学系,埃克塞特大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI
AI总结 本研究通过比较五种视觉-语言模型与600名人类在视觉世界范式中的表现,发现添加视觉上下文显著提升模型与人类在预测评分上的一致性,且注意力机制而非模型规模是主要驱动因素。
Comments 39 pages, 6 Figures, published in NPJ Artificial Intelligence