Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
通过语义-空间解耦解决前馈新视角合成变换器中的表示歧义
机构 * Institute of Trustworthy Embodied Artificial Intelligence (TEAI)(可信具身人工智能研究所) ; Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身人工智能重点实验室) ; Sch. of Artificial Intelligence & Sch. of Computer Science, Shanghai Jiao Tong University(上海交通大学人工智能学院与计算机科学学院) ; University of Science and Technology of China(中国科学技术大学)
专题命中 新视角合成 :novel view synthesis(title,abstract);分类 cs.CV
AI总结 本文提出通过语义-空间解耦解决前馈新视角合成变换器中的表示歧义问题,通过分离语义和空间令牌,保持两者的显式表示并利用共享注意力路由保持跨分支交互,同时引入可选分类监督和双向调制以提高交互效果,从而在解码器-only和编码器-解码器前馈NVS模型中实现一致的改进。
Comments 24 pages, 11 figures, 4 tables. Project page: https://hangzay.github.io/ssd_lvsm/