S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
S$^2$-VLA:面向长时程操作的基于状态空间的视觉-语言-动作模型
机构 * School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) ; State Key Laboratory of Submarine Geoscience, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院海底科学国家重点实验室)
专题命中 机器人基础模型 :manipulation(title,abstract);robotic(abstract);分类 cs.RO、cs.AI
AI总结 针对长时程操作任务中累积误差导致性能下降的问题,提出S$^2$-VLA框架,通过状态空间引导的自适应注意力机制动态融合视觉、语言和动作信息,在LIBERO等基准上超越7B模型。
Comments Accepted to IJCAI 2026