MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation
MV-Actor:对齐多视角语义与空间感知以实现双臂操作
机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) ; Institute for AI Industry Research (AIR), Tsinghua University(清华大学智能产业研究院) ; AIR Wuxi Innovation Center, Tsinghua University(清华大学智能产业研究院无锡创新中心)
AI总结 提出MV-Actor框架,通过多视角语义交互和语义-空间令牌交互统一语义与空间表示,并利用引导度量深度修复模块处理深度噪声,在PerAct2基准上达到87.8%平均成功率。
Comments 14 pages,9 figures