Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding
感知空间:面向高效准确3D场景理解的自我运动感知视频表示
机构 * Department of Computer Science(计算机科学系) ; University of Michigan(密歇根大学) ; University of Michigan Ann Arbor(密歇根大学安阿伯分校)
专题命中 视觉推理 :MLLM(summary_cn,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 本文提出Motion-MLLM框架,结合IMU数据与视觉特征,通过运动-视觉关键帧过滤模块和异构跨模态融合模块,提升3D场景理解与空间推理的效率和准确性。
Comments 22 pages, 10 figures