Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions
超越视觉歧义:通过详细长文本引导具有挑战性场景下的鲁棒单目深度估计
机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
专题命中 其他LLM :language model(abstract,abstract_cn)
AI总结 该研究针对单目深度估计的视觉歧义问题,提出CapDepth框架,通过详细长文本引导,在非朗伯表面和恶劣天气下的深度误差较现有最优方法分别降低25.0%和22.0%。
Comments Accepted to ACM MM 2026