StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
StoryTeller:用于长格式音频描述的免训练叙事接地
机构 * Dartmouth College(达特茅斯学院)
专题命中 长视频与时序推理 :video-language(abstract);分类 cs.CV
AI总结 研究针对长格式音频描述问题,提出免训练的StoryTeller框架。它通过维护叙事记忆跨场景传递信息,仅用原始视频和标题,经语义过滤等确保信息准确。引入StoryAD-QA基准测试,实验证明该框架显著提升了叙事相关能力。
Comments Accepted to the European Conference on Computer Vision (ECCV) 2026