PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
PPLLaVA:通过提示引导的视频序列理解
机构 * Peng Cheng Laboratory(鹏城实验室) ; Peking University(北京大学) ; XPeng Inc.(小鹏汽车有限公司) ; Ministry of Industry and Information Technology of the People’s Republic of China, China Electronics Standardization Institute, Beijing, China(中华人民共和国工业和信息化部,中国电子标准化研究院,北京,中国)
AI总结 PPLLaVA通过提示引导的池化策略实现视频序列高效理解,减少18倍token,提升推理效率并在多种视频理解任务中取得最佳性能。
Comments Accepted to ICLR' 26