Scalable and Explainable Learner-Video Interaction Prediction using Multimodal Large Language Models
基于多模态大语言模型的可扩展且可解释的学习者-视频交互预测
机构 * EPFL(瑞士联邦理工学院洛桑)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 本文提出利用多模态大语言模型预测学习者观看、暂停、跳过和回放行为,以评估视频内容的认知负荷,通过7700万次视频控制事件验证了模型的可扩展性和解释性。
Comments Accepted as long paper to the 27th International Conference on Artificial Intelligence in Education (AIED 2026)