Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark
易完成,难选择:探究大型语言模型在ProverbIT基准上的性能
AI总结 该研究构建了意大利谚语基准ProverbIT,评估13个前沿LLMs的谚语处理能力,发现LLMs虽能完成谚语任务,但在无正确答案的选择题中性能骤降,暴露其依赖记忆而非深层语义理解的局限。
Journal ref Proceedings of the Eleventh Italian Conference on Computational Linguistics (CLiC-it 2025), pages 722-734