Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models
使用视觉-语言模型进行意大利议会演讲的转录与识别
机构 * Università degli Studi di Milano(米兰大学) ; Department of Social and Political Sciences(社会科学系) ; Department of Literary Studies, Philology and Linguistics(文学研究、语言学与语言学系) ; Department of Computer Science(计算机科学系)
专题命中 文档图表理解 :vision-language model(title,abstract);分类 cs.AI
AI总结 本文提出基于视觉-语言模型的 pipeline,用于自动转录、语义分割和实体链接意大利议会演讲,提升转录质量和发言者标注。
Comments to be published in: ParlaCLARIN V: Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora, organized within the 15th Language Resource and Evaluation Conference (2026)