Towards named entity annotation of Latvian National Library corpus

Paikens, Peteris; Auzina, Ilze; Garkaje, Ginta; Paegle, Madara

doi:10.3233/978-1-61499-133-5-169

Abstract

The paper describes a work in progress of building a catalogue of named entities – people, places and organizations – based on a recently digitized large (4.5 billion tokens) Latvian corpus. The authors propose an annotation standard for markup of named entities within Latvian corpus, according to which a representative set of documents (150 000 words) are manually annotated. This corpus is used for training and evaluation of an automated named entity recognition system based on Stanford CRF classifier, achieving an F-score of up to 81%. The named entities indexed within the Latvian National Library corpus and the annnotated documents are publicly available for linguistic and historical research online.

This website uses cookies

This website uses cookies