Nikola Ljubešić
2020
Gigafida 2.0: The Reference Corpus of Written Standard Slovene
Simon Krek
|
Špela Arhar Holdt
|
Tomaž Erjavec
|
Jaka Čibej
|
Andraz Repar
|
Polona Gantar
|
Nikola Ljubešić
|
Iztok Kosem
|
Kaja Dobrovoljc
Proceedings of The 12th Language Resources and Evaluation Conference
We describe a new version of the Gigafida reference corpus of Slovene. In addition to updating the corpus with new material and annotating it with better tools, the focus of the upgrade was also on its transformation from a general reference corpus, which contains all language variants including non-standard language, to the corpus of standard (written) Slovene. This decision could be implemented as new corpora dedicated specifically to non-standard language emerged recently. In the new version, the whole Gigafida corpus was deduplicated for the first time, which facilitates automatic extraction of data for the purposes of compilation of new lexicographic resources such as the collocations dictionary and the thesaurus of Slovene.
CoSimLex: A Resource for Evaluating Graded Word Similarity in Context
Carlos Santos Armendariz
|
Matthew Purver
|
Matej Ulčar
|
Senja Pollak
|
Nikola Ljubešić
|
Mark Granroth-Wilding
Proceedings of The 12th Language Resources and Evaluation Conference
State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous measures of meaning similarity. This paper describes an effort to build a new dataset, CoSimLex, intended to fill this gap. Building on the standard pairwise similarity task of SimLex-999, it provides context-dependent similarity measures; covers not only discrete differences in word sense but more subtle, graded changes in meaning; and covers not only a well-resourced language (English) but a number of less-resourced languages. We define the task and evaluation metrics, outline the dataset collection methodology, and describe the status of the dataset so far.
Search
Co-authors
- Simon Krek 1
- Špela Arhar Holdt 1
- Tomaž Erjavec 1
- Jaka Čibej 1
- Andraz Repar 1
- show all...
Venues
- LREC2