FlorUniTo@TRAC-2: Retrofitting Word Embeddings on an Abusive Lexicon for Aggressive Language Detection
Anna Koufakou, Valerio Basile, Viviana Patti
Abstract
This paper describes our participation to the TRAC-2 Shared Tasks on Aggression Identification. Our team, FlorUniTo, investigated the applicability of using an abusive lexicon to enhance word embeddings towards improving detection of aggressive language. The embeddings used in our paper are word-aligned pre-trained vectors for English, Hindi, and Bengali, to reflect the languages in the shared task data sets. The embeddings are retrofitted to a multilingual abusive lexicon, HurtLex. We experimented with an LSTM model using the original as well as the transformed embeddings and different language and setting variations. Overall, our systems placed toward the middle of the official rankings based on weighted F1 score. However, the results on the development and test sets show promising improvements across languages, especially on the misogynistic aggression sub-task.- Anthology ID:
- 2020.trac-1.17
- Volume:
- Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying
- Month:
- May
- Year:
- 2020
- Address:
- Marseille, France
- Venues:
- LREC | TRAC | WS
- SIG:
- Publisher:
- European Language Resources Association (ELRA)
- Note:
- Pages:
- 106–112
- URL:
- https://www.aclweb.org/anthology/2020.trac-1.17
- DOI:
- PDF:
- https://www.aclweb.org/anthology/2020.trac-1.17.pdf
You can write comments here (and agree to place them under CC-by). They are not guaranteed to stay and there is no e-mail functionality.