Santiago Castro
2020
LifeQA: A Real-life Dataset for Video Question Answering
Santiago Castro
|
Mahmoud Azab
|
Jonathan Stroud
|
Cristina Noujaim
|
Ruoyao Wang
|
Jia Deng
|
Rada Mihalcea
Proceedings of The 12th Language Resources and Evaluation Conference
We introduce LifeQA, a benchmark dataset for video question answering that focuses on day-to-day real-life situations. Current video question answering datasets consist of movies and TV shows. However, it is well-known that these visual domains are not representative of our day-to-day lives. Movies and TV shows, for example, benefit from professional camera movements, clean editing, crisp audio recordings, and scripted dialog between professional actors. While these domains provide a large amount of data for training models, their properties make them unsuitable for testing real-life question answering systems. Our dataset, by contrast, consists of video clips that represent only real-life scenarios. We collect 275 such video clips and over 2.3k multiple-choice questions. In this paper, we analyze the challenging but realistic aspects of LifeQA, and we apply several state-of-the-art video question answering models to provide benchmarks for future research. The full dataset is publicly available at https://lit.eecs.umich.edu/lifeqa/.
HAHA 2019 Dataset: A Corpus for Humor Analysis in Spanish
Luis Chiruzzo
|
Santiago Castro
|
Aiala Rosá
Proceedings of The 12th Language Resources and Evaluation Conference
This paper presents the development of a corpus of 30,000 Spanish tweets that were crowd-annotated with humor value and funniness score. The corpus contains approximately 38.6% of humorous tweets with an average score of 2.04 in a scale from 1 to 5 for the humorous tweets. The corpus has been used in an automatic humor recognition and analysis competition, obtaining encouraging results from the participants.
Search
Co-authors
- Mahmoud Azab 1
- Jonathan Stroud 1
- Cristina Noujaim 1
- Ruoyao Wang 1
- Jia Deng 1
- show all...
Venues
- LREC2