Loading...
Thumbnail Image
Item

Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities

Yaneva, Victoria
Orăsan, Constantin
Evans, Richard
Rohanian, Omid
Alternative
Abstract
Given the lack of large user-evaluated corpora in disability-related NLP research (e.g. text simplification or readability assessment for people with cognitive disabilities), the question of choosing suitable training data for NLP models is not straightforward. The use of large generic corpora may be problematic because such data may not reflect the needs of the target population. At the same time, the available user-evaluated corpora are not large enough to be used as training data. In this paper we explore a third approach, in which a large generic corpus is combined with a smaller population-specific corpus to train a classifier which is evaluated using two sets of unseen user-evaluated data. One of these sets, the ASD Comprehension corpus, is developed for the purposes of this study and made freely available. We explore the effects of the size and type of the training data used on the performance of the classifiers, and the effects of the type of the unseen test datasets on the classification performance.
Citation
Journal
Research Unit
DOI
PubMed ID
PubMed Central ID
Embedded videos
Type
Conference contribution
Language
en
Description
The 12th Workshop on Innovative Use of NLP for Building Educational Applications, 8th September 2017 Copenhagen, Denmark.
Series/Report no.
ISSN
EISSN
ISBN
ISMN
Gov't Doc #
Sponsors
Rights
Research Projects
Organizational Units
Journal Issue
Embedded videos