Using machine learning in learner corpus research: error annotation and (extra-)linguistic metadata as proxies for approaching the interlanguage
preprint
OA: closed
Abstract
This article investigates the potential of using error-annotated L2 utterances and their associated language profile to more accurately assess the language proficiency of an L2 learner. The main goal of this paper is to demonstrate the use of machine learning methods for uncovering features of the learner's interlanguage and their proficiency level, as well as to answer three questions: a) is it feasible to use error-annotated data along with linguistic profile metadata to trace patterns related to aspects of learners’ interlanguage? b) How important may error annotation be in estimating the learner’s proficiency level? c) Which of the (extra-)linguistic profile metadata stand out in contributing to the correct classification of the proficiency level? GLCII, the largest online freely available L2-Greek corpus, is utilized as an example corpus to illustrate the approach proposed in this paper with which for any L2, central questions can potentially be addressed by the relevant machine learning method so that, in a second stage, useful conclusion can be drawn. As a final point, we urge the implementation of ML techniques as a way of expediting research in a variety of related areas in applied linguistics.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00