Generative Data Imputation for Sparse Learner Performance Data Using Generative Adversarial Imputation Networks

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study developed a Generative Adversarial Imputation Network (GAIN) to address sparse learner performance data in intelligent tutoring systems, demonstrating improved imputation accuracy and preservation of learning characteristics across multiple ITS datasets.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

As learners engage with Intelligent Tutoring Systems (ITSs) by responding to a series of questions, their performance data, such as correct or incorrect responses, is crucial for assessing and predicting their knowledge states through analysis and modeling. However, data sparsity, often arising from skipped or incomplete responses, poses challenges for accurately assessing learning and delivering personalized instruction. To address this, we propose a generative data imputation method based on Generative Adversarial Imputation Networks (GAIN) to complete missing learning performance data. Our approach employs a three-dimensional (3D) framework structured by learners, questions, and attempts, with an adaptable design along the attempts dimension to manage varying sparsity levels. Enhanced by convolutional neural networks in the input and output layers and optimized with a least squares loss function, our GAIN-based method aligns the input and output shapes with the dimensions of question-attempt matrices across the learners' dimension. Extensive experiments on datasets from three types of ITSs, including AutoTutor Adult Reading Comprehension (ARC), ASSISTments and MATHia, demonstrate that our approach generally outperforms baseline methods, e.g., tensor factorization-based methods and other Generative Adversarial Network (GAN) variants, in imputation accuracy across different setting of maximum attempts. Bayesian Knowledge Tracing (BKT) modeling further validates the imputed data's efficacy by estimating learning parameters, including initial knowledge P (L0), learning rate P (T), guess rate P (G), and slip rate P (S). Results reveal that the imputed data not only enhances model fit but also closely aligns with the original sparse distributions by capturing underlying learning behaviors, indicating greater reliability in learner assessments. Kullback-Leibler (KL) divergence measurements of all these learning parameters confirm that the imputed data effectively preserve essential learning characteristics, maintaining low divergence
Full text 3,218 characters · extracted from oa-doi-fallback · 2 sections · click to expand

Abstract

As learners engage with Intelligent Tutoring Systems (ITSs) by responding to a series of questions, their performance data, such as correct or incorrect responses, is crucial for assessing and predicting their knowledge states through analysis and modeling. However, data sparsity, often arising from skipped or incomplete responses, poses challenges for accurately assessing learning and delivering personalized instruction. To address this, we propose a generative data imputation method based on Generative Adversarial Imputation Networks (GAIN) to complete missing learning performance data. Our approach employs a three-dimensional (3D) framework structured by learners, questions, and attempts, with an adaptable design along the attempts dimension to manage varying sparsity levels. Enhanced by convolutional neural networks in the input and output layers and optimized with a least squares loss function, our GAIN-based method aligns the input and output shapes with the dimensions of question-attempt matrices across the learners' dimension. Extensive experiments on datasets from three types of ITSs, including AutoTutor Adult Reading Comprehension (ARC), ASSISTments and MATHia, demonstrate that our approach generally outperforms baseline methods, e.g., tensor factorization-based methods and other Generative Adversarial Network (GAN) variants, in imputation accuracy across different setting of maximum attempts. Bayesian Knowledge Tracing (BKT) modeling further validates the imputed data's efficacy by estimating learning parameters, including initial knowledge P (L0), learning rate P (T), guess rate P (G), and slip rate P (S). Results reveal that the imputed data not only enhances model fit but also closely aligns with the original sparse distributions by capturing underlying learning behaviors, indicating greater reliability in learner assessments. Kullback-Leibler (KL) divergence measurements of all these learning parameters confirm that the imputed data effectively preserve essential learning characteristics, maintaining low divergence Supplementary Material File (ieeetlt_generativedataimputation_draft_ (9).pdf) - Download - 13.48 MB Information & Authors Information Version history Copyright This work is licensed under a Non Exclusive No Reuse License.

Keywords

Authors Funding Information R305A200413 Metrics & Citations Metrics Article Usage 198views 75downloads Citations Download citation Liang Zhang, Jionghao Lin, John Sabatini, et al. Generative Data Imputation for Sparse Learner Performance Data Using Generative Adversarial Imputation Networks. Authorea. 14 March 2025. DOI: https://doi.org/10.22541/au.174197371.16854905/v1 DOI: https://doi.org/10.22541/au.174197371.16854905/v1 If you have the appropriate software installed, you can download article citation data to the citation manager of your choice. Simply select your manager software from the list below and click Download. For more information or tips please see 'Downloading to a citation manager' in the Help menu.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00