Using set visualization techniques to investigate and explain patterns of missing values in electronic health records

preprint OA: gold CC-BY-ND-4.0
📄 Open PDF View at publisher

Abstract

ABSTRACT Objectives Missing data is the most common data quality issue in electronic health records (EHRs). Checks are typically limited to counting the number of missing values in individual fields, but researchers and organisations need to understand multi-field missing data patterns, and counts or numerical summaries are poorly suited to that. This study shows how set-based visualization enables multi-field missing data patterns to be discovered and investigated. Design Development and evaluation of interactive set visualization techniques to find patterns of missing data and generate actionable insights. Setting and participants Anonymised Admitted Patient Care health records for NHS hospitals and independent sector providers in England. The visualization and data mining software was run over 16 million records and 86 fields in the dataset. Results The dataset contained 960 million missing values. Set visualization bar charts showed how those values were distributed across the fields, including several fields that, unexpectedly, were not complete. Set intersection heatmaps revealed unexpected gaps in diagnosis, operation and date fields. Information gain ratio and entropy calculations allowed us to identify the origin of each unexpected pattern, in terms of the values of other fields. Conclusions Our findings show how set visualization reveals important insights about multi-field missing data patterns in large EHR datasets. The study revealed both rare and widespread data quality issues that were previously unknown to an epidemiologist, and allowed a particular part of a specific hospital to be pinpointed as the origin of rare issues that NHS Digital did not know exist. ARTICLE SUMMARY Strengths and limitations of this study This study demonstrates the utility of interactive set visualization techniques for finding and explaining patterns of missing values in electronic health records, irrespective of whether those patterns are common or rare. The techniques were evaluated in a case study with a large (16-million record; 86 field) Admitted Patient Care dataset from NHS hospitals. There was only one data table in the dataset. However, ways to adapt the techniques for longitudinal data and relational databases are described. The evaluation only involved one dataset, but that was from a national organisation that provides many similar datasets each year to researchers and organisations.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-ND-4.0