The Measurement Reliability Crisis in Event-Related Potential Research: Evidence from Bilingual Language Control

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher

Abstract

Event-related potential (ERP) research exhibits a critical methodological blind spot: inter-condition correlations (ρXY), the parameter that mathematically determines difference wave reliability via the Lord–Novick formula, are systematically unmeasured. This creates fundamental interpretive ambiguity because Classical Test Theory demonstrates that difference wave reliability is inversely related to ρXY, with high correlations suppressing reliability even when constituent waveforms are psychometrically sound. Systematic examination of four foundational ERP studies in bilingual language switching (2001–2016, Ncitations ≈ 1,575) confirms this pattern: none report difference wave reliability or inter-condition correlations, despite these measures being the primary dependent variables isolating language control processes. Sensitivity analysis reveals that this vulnerability extends across the entire plausible parameter space: even when inter-condition correlations are moderate (ρXY = .50–.60), achieving adequate difference wave reliability (ρDD′ ≥ .70) requires constituent reliabilities approaching the ceiling of current ERP methodology (ρ = .85–.88). When experimental control produces high correlations (ρXY ≥ .70, a plausible inference from structural analysis and cross-domain precedent in fMRI and other ERP components, though empirically unverified in language switching), the degradation becomes catastrophic: constituent reliabilities of .85 yield difference wave reliabilities near zero. The conspicuous absence of empirical ρXY measurements and reliability verification means the field cannot determine whether two decades of N2 debate reflect genuine neurocognitive complexity, measurement instability, or both in unknown proportions. If inter-condition correlations prove high in tightly controlled paradigms (as rigorous experimental design would predict), inconsistent findings become expected outcomes of unreliable measurement rather than theoretical puzzles requiring reconciliation. However, problems persist even if correlations are moderate, indicating systematic vulnerability regardless of precise parameter values. Three reforms can address this gap: mandatory reporting of difference wave reliability and inter-condition correlations (ρ≥ .70 for group comparisons, ρ≥ .80 for individual differences research), adoption of Generalizability Theory frameworks that partition variance across multiple sources, and single-trial analytical methods that avoid categorical subtraction. These principles extend beyond bilingual research to all ERP domains using difference waves. Systematic psychometric verification should precede theoretical elaboration; the measurement revolution Parsons (2022) called for requires not novel techniques but rigorous application of established principles.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0