Early economic evaluation of retrieval-layer correction in clinical RAG: a decision-uncertainty framework

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Embedding geometry degradation is common in clinical retrieval-augmented generation (RAG) systems and clearly reduces retrieval accuracy. Corpus-only ZCA whitening is a no-retraining correction that boosts retrieval accuracy on diverse clinical text, but its cost-effectiveness depends on whether these improvements lead to better clinical outcomes, a connection that has not yet been empirically confirmed in RAG settings. Objective To quantify the conditions where a low-cost retrieval-layer intervention could be economically viable, and to identify the empirical parameters whose measurement would most decrease decision uncertainty. Methods An exploratory decision model with explicit structural gating was developed from a healthcare system perspective (Norwegian reference case, 4% discount rate, 5-year horizon). Whitening effectiveness was modeled across two corpus branches: beneficial on heterogeneous corpora (base-case ΔMRR = + 0.221); harmful on homogeneous corpora (ΔMRR = − 0.05). The surrogate link from retrieval improvement to diagnostic accuracy (α) was empirically estimated from the DiReCT dataset (MIMIC-IV-Ext-DiReCT, NeurIPS 2024): 511 physician-annotated clinical notes from MIMIC-IV, with ZCA whitening applied to ClinicalBERT embeddings and measuring change in primary discharge diagnosis retrieval accuracy. The primary outputs are scale-independent: minimum annual query volume (N*) for cost-effectiveness, and outcomes per 1,000 queries. Results A DiReCT-based retrieval experiment estimated an empirical α = 1.111 (95% CI [1.014, 2.541]; ClinicalBERT, PDD-level) in a diagnosis-label retrieval setting, replacing the transported Tao et al. CDSS estimates (0.36) as the base case; the Tao et al. estimate is maintained as the conservative scenario. The experiment used 343 MIMIC-IV clinical notes with sufficient text content (from the full DiReCT dataset of 511 annotated notes). The minimum N* for whitening to cover its €800 implementation cost is 6 annual queries at the base-case parameters and 18 at the conservative α = 0.36, thresholds that are low compared to typical institutional deployment scales. Per 1,000 annual queries, whitening prevents 4.74 adverse diagnostic events (base case) or 1.53 (conservative), resulting in €253,008 or €81,983 in healthcare savings over 5 years, respectively. These estimates depend on whether improvements in diagnosis-label retrieval accuracy translate into actual clinician diagnostic performance, a structural assumption the DiReCT experiment does not itself address. Conclusions This framework shows that whitening appears economically plausible across the modelled cost structure. The DiReCT experiment provides an empirical α estimate in a clinical-note retrieval task with diagnosis-label relevance, substantially above the previously transported CDSS estimate (α = 0.36), which is retained as the conservative scenario. The remaining structural uncertainty, whether diagnosis-label retrieval translates to clinician diagnostic performance, would require a case-level linkage study with adequate causal identification to resolve.
Full text 115,474 characters · extracted from preprint-html · click to expand
Early economic evaluation of retrieval-layer correction in clinical RAG: a decision-uncertainty framework | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Early economic evaluation of retrieval-layer correction in clinical RAG: a decision-uncertainty framework Yngve Mikkelsen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9237671/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Embedding geometry degradation is common in clinical retrieval-augmented generation (RAG) systems and clearly reduces retrieval accuracy. Corpus-only ZCA whitening is a no-retraining correction that boosts retrieval accuracy on diverse clinical text, but its cost-effectiveness depends on whether these improvements lead to better clinical outcomes, a connection that has not yet been empirically confirmed in RAG settings. Objective To quantify the conditions where a low-cost retrieval-layer intervention could be economically viable, and to identify the empirical parameters whose measurement would most decrease decision uncertainty. Methods An exploratory decision model with explicit structural gating was developed from a healthcare system perspective (Norwegian reference case, 4% discount rate, 5-year horizon). Whitening effectiveness was modeled across two corpus branches: beneficial on heterogeneous corpora (base-case ΔMRR = + 0.221); harmful on homogeneous corpora (ΔMRR = − 0.05). The surrogate link from retrieval improvement to diagnostic accuracy (α) was empirically estimated from the DiReCT dataset (MIMIC-IV-Ext-DiReCT, NeurIPS 2024): 511 physician-annotated clinical notes from MIMIC-IV, with ZCA whitening applied to ClinicalBERT embeddings and measuring change in primary discharge diagnosis retrieval accuracy. The primary outputs are scale-independent: minimum annual query volume (N*) for cost-effectiveness, and outcomes per 1,000 queries. Results A DiReCT-based retrieval experiment estimated an empirical α = 1.111 (95% CI [1.014, 2.541]; ClinicalBERT, PDD-level) in a diagnosis-label retrieval setting, replacing the transported Tao et al. CDSS estimates (0.36) as the base case; the Tao et al. estimate is maintained as the conservative scenario. The experiment used 343 MIMIC-IV clinical notes with sufficient text content (from the full DiReCT dataset of 511 annotated notes). The minimum N* for whitening to cover its €800 implementation cost is 6 annual queries at the base-case parameters and 18 at the conservative α = 0.36, thresholds that are low compared to typical institutional deployment scales. Per 1,000 annual queries, whitening prevents 4.74 adverse diagnostic events (base case) or 1.53 (conservative), resulting in €253,008 or €81,983 in healthcare savings over 5 years, respectively. These estimates depend on whether improvements in diagnosis-label retrieval accuracy translate into actual clinician diagnostic performance, a structural assumption the DiReCT experiment does not itself address. Conclusions This framework shows that whitening appears economically plausible across the modelled cost structure. The DiReCT experiment provides an empirical α estimate in a clinical-note retrieval task with diagnosis-label relevance, substantially above the previously transported CDSS estimate (α = 0.36), which is retained as the conservative scenario. The remaining structural uncertainty, whether diagnosis-label retrieval translates to clinician diagnostic performance, would require a case-level linkage study with adequate causal identification to resolve. Health Economics and Outcomes Research retrieval-augmented generation clinical informatics early economic evaluation value of information embedding geometry ZCA whitening decision uncertainty Figures Figure 1 1. Why is retrieval infrastructure a health economics problem Retrieval-augmented generation (RAG) has become a key approach for grounding large language model (LLM) outputs in verified clinical evidence [ 1 , 2 ]. In a clinical RAG pipeline, an embedding model converts queries and clinical documents into vector representations; cosine similarity then finds the most relevant documents to provide as context to the LLM. Retrieval quality depends on embedding geometry: when embeddings cluster into a narrow cone (anisotropy), cosine similarity loses its discrimination ability, causing retrieval to fail regardless of the quality of generation [ 3 ]. A companion benchmarking study found that context variables together explain 49.0% of the variance in MRR@10 across 294 experimental conditions on three clinical corpora. It also showed that domain-specific models (BioBERT [ 4 ], ClinicalBERT [ 5 ]) exhibited nearly degenerate geometry despite biomedical pretraining [ 6 ]. A layer-level analysis across 1,400 conditions revealed a U-shaped performance curve, with mid-layer collapse and variable recovery at the final layer [ 7 ]. Corpus-only ZCA whitening, a linear transformation applied to document embeddings without retraining the model, improved MRR@10 by + 0.16 to + 0.27 on degraded models across diverse clinical text conditions but reduced performance on structurally uniform corpora. If improvements in retrieval lead to better clinical outcomes, these empirically demonstrated gains could have economic value. Diagnostic errors are estimated to occur in 10–15% of clinical encounters and contribute to 6–17% of adverse events [ 8 , 9 ]. Clinical decision support systems (CDSS) embedded in electronic health records have been shown to decrease unnecessary healthcare utilization in multiple systematic reviews [ 10 ]. The question of whether RAG retrieval quality significantly influences clinical outcomes and whether investing in infrastructure-layer improvements is justified is therefore an HEOR question, not just a technical one. Although ZCA whitening is a linear algebra operation on a vector database rather than a traditional health technology, its downstream effects on a diagnostic AI pipeline that patients interact with place it within the realm of early technology assessment for digital health interventions. The NICE Evidence Standards Framework for Digital Health Technologies [ 11 ] and a systematic review of valuation methods for digital health interventions [ 12 ] both include infrastructure-level components whose clinical impact needs prospective validation, which is the approach adopted here. However, conventional cost-effectiveness analysis is premature at this stage of the evidence. No study has empirically connected improvements in retrieval metrics in a clinical RAG pipeline to measurable changes in diagnostic accuracy or patient outcomes. The surrogate chain from MRR@10 to adverse diagnostic events depends on an unmeasured link. This paper does not claim otherwise. Instead, it asks a question that can be answered with current evidence: under what conditions could a low-cost retrieval-layer intervention be economically attractive, and which empirical parameters most influence that conclusion? 2. Objective This paper constructs an exploratory decision model to answer three related questions: What is the minimum annual query volume (N*) at which ZCA whitening recovers its implementation cost, across a range of plausible surrogate link values? What values of the three most uncertain parameters, the surrogate link (α), harm magnitude in homogeneous corpora, and clinician adoption, would make whitening not economically attractive? What structural uncertainties are not captured by probabilistic sensitivity analysis, and what case-level linkage study would most efficiently reduce decision uncertainty? The fine-tuning comparator (a more resource-intensive model adaptation option) is addressed in the Supplement. The primary comparison is whitening versus no intervention. 3. Methods 3.1 Model structure and population An exploratory decision tree was created for a single deployment choice: apply corpus-only whitening or do nothing. The model population consists of clinical RAG deployments that have been checked for embedding degradation and corpus type, as detailed below. Fine-tuning as an active comparator is addressed in the Supplement. Whitening effectiveness is modeled across two branches of the degraded-model population. In the degraded-and-heterogeneous branch (estimated at 30.2% of all new deployments: 0.45 × 0.67), whitening improves retrieval accuracy. In the degraded-and-homogeneous branch (14.9%: 0.45 × 0.33), whitening causes a slight reduction in retrieval accuracy (ΔMRR = − 0.05, base case). This two-branch structure was introduced to address a directional bias in previous versions of this model, which set the harm from the homogeneous corpus to zero. The harm branch draws on empirical data showing that whitening degraded performance on all eight structurally uniform corpus conditions tested [ 7 ]. In 55% of deployments with non-degraded models, no intervention or benefit is observed across all arms. The surrogate link α relates the improvement in MRR@10 to the gain in diagnosis-label retrieval accuracy, which is defined as a top-1 match between the retrieved document’s primary discharge diagnosis and the query’s PDD. This is different from clinician diagnostic accuracy, as discussed in Limitation 1, and remains a key structural uncertainty in the model. In this version, α is based on an empirical estimate from a DiReCT-based diagnosis-label retrieval experiment (α = 1.111, ClinicalBERT, MIMIC-IV clinical notes), with the Tao et al. CDSS estimate (α = 0.36) kept as the conservative lower bound. Neither estimate has been validated in a live clinical workflow; this structural issue is discussed in Section 4.3 . Healthcare system perspective, Norwegian reference case [ 13 ]. Five-year time horizon, 4% annual discount rate. All costs in EUR. Parameter sources and distributions are summarized in Table 1 . 3.2 Primary outputs: scale-independent results All primary results are expressed per 1,000 annual queries or as minimum query-volume thresholds to separate findings from the highly uncertain deployment-scale parameter N. Absolute costs for illustrative deployment sizes are shown in Supplement Table S1. The reason is that N varies by about five orders of magnitude across plausible deployment scenarios (from a small clinic to a large health system), making absolute estimates useful as rough examples rather than precise forecasts. 3.3 Threshold analyses Three threshold analyses determine the parameter values where whitening becomes economically unviable. For each threshold, two figures are presented: the threshold value and the difference between the base-case parameter and the threshold. A large difference indicates robustness; a small difference points to a decision-sensitive zone. The surrogate link threshold (α*) is the value of α below which whitening does not recoup its implementation cost at a specific N. The harm threshold is the homogeneous-corpus ΔMRR magnitude above which whitening fails. The adoption threshold is the minimum clinician adoption rate needed for a positive net monetary benefit. 3.4 Structural uncertainty analysis Probabilistic sensitivity analysis (PSA, 10,000 iterations) samples parameter uncertainty within the model’s fixed causal structure. It does not address structural uncertainty, which concerns whether the assumed relationships are valid. Table 2 formally describes the structural assumptions, their possible failure modes, the direction of bias if they fail, and the type of study needed to test them. This table is the equivalent, for a RAG infrastructure model, of what a formal validity assessment provides for a traditional surrogate endpoint [ 14 ]. 3.5 Sensitivity analyses One-way sensitivity analysis (tornado diagram) varied each parameter between its 5th- and 95th-percentile values. PSA used truncated normal distributions for the surrogate link (α, anchored at the DiReCT empirical estimate, bounds [0.36, 2.54]) and the harm parameter (ΔMRR_harm, bounds [− 0.20, 0.0]), and beta and gamma distributions for probabilities and costs, respectively. A joint chain sensitivity analysis is not included but is a recognized limitation: the tornado diagram treats parameters independently, whereas the real structural uncertainty lies in the multiplicative chain (p_degraded × p_heterogeneous × ΔMRR × α × p_adopt × p_ADE × cost_ADE). Readers should note that the individually low N* thresholds are an artifact of this chain operating at base-case values across all links; a joint stress test that sets multiple links simultaneously to pessimistic values would yield higher N*. A discount rate of 4% is used, following the Norwegian reference case [ 13 ]. As a sensitivity, 3.0% and 3.5% (NICE and ISPOR international reference rates) were tested; results are presented in Supplement Table S1. A societal perspective scenario using NPE-derived costs (Norsk pasientskadeerstatning annual reports 2021–2025, average approved payout NOK 765,481 ÷ 11.50 NOK/EUR = €66,564) is reported separately; NPE payouts include pain/suffering, lost earnings, and future care and are not equivalent to direct healthcare system costs. All analyses were conducted in Python 3.10; code is available at the companion repository [ 15 ]. Table 1 Model parameters, base-case values, PSA distributions, and evidence sources. Parameter Base PSA distribution Source / note Effectiveness Mean ΔMRR@10 — whitening, heterogeneous corpus + 0.221 N(0.221, 0.035) Mikkelsen 2026 JAMIA Table 4 : mean of four degraded models [ 7 ] Whitening ΔMRR — homogeneous corpus (harm; negative) −0.05 Trunc N(− 0.05, 0.04) [− 0.20, 0.0] Mikkelsen 2026 JAMIA Suppl. Table S2: harmful on all 8 Synthetic corpus conditions. Magnitude uncertain; −0.05 is conservative [ 7 ] P(corpus heterogeneous) 0.67 Beta(8, 4) Mikkelsen 2026 JAMIA: whitening positive on 16/24 conditions (heterogeneous corpora only) [ 7 ] P(model degraded) 0.45 Beta(5, 6) 4 of 11 model configurations showed degraded geometry in companion study [ 7 ] Surrogate link (empirically estimated in diagnosis-label retrieval task; clinical workflow link unvalidated) Surrogate link α (Δ diagnosis-label retrieval accuracy per unit ΔMRR@10; see Limitation 1 for distinction from clinical diagnostic accuracy) 1.111 Trunc N(1.11, 0.37) [0.36, 2.54] DiReCT experiment (MIMIC-IV-Ext-DiReCT, NeurIPS 2024): ZCA whitening on ClinicalBERT embeddings, 343 MIMIC-IV clinical notes, PDD-level retrieval accuracy. α = 1.111, 95% CI [1.014, 2.541]. PSA lower bound = 0.36 (Tao et al. 2020 conservative floor). α > 1 is mechanistically expected: whitening produces top-1 rank promotions, where Δacc@1 is binary but ΔMRR is fractional. Clinician adoption rate 0.60 Beta(9, 6) Kawamoto et al. BMJ 2005 systematic review of CDSS features and trial outcomes [ 15 ]. Value of 0.60 is a constructed assumption; no single empirical study directly supports this figure for RAG-based diagnostic tools. The threshold analysis confirms adoption uncertainty is not binding at any modelled deployment scale. Clinical outcomes P(adverse event | diagnostic error) 0.12 Beta(6, 44) 6–17% of adverse events involve diagnostic error (midpoint 12%) [ 8 ] QALY loss per ADE 0.06 Gamma(1.2, 0.05) Severity-weighted: fatal (5%), serious (15%), moderate (30%), minor (50%). Constructed from IOM mortality data [ 16 ]. Healthcare cost per ADE — system perspective (€) €12,000 Gamma(2.0, 6000) Norwegian DRG + OECD diagnostic safety report [ 8 ]. Excess hospitalisation 2–5 days. Range €5k–30k in PSA. ADE cost — societal perspective (NPE) (€) €66,564 Gamma(56.4, 1180) NPE annual reports 2021–2025: NOK 765,481 ÷ 11.50 NOK/EUR. Societal perspective; includes pain/suffering, lost earnings. Sensitivity scenario only. Technology costs Whitening: one-time implementation (€) €800 Gamma(2.0, 400) 1–2 engineer-days. Marginal compute ≈ 0 (NumPy matrix multiply) [ 15 ] Whitening: annual maintenance (€) €150 Gamma(1.5, 100) Periodic corpus refit as document index evolves All monetary values in EUR. N/A = not applicable in PSA (fixed for scenario analysis only). 4. Results 4.1 Scale-independent primary results Table 2 shows the main results per 1,000 annual queries, excluding fixed implementation costs. These findings are independent of assumptions about deployment scale. At baseline parameters (α = 1.111, based on the empirical DiReCT estimate), whitening prevents 4.74 adverse diagnostic events (ADEs) annually and saves €253,008 in healthcare costs over 5 years for each 1,000 queries. With the Tao et al. [ 17 ] conservative α = 0.36, these numbers decrease to 1.53 ADEs and €81,983, which are the same as in the previous version of this model, now shown as a scenario rather than the baseline. The decrease from the zero-harm estimate reflects active modeling of whitening’s impact on retrieval quality in uniform corpora. Table 2 Primary results per 1,000 annual queries (5-year horizon, healthcare system perspective). Excludes fixed implementation costs (€800 one-time). Metric Zero-harm (naïve) Base case (α = 1.111, harm = − 0.05) Conservative (α = 0.36, Tao et al.) ADEs averted per year (net) 5.33 4.74 1.53 Healthcare saving 5 year (€) €284,738 €253,008 €81,983 QALY gain 5yr 1.4216 1.2646 0.4099 Zero-harm column sets homogeneous-corpus harm to zero. Base-case column reflects harm modelling. Conservative column uses Tao et al. transported α = 0.36 as sensitivity scenario. 4.2 Threshold analyses Table 3 shows the three main threshold analyses. The main finding is the imbalance between parameter and structural uncertainty. Table 3 Threshold analyses: parameter values at which whitening is no longer cost-effective (system perspective, N = 133,000/yr unless stated). Parameter Base case Threshold value Margin Interpretation Surrogate link α 1.111 0.00002 1.111 Even a negligible link suffices at N = 133k. Decision hinges on whether α = 0 exactly — a structural, not parametric, question. Harm magnitude |ΔMRR| in homogeneous corpus 0.05 0.449 0.399 Harm would need to be 9× larger than base-case to make whitening not cost-effective. Robust. Clinician adoption p_adopt 0.60 0.00001 0.600 Effectively zero threshold at N = 133k given low tech cost. Minimum N* (α = 1.111, harm = − 0.05) 133,000/yr 6 queries/yr N/A Breakeven query volume at base-case α. Lower than v3 (was 18) because empirical α is larger. Minimum N* (α = 0.36, Tao et al.) 133,000/yr 18 queries/yr N/A Breakeven at Tao et al. conservative α — now a scenario, not base case. Minimum N* (α = 0.15, stress-test) 133,000/yr 43 queries/yr N/A Breakeven under extreme stress-test. Still below any real deployment. Table 4 presents minimum N* across a grid of α and harm values, enabling readers to apply their own priors about the surrogate link. Table 4 Minimum annual query volume N* for whitening to recover €800 implementation cost, by surrogate link (α) and homogeneous-corpus harm magnitude. α Harm = − 0.05 Harm = − 0.20 Harm = − 0.45 Decision implication 0.05 (extreme stress) 129 207 ∞ Fails only at catastrophic harm + negligible link 0.15 (stress-test) 43 69 ∞ Below the scale of most deployed systems even under extreme assumptions 0.36 (Tao et al. conservative) 18 29 ∞ Old base case — now conservative scenario 1.111 (DiReCT base case) 6 9 ∞ New base case: any deployment with > 9 queries/yr 2.54 (DiReCT 97.5th pct) 3 4 ∞ Upper CI bound: effectively all deployments ∞ = harm exceeds benefit regardless of N at this α. DiReCT 95% CI for α: [1.014, 2.541]. Tao et al. α = 0.36 is retained as conservative scenario (lower PSA bound). 4.3 Structural uncertainty analysis Table 5 details the structural assumptions underlying the model. These are assumptions that PSA cannot test because it assesses parameter uncertainty within a fixed causal framework. Each row lists an assumption, its potential failure mode, the bias direction if it fails, and the type of evidence that would test it. Table 5 Structural uncertainties not resolved by probabilistic sensitivity analysis. Structural assumption How it could fail Bias direction if false Study type required to test ΔMRR improvement in retrieval translates to improved diagnostic accuracy (surrogate link α > 0) RAG generation may compensate for retrieval quality; clinicians may not act on retrieved context; context relevance may not determine diagnosis; diagnosis-label retrieval (DiReCT) may not translate to clinician diagnostic performance Overstates benefit; entire economic case collapses if α = 0 Case-level linkage study: retrieval quality (MRR@10) measured alongside diagnostic accuracy outcomes on the same clinical encounters in a live RAG deployment. May be prospective or retrospective with logged retrieval data. Single highest-priority evidence need. Harm in homogeneous branch is proportional to α (same surrogate multiplier) Harm from whitening on uniform corpora may operate through a different mechanism than benefit Understates or overstates harm depending on direction Experimental: measure ΔADE rate on deployments identified as using whitening on homogeneous corpora No hallucination-induced harm from higher-ranked but incorrect documents A document retrieved more confidently by a whitened model may be more convincingly wrong Overstates benefit; could produce net harm Prospective: measure hallucination and incorrect-document-surfacing rates pre/post whitening Fine-tuning effectiveness (ΔMRR = 0.35) is transportable to new deployments Fine-tuning on available clinical corpora may not generalise to deployment corpus Overstates or understates comparative advantage of fine-tuning (supplement only) Head-to-head pre/post study of fine-tuned vs whitened models on deployment corpus Clinician adoption is independent of retrieval improvement magnitude Higher retrieval quality may increase adoption non-linearly; adoption may be zero in fully automated pipelines Could under- or overstate benefit depending on adoption curve shape Observational: measure clinician engagement with RAG-surfaced evidence as function of retrieval quality Deployment prevalence weights (p_degraded, p_heterogeneous) reflect real deployment populations Models are not randomly selected; heterogeneous corpora may be over- or under-represented Could over- or understate population-level impact Survey study: audit of embedding model and corpus choices in deployed clinical RAG systems 4.4 Probabilistic sensitivity analysis The PSA (10,000 iterations) classified whitening as dominant (cost-saving and QALY-gaining) in 99.7% of iterations. The PSA distribution for α is centered on the empirical estimate from DiReCT (Truncated N(1.11, 0.37), bounds [0.36, 2.54]), with the lower bound set at the conservative value from Tao et al. The remaining 0.3% of iterations are non-dominant, representing extreme cases of strong homogeneous-corpus harm and low surrogate link at the PSA boundary. An important caveat is that the PSA distribution for α intentionally excludes zero. The possibility that α = 0—indicating that retrieval improvement produces no diagnostic benefit—is not sampled in PSA because it falls outside the empirically bounded distribution. The 99.7% dominance rate should be viewed as conditional: “if α is substantively positive, whitening dominates with high probability.” It does not address the prior probability that α is positive, which is the main structural question discussed in Section 4.3 . The cost-effectiveness acceptability curve (CEAC) showed over a 99% probability of cost-effectiveness above €5,000/QALY, reflecting the same condition. The DiReCT experiment offers empirical support for the lower bound of the α distribution but does not clarify whether diagnosis-label retrieval in annotated notes improves clinician diagnostic performance in real workflows. That structural question remains unchanged by PSA. 4.5 Societal perspective and NPE scenario From a societal perspective, using NPE-derived costs (EUR 66,564 average from Norsk pasientskadeerstatning annual reports 2021–2025), the minimum N* drops to 2 annual queries under base-case parameters. The societal scenario strengthens the economic case but is based on a fundamentally different cost concept: compensation payouts, including pain, suffering, and lost earnings, which are not suitable as the primary estimate from a healthcare system perspective. It is reported to demonstrate the upper limit of economic attractiveness if payers account for societal costs. 5. Discussion 5.1 What the framework shows The main finding is asymmetric: within-model parameter uncertainty is managed at low query volumes under the modeled cost structure (N* = 6–43 depending on assumptions), while uncertainty about the structure of the surrogate chain remains unresolved by the model. This asymmetry itself offers insight. It shows that if a decision-maker accepts the surrogate assumption at any plausible level, whitening becomes economically viable across the modeled deployment range without further analysis. Conversely, if they do not accept it, no amount of modeling or parameter sensitivity analysis will change that. This presents a different framing than asking “is whitening cost-effective?” The answer to that question depends entirely on a single unverified assumption. The correct framing is: “if the surrogate link exists at any plausible magnitude (α ≥ 0.05), whitening is cost-effective across a broad range of modeled deployment scales; the real question is whether the link exists.” The threshold and N* analyses in Tables 3 and 4 clarify this structure and enable readers with different priors about the surrogate to draw their own conclusions. A health economist who believes α might be near zero will see the model as underpowered for adoption; someone who thinks any RAG retrieval improvement provides downstream diagnostic benefits will see the model as supporting a low-risk intervention. The paper does not settle this disagreement; it maps it. 5.2 Relationship to VoI methodology The framework described here is closely related to value-of-information (VoI) analysis, as it determines which uncertainty most influences the decision outcome [ 18 ]. The expected value of perfect information (EVPI) about the surrogate link is conceptually the current value of the net monetary benefit if the link could be confirmed; if the link is absent or negligible, this benefit drops to zero. Because the decision outcome mainly depends on whether α is practically positive rather than its exact size, this approach aligns well with expected value of sample information (EVSI): what would a case-level linkage study with sufficient causal identification add to the posterior belief about α? A full EVSI analysis is beyond the scope of this paper because it would require a prior distribution over the structural question itself, not just over the parameters within the model. Instead, the framework offers a qualitative EVSI argument: a study that prospectively measures diagnostic accuracy as a function of retrieval quality in a clinical RAG deployment would either confirm the surrogate link, thereby supporting the economic case, or refute it, invalidating the model. The cost of that study, compared to the potential healthcare savings if the link holds, determines the decision horizon. 5.3 Limitations First, the surrogate link α is now based on an empirical estimate from the DiReCT retrieval experiment (α = 1.111, ClinicalBERT, PDD-level, MIMIC-IV clinical notes), with the PSA distribution bounded below by the Tao et al. transported CDSS estimate (0.36). However, the DiReCT experiment measures diagnosis-label retrieval accuracy, not clinician diagnostic performance or patient outcomes. Retrieving a document whose primary discharge diagnosis matches the query is not the same as a clinician using that document to make an accurate diagnosis. The truncated normal PSA distribution (bounds [0.36, 2.54]) reflects empirically grounded parameter uncertainty but does not resolve the structural question of whether the retrieval-to-diagnosis pathway functions in an active clinical workflow. Second, the harm branch is supported by empirical data, as whitening was harmful across all eight synthetic corpus conditions [ 7 ], though the magnitude remains uncertain. The PSA distribution (truncated normal, bounds [− 0.20, 0.0]) represents a range of plausible harm magnitudes but does not account for the potential of catastrophic harm from whitening amplifying noise in severely homogeneous corpora. The threshold analysis (Table 3 , row 2) indicates that harm would need to reach − 0.449 to change the base-case outcome, offering a considerable margin but not guaranteeing complete certainty. Third, the model does not include a hallucination-offset branch. If whitening causes a higher-ranked but incorrect document to surface more confidently, the generation stage may produce more authoritative yet wrong clinical information. This mechanism is not modeled and would operate in the same way as the harm branch, potentially amplifying it. This is the most significant unresolved bias in the model and is highlighted in Table 5 , row 3. Fourth, clinician adoption is regarded as a fixed parameter rather than as a function of retrieval quality and is based on pre-LLM CDSS literature. Adoption in fully automated RAG pipelines (where the clinician does not explicitly interact with retrieved documents) may differ structurally. The threshold analysis (Table 3 , row 3) confirms that adoption uncertainty is not restrictive at typical deployment scales, but the mechanism might not be well-specified. Fifth, the model parameters for deployment prevalence (p_degraded, p_heterogeneous) are derived from a controlled experiment with 11 model configurations and three corpora. These weights might not reflect the distribution of models and corpora in real-world clinical RAG deployments, which are not audited in the published literature. Sixth, the QALY loss per ADE (0.06) is entirely based on assumed severity weights (5% fatal, 15% serious, 30% moderate, 50% minor) applied to IOM mortality estimates [ 16 ]. No empirical QALY estimates derived from diagnostic error patient populations are used. This means the QALY results reported in Tables 2 and S1 are based on constructed inputs and should be interpreted accordingly. The incremental cost results (ic_BvA) depend solely on monetary costs and are unaffected by this limitation; the QALY and ICER results fully reflect it. Seventh, the evidence supporting effectiveness is self-referential: the ΔMRR values, degradation prevalence (p_degraded), and corpus heterogeneity proportions (p_heterogeneous) are all derived from companion papers by the same author [ 6 , 7 ], neither of which has undergone peer review at the time of submission. This means the primary economic inputs rely on unpublished empirical work. This presents a significant limitation for a journal expecting parameter inputs to have independent empirical support. The model parameters should be considered provisional until the companion papers are independently reviewed, and the economic conclusions should be reevaluated once peer-reviewed estimates become available. 5.4 Evidence priorities Table 5 ranks structural assumptions by decision relevance. The most important evidence need is a case-level linkage study that measures both retrieval quality (MRR@10 or equivalent) and diagnostic accuracy outcomes within the same clinical encounters, during the deployment of a RAG system. The study can be prospective or retrospective, as long as retrieval exposure, timing, and diagnostic outcomes are measured with enough causal clarity. A prospective before-and-after or stepped-wedge design would provide the clearest identification, but a well-designed retrospective cohort with logged retrieval data and linked outcome records could also significantly reduce the main structural uncertainty. The second-priority study is an audit of deployed clinical RAG systems to assess the prevalence of degraded embedding geometry and corpus heterogeneity. This would replace the constructed parameters p_degraded and p_heterogeneous with empirical population-level estimates, enhancing the transportability of the model’s population-level conclusions. The third priority is to assess the extent of whitening harm on structurally uniform corpora in live deployments. The current harm estimate (ΔMRR = − 0.05) is based on a synthetic corpus that might be more homogeneous than typical real-world collections. 6. Conclusion This early economic framework indicates that corpus-only ZCA whitening is financially viable across the modeled cost structure, provided the surrogate chain from retrieval improvement to diagnostic accuracy remains valid. The minimum query volume needed for cost-effectiveness ranges from 6 to 43 queries per year across the plausible α spectrum (from the base case to stress tests), thresholds that are low compared to the scale of typical institutional deployments. The harm component (modeled at ΔMRR = − 0.05 for homogeneous corpora) does not significantly alter this conclusion, given the low technology costs relative to healthcare savings. The paper’s main contribution is not to demonstrate cost-effectiveness but to map the structure of decision uncertainty. The DiReCT experiment strengthens α’s empirical foundation—shifting it from an estimated value borrowed from a CDSS to a direct measurement in a clinical-note retrieval task with diagnosis-label relevance. However, it does not determine whether retrieval improvements lead to better clinician diagnostic performance. Parameter sensitivity analysis remains inconclusive within the model’s structure. Structural uncertainty is decisive: the entire economic argument depends on whether a ΔMRR improvement in a clinical RAG pipeline results in better clinician diagnostic accuracy in practice. Addressing this requires a case-level linkage study with proper causal identification. This framework outlines what such a study should measure and how its findings would influence the economic conclusion. Declarations Author contributions Yngve Mikkelsen: Conceptualisation, Methodology, Software, Formal analysis, Investigation, writing, original draft, review & editing. Competing interests The author is the Chief Medical Officer at AlgiPharma AS. AlgiPharma has no products or commercial interests related to embedding models, RAG, or clinical informatics systems. No other competing interests are declared. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Data and code availability All model code is available at: https://github.com/yngvemikkelsen/clinical-rag-he-model. The underlying embedding benchmark data and clinical-embedding-fix library are available at the companion repositories cited as references [6] and [15]. Generative AI disclosure Claude (Anthropic) was used as a coding assistant and for initial manuscript drafting. Grammarly was used for grammar checking. All content was critically reviewed and revised by the author, who takes full responsibility for scientific content and conclusions. References Lewis P, Perez E, Piktus A, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems. 2020;33:9459–74. Gao Y, Xiong Y, Gao X. Retrieval-augmented generation for large language models: A survey. arXiv. 2023. doi: 10.48550/arXiv.2312.10997. Ethayarajh K. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. Proceedings of EMNLP-IJCNLP. 2019:55–65. Lee J, Yoon W, Kim S. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36(4):1234–40. Alsentzer E, Murphy J, Boag W. Publicly available clinical BERT embeddings. Proceedings of the 2nd Clinical NLP Workshop. 2019:72–8. Mikkelsen Y. Context matters more than model choice: A multi-corpus benchmark of embedding models for clinical retrieval-augmented generation. JMIR Preprints. 2026. doi: 10.2196/preprints.94241. Mikkelsen Y. Context or tuning? Layer-level analysis of embedding degradation in clinical document retrieval. doi. Slawomirski L. The economics of diagnostic safety. Paris: OECD Publishing; 2025. Newman-Toker DE, McDonald KM, Meltzer DO. How much diagnostic safety can we afford, and how should we decide? A health economics perspective. BMJ Quality & Safety. 2013;22(Suppl 2):ii11. Lewkowicz D, Wohlbrandt A, Boettinger E. Economic impact of clinical decision support interventions based on electronic health records. BMC Health Serv Res. 2020;20(1):871. National Institute for Health and Care Excellence. Evidence standards framework for digital health technologies. London: NICE; 2019. Kolasa K, Kozinski G. How to Value Digital Health Interventions? A Systematic Literature Review. Int J Environ Res Public Health. 2020;17(6). Norwegian Medicines Agency. Guidelines for the submission of documentation for single technology assessment. 2023. Ciani O, Grigore B, Taylor RS. Development of a framework and decision tool for the evaluation of health technologies based on surrogate endpoint evidence. Health Econ. 2022;31 Suppl 1(Suppl 1):44–72. Mikkelsen Y. Clinical-embedding-fix: post-hoc corrections for degenerate clinical text embeddings in RAG pipelines. SoftwareX. 2026 [submitted]. doi. Institute of Medicine Committee on Quality of Health Care in A. In: Kohn LT, Corrigan JM, Donaldson MS, editors. To Err is Human: Building a Safer Health System. Washington (DC): National Academies Press (US) Copyright 2000 by the National Academy of Sciences. All rights reserved.; 2000. Tao L, Zhang C, Zeng L, Zhu S, Li N, Li W, et al. Accuracy and Effects of Clinical Decision Support Systems Integrated With BMJ Best Practice-Aided Diagnosis: Interrupted Time Series Study. JMIR Med Inform. 2020;8(1):e16912. Fenwick E, Steuten L, Knies S, Ghabri S, Basu A, Murray JF, et al. Value of Information Analysis for Research Decisions-An Introduction: Report 1 of the ISPOR Value of Information Analysis Emerging Good Practices Task Force. Value Health. 2020;23(2):139–50. Wang L, Yang N, Huang X, et al. Improving text embeddings with large language models. 2024. Husereau D, Drummond M, Augustovski F, de Bekker-Grob E, Briggs AH, Carswell C, et al. Consolidated Health Economic Evaluation Reporting Standards (CHEERS) 2022 Explanation and Elaboration: A Report of the ISPOR CHEERS II Good Practices Task Force. Value Health. 2022;25(1):10–31. Additional Declarations The authors declare no competing interests. Supplementary Files Supplement.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9237671","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":612878383,"identity":"f6cd6a66-7699-43f4-be00-5c2dcbaed430","order_by":0,"name":"Yngve Mikkelsen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHElEQVRIie3RMWuDQBTA8SdCu7zE9aQl9iNccG3JVzE4uJipUBw6CAVdhK4HGfoVhELnk4NMNxdBB6XQOdkslBINhEI5ScYO958O5Xe+hwA63T8OCYLB+8MMyO9T8xSBgbhnEziSZXyK0Cp9J1t4u7YzwYsuqoOXddrALgHHitGlKlLLB5tBhVeTxBMoP1d5LanBJMwZR9dTkTL0XezJzEIqjESschKCiREYOaDLx8j3gVjb4utHBA4LGhMpLMZJUHzAYbAM+CQWHpQeHb6yHIhqMLsMjTajFdrZhgrciPmwS8Ek8Zm4uFetPy2DhndRtSDSb9vuUTjOOm2bXXJ795w+vRIFueH92PD3sn4FMvojnfiyUb/R6XQ63bE9Ti1njHT9LXQAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0003-1543-3805","institution":"Saïd Business School, University of Oxford, Oxford, United Kingdom","correspondingAuthor":true,"prefix":"","firstName":"Yngve","middleName":"","lastName":"Mikkelsen","suffix":""}],"badges":[],"createdAt":"2026-03-26 19:46:16","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9237671/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9237671/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105707212,"identity":"9b4f7b8a-f16a-493f-98f4-cb92175c73b9","added_by":"auto","created_at":"2026-03-30 07:13:45","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":337954,"visible":true,"origin":"","legend":"\u003cp\u003ePSA and scenario results. (A) Cost-effectiveness plane: B versus A (whitening versus no correction), 2,000 PSA iterations; diamond indicates the base case, dashed line represents €60k/QALY WTP. (B) Cost-effectiveness acceptability curves. (C) Tornado diagram showing one-way sensitivity analysis on incremental cost. (D) Net monetary benefit by scenario.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9237671/v1/2ecc3ea5647cc10c72428a24.png"},{"id":105728981,"identity":"12486426-62f5-4809-8d09-b0d25421fa30","added_by":"auto","created_at":"2026-03-30 11:13:09","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1253728,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9237671/v1/616cc6f7-3d2e-4160-8cb4-510ac19812d8.pdf"},{"id":105707213,"identity":"72b8f89c-17a2-4e2a-9e04-2d4e6cf84586","added_by":"auto","created_at":"2026-03-30 07:13:46","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":32863,"visible":true,"origin":"","legend":"","description":"","filename":"Supplement.docx","url":"https://assets-eu.researchsquare.com/files/rs-9237671/v1/6b8294be177bd02097ea9ce6.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eEarly economic evaluation of retrieval-layer correction in clinical RAG: a decision-uncertainty framework\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"1. Why is retrieval infrastructure a health economics problem","content":"\u003cp\u003eRetrieval-augmented generation (RAG) has become a key approach for grounding large language model (LLM) outputs in verified clinical evidence [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. In a clinical RAG pipeline, an embedding model converts queries and clinical documents into vector representations; cosine similarity then finds the most relevant documents to provide as context to the LLM. Retrieval quality depends on embedding geometry: when embeddings cluster into a narrow cone (anisotropy), cosine similarity loses its discrimination ability, causing retrieval to fail regardless of the quality of generation [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eA companion benchmarking study found that context variables together explain 49.0% of the variance in MRR@10 across 294 experimental conditions on three clinical corpora. It also showed that domain-specific models (BioBERT [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], ClinicalBERT [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]) exhibited nearly degenerate geometry despite biomedical pretraining [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. A layer-level analysis across 1,400 conditions revealed a U-shaped performance curve, with mid-layer collapse and variable recovery at the final layer [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Corpus-only ZCA whitening, a linear transformation applied to document embeddings without retraining the model, improved MRR@10 by +\u0026thinsp;0.16 to +\u0026thinsp;0.27 on degraded models across diverse clinical text conditions but reduced performance on structurally uniform corpora.\u003c/p\u003e \u003cp\u003eIf improvements in retrieval lead to better clinical outcomes, these empirically demonstrated gains could have economic value. Diagnostic errors are estimated to occur in 10\u0026ndash;15% of clinical encounters and contribute to 6\u0026ndash;17% of adverse events [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Clinical decision support systems (CDSS) embedded in electronic health records have been shown to decrease unnecessary healthcare utilization in multiple systematic reviews [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. The question of whether RAG retrieval quality significantly influences clinical outcomes and whether investing in infrastructure-layer improvements is justified is therefore an HEOR question, not just a technical one. Although ZCA whitening is a linear algebra operation on a vector database rather than a traditional health technology, its downstream effects on a diagnostic AI pipeline that patients interact with place it within the realm of early technology assessment for digital health interventions. The NICE Evidence Standards Framework for Digital Health Technologies [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] and a systematic review of valuation methods for digital health interventions [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] both include infrastructure-level components whose clinical impact needs prospective validation, which is the approach adopted here.\u003c/p\u003e \u003cp\u003eHowever, conventional cost-effectiveness analysis is premature at this stage of the evidence. No study has empirically connected improvements in retrieval metrics in a clinical RAG pipeline to measurable changes in diagnostic accuracy or patient outcomes. The surrogate chain from MRR@10 to adverse diagnostic events depends on an unmeasured link. This paper does not claim otherwise. Instead, it asks a question that can be answered with current evidence: under what conditions could a low-cost retrieval-layer intervention be economically attractive, and which empirical parameters most influence that conclusion?\u003c/p\u003e"},{"header":"2. Objective","content":"\u003cp\u003eThis paper constructs an exploratory decision model to answer three related questions:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eWhat is the minimum annual query volume (N*) at which ZCA whitening recovers its implementation cost, across a range of plausible surrogate link values?\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eWhat values of the three most uncertain parameters, the surrogate link (α), harm magnitude in homogeneous corpora, and clinician adoption, would make whitening not economically attractive?\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eWhat structural uncertainties are not captured by probabilistic sensitivity analysis, and what case-level linkage study would most efficiently reduce decision uncertainty?\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe fine-tuning comparator (a more resource-intensive model adaptation option) is addressed in the Supplement. The primary comparison is whitening versus no intervention.\u003c/p\u003e"},{"header":"3. Methods","content":"\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Model structure and population\u003c/h2\u003e \u003cp\u003eAn exploratory decision tree was created for a single deployment choice: apply corpus-only whitening or do nothing. The model population consists of clinical RAG deployments that have been checked for embedding degradation and corpus type, as detailed below. Fine-tuning as an active comparator is addressed in the Supplement.\u003c/p\u003e \u003cp\u003eWhitening effectiveness is modeled across two branches of the degraded-model population. In the degraded-and-heterogeneous branch (estimated at 30.2% of all new deployments: 0.45 \u0026times; 0.67), whitening improves retrieval accuracy. In the degraded-and-homogeneous branch (14.9%: 0.45 \u0026times; 0.33), whitening causes a slight reduction in retrieval accuracy (ΔMRR\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.05, base case). This two-branch structure was introduced to address a directional bias in previous versions of this model, which set the harm from the homogeneous corpus to zero. The harm branch draws on empirical data showing that whitening degraded performance on all eight structurally uniform corpus conditions tested [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. In 55% of deployments with non-degraded models, no intervention or benefit is observed across all arms.\u003c/p\u003e \u003cp\u003eThe surrogate link α relates the improvement in MRR@10 to the gain in diagnosis-label retrieval accuracy, which is defined as a top-1 match between the retrieved document\u0026rsquo;s primary discharge diagnosis and the query\u0026rsquo;s PDD. This is different from clinician diagnostic accuracy, as discussed in Limitation 1, and remains a key structural uncertainty in the model. In this version, α is based on an empirical estimate from a DiReCT-based diagnosis-label retrieval experiment (α\u0026thinsp;=\u0026thinsp;1.111, ClinicalBERT, MIMIC-IV clinical notes), with the Tao et al. CDSS estimate (α\u0026thinsp;=\u0026thinsp;0.36) kept as the conservative lower bound. Neither estimate has been validated in a live clinical workflow; this structural issue is discussed in Section \u003cspan refid=\"Sec12\" class=\"InternalRef\"\u003e4.3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eHealthcare system perspective, Norwegian reference case [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Five-year time horizon, 4% annual discount rate. All costs in EUR. Parameter sources and distributions are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Primary outputs: scale-independent results\u003c/h2\u003e \u003cp\u003eAll primary results are expressed per 1,000 annual queries or as minimum query-volume thresholds to separate findings from the highly uncertain deployment-scale parameter N. Absolute costs for illustrative deployment sizes are shown in Supplement Table S1. The reason is that N varies by about five orders of magnitude across plausible deployment scenarios (from a small clinic to a large health system), making absolute estimates useful as rough examples rather than precise forecasts.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Threshold analyses\u003c/h2\u003e \u003cp\u003eThree threshold analyses determine the parameter values where whitening becomes economically unviable. For each threshold, two figures are presented: the threshold value and the difference between the base-case parameter and the threshold. A large difference indicates robustness; a small difference points to a decision-sensitive zone.\u003c/p\u003e \u003cp\u003eThe surrogate link threshold (α*) is the value of α below which whitening does not recoup its implementation cost at a specific N. The harm threshold is the homogeneous-corpus ΔMRR magnitude above which whitening fails. The adoption threshold is the minimum clinician adoption rate needed for a positive net monetary benefit.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Structural uncertainty analysis\u003c/h2\u003e \u003cp\u003eProbabilistic sensitivity analysis (PSA, 10,000 iterations) samples parameter uncertainty within the model\u0026rsquo;s fixed causal structure. It does not address structural uncertainty, which concerns whether the assumed relationships are valid. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e formally describes the structural assumptions, their possible failure modes, the direction of bias if they fail, and the type of study needed to test them. This table is the equivalent, for a RAG infrastructure model, of what a formal validity assessment provides for a traditional surrogate endpoint [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Sensitivity analyses\u003c/h2\u003e \u003cp\u003eOne-way sensitivity analysis (tornado diagram) varied each parameter between its 5th- and 95th-percentile values. PSA used truncated normal distributions for the surrogate link (α, anchored at the DiReCT empirical estimate, bounds [0.36, 2.54]) and the harm parameter (ΔMRR_harm, bounds [\u0026minus;\u0026thinsp;0.20, 0.0]), and beta and gamma distributions for probabilities and costs, respectively. A joint chain sensitivity analysis is not included but is a recognized limitation: the tornado diagram treats parameters independently, whereas the real structural uncertainty lies in the multiplicative chain (p_degraded \u0026times; p_heterogeneous\u0026thinsp;\u0026times;\u0026thinsp;ΔMRR\u0026thinsp;\u0026times;\u0026thinsp;α\u0026thinsp;\u0026times;\u0026thinsp;p_adopt \u0026times; p_ADE \u0026times; cost_ADE). Readers should note that the individually low N* thresholds are an artifact of this chain operating at base-case values across all links; a joint stress test that sets multiple links simultaneously to pessimistic values would yield higher N*. A discount rate of 4% is used, following the Norwegian reference case [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. As a sensitivity, 3.0% and 3.5% (NICE and ISPOR international reference rates) were tested; results are presented in Supplement Table S1. A societal perspective scenario using NPE-derived costs (Norsk pasientskadeerstatning annual reports 2021\u0026ndash;2025, average approved payout NOK 765,481\u0026thinsp;\u0026divide;\u0026thinsp;11.50 NOK/EUR = \u0026euro;66,564) is reported separately; NPE payouts include pain/suffering, lost earnings, and future care and are not equivalent to direct healthcare system costs. All analyses were conducted in Python 3.10; code is available at the companion repository [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameters, base-case values, PSA distributions, and evidence sources.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e Parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBase\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePSA distribution\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSource / note\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003eEffectiveness\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMean ΔMRR@10 \u0026mdash; whitening, heterogeneous corpus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e+\u0026thinsp;0.221\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eN(0.221, 0.035)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMikkelsen 2026 JAMIA Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e: mean of four degraded models [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWhitening ΔMRR \u0026mdash; homogeneous corpus (harm; negative)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026minus;0.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTrunc N(\u0026minus;\u0026thinsp;0.05, 0.04) [\u0026minus;\u0026thinsp;0.20, 0.0]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMikkelsen 2026 JAMIA Suppl. Table S2: harmful on all 8 Synthetic corpus conditions. Magnitude uncertain; \u0026minus;0.05 is conservative [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eP(corpus heterogeneous)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBeta(8, 4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMikkelsen 2026 JAMIA: whitening positive on 16/24 conditions (heterogeneous corpora only) [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eP(model degraded)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBeta(5, 6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e4 of 11 model configurations showed degraded geometry in companion study [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSurrogate link (empirically estimated in diagnosis-label retrieval task; clinical workflow link unvalidated)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSurrogate link α (Δ diagnosis-label retrieval accuracy per unit ΔMRR@10; see Limitation 1 for distinction from clinical diagnostic accuracy)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.111\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTrunc N(1.11, 0.37) [0.36, 2.54]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDiReCT experiment (MIMIC-IV-Ext-DiReCT, NeurIPS 2024): ZCA whitening on ClinicalBERT embeddings, 343 MIMIC-IV clinical notes, PDD-level retrieval accuracy. α\u0026thinsp;=\u0026thinsp;1.111, 95% CI [1.014, 2.541]. PSA lower bound\u0026thinsp;=\u0026thinsp;0.36 (Tao et al. 2020 conservative floor). α\u0026thinsp;\u0026gt;\u0026thinsp;1 is mechanistically expected: whitening produces top-1 rank promotions, where Δacc@1 is binary but ΔMRR is fractional.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClinician adoption rate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBeta(9, 6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eKawamoto et al. BMJ 2005 systematic review of CDSS features and trial outcomes [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Value of 0.60 is a constructed assumption; no single empirical study directly supports this figure for RAG-based diagnostic tools. The threshold analysis confirms adoption uncertainty is not binding at any modelled deployment scale.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eClinical outcomes\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eP(adverse event | diagnostic error)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBeta(6, 44)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6\u0026ndash;17% of adverse events involve diagnostic error (midpoint 12%) [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQALY loss per ADE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGamma(1.2, 0.05)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSeverity-weighted: fatal (5%), serious (15%), moderate (30%), minor (50%). Constructed from IOM mortality data [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHealthcare cost per ADE \u0026mdash; system perspective (\u0026euro;)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026euro;12,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGamma(2.0, 6000)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNorwegian DRG\u0026thinsp;+\u0026thinsp;OECD diagnostic safety report [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Excess hospitalisation 2\u0026ndash;5 days. Range \u0026euro;5k\u0026ndash;30k in PSA.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eADE cost \u0026mdash; societal perspective (NPE) (\u0026euro;)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026euro;66,564\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGamma(56.4, 1180)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNPE annual reports 2021\u0026ndash;2025: NOK 765,481\u0026thinsp;\u0026divide;\u0026thinsp;11.50 NOK/EUR. Societal perspective; includes pain/suffering, lost earnings. Sensitivity scenario only.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTechnology costs\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWhitening: one-time implementation (\u0026euro;)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026euro;800\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGamma(2.0, 400)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1\u0026ndash;2 engineer-days. Marginal compute\u0026thinsp;\u0026asymp;\u0026thinsp;0 (NumPy matrix multiply) [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWhitening: annual maintenance (\u0026euro;)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u0026euro;150\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGamma(1.5, 100)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePeriodic corpus refit as document index evolves\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003e\u003cem\u003eAll monetary values in EUR. N/A\u0026thinsp;=\u0026thinsp;not applicable in PSA (fixed for scenario analysis only).\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. Results","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n \u003ch2\u003e4.1 Scale-independent primary results\u003c/h2\u003e\n \u003cp\u003eTable \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows the main results per 1,000 annual queries, excluding fixed implementation costs. These findings are independent of assumptions about deployment scale. At baseline parameters (\u0026alpha;\u0026thinsp;=\u0026thinsp;1.111, based on the empirical DiReCT estimate), whitening prevents 4.74 adverse diagnostic events (ADEs) annually and saves \u0026euro;253,008 in healthcare costs over 5 years for each 1,000 queries. With the Tao et al. [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] conservative \u0026alpha;\u0026thinsp;=\u0026thinsp;0.36, these numbers decrease to 1.53 ADEs and \u0026euro;81,983, which are the same as in the previous version of this model, now shown as a scenario rather than the baseline. The decrease from the zero-harm estimate reflects active modeling of whitening\u0026rsquo;s impact on retrieval quality in uniform corpora.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003ePrimary results per 1,000 annual queries (5-year horizon, healthcare system perspective). Excludes fixed implementation costs (\u0026euro;800 one-time).\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eMetric\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eZero-harm (na\u0026iuml;ve)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eBase case (\u0026alpha;\u0026thinsp;=\u0026thinsp;1.111, harm\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.05)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eConservative (\u0026alpha;\u0026thinsp;=\u0026thinsp;0.36, Tao et al.)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eADEs averted per year (net)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e5.33\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e4.74\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e1.53\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eHealthcare saving 5\u0026nbsp;year (\u0026euro;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e\u0026euro;284,738\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e\u0026euro;253,008\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u0026euro;81,983\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eQALY gain 5yr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e1.4216\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e1.2646\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e0.4099\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e\n \u003cp\u003e\u003cem\u003eZero-harm column sets homogeneous-corpus harm to zero. Base-case column reflects harm modelling. Conservative column uses Tao et al. transported \u0026alpha;\u0026thinsp;=\u0026thinsp;0.36 as sensitivity scenario.\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003e4.2 Threshold analyses\u003c/h2\u003e\n \u003cp\u003eTable \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows the three main threshold analyses. The main finding is the imbalance between parameter and structural uncertainty.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eThreshold analyses: parameter values at which whitening is no longer cost-effective (system perspective, N\u0026thinsp;=\u0026thinsp;133,000/yr unless stated).\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"5\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eParameter\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eBase case\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eThreshold value\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eMargin\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eInterpretation\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eSurrogate link \u0026alpha;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e1.111\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e0.00002\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e1.111\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eEven a negligible link suffices at N\u0026thinsp;=\u0026thinsp;133k. Decision hinges on whether \u0026alpha;\u0026thinsp;=\u0026thinsp;0 exactly \u0026mdash; a structural, not parametric, question.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eHarm magnitude |\u0026Delta;MRR| in homogeneous corpus\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e0.449\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e0.399\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eHarm would need to be 9\u0026times; larger than base-case to make whitening not cost-effective. Robust.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eClinician adoption p_adopt\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e0.60\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e0.00001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e0.600\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eEffectively zero threshold at N\u0026thinsp;=\u0026thinsp;133k given low tech cost.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eMinimum N* (\u0026alpha;\u0026thinsp;=\u0026thinsp;1.111, harm\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.05)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e133,000/yr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e6 queries/yr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eN/A\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eBreakeven query volume at base-case \u0026alpha;. Lower than v3 (was 18) because empirical \u0026alpha; is larger.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eMinimum N* (\u0026alpha;\u0026thinsp;=\u0026thinsp;0.36, Tao et al.)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e133,000/yr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e18 queries/yr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eN/A\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eBreakeven at Tao et al. conservative \u0026alpha; \u0026mdash; now a scenario, not base case.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eMinimum N* (\u0026alpha;\u0026thinsp;=\u0026thinsp;0.15, stress-test)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e133,000/yr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e43 queries/yr\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eN/A\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eBreakeven under extreme stress-test. Still below any real deployment.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eTable \u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e presents minimum N* across a grid of \u0026alpha; and harm values, enabling readers to apply their own priors about the surrogate link.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eMinimum annual query volume N* for whitening to recover \u0026euro;800 implementation cost, by surrogate link (\u0026alpha;) and homogeneous-corpus harm magnitude.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"5\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u0026alpha;\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eHarm\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.05\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eHarm\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.20\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eHarm\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.45\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eDecision implication\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e0.05 (extreme stress)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e129\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e207\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u0026infin;\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eFails only at catastrophic harm\u0026thinsp;+\u0026thinsp;negligible link\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e0.15 (stress-test)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e43\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u0026infin;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eBelow the scale of most deployed systems even under extreme assumptions\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e0.36 (Tao et al. conservative)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u0026infin;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eOld base case \u0026mdash; now conservative scenario\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e1.111 (DiReCT base case)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u0026infin;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eNew base case: any deployment with \u0026gt;\u0026thinsp;9 queries/yr\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2.54 (DiReCT 97.5th pct)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u0026infin;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eUpper CI bound: effectively all deployments\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003ctfoot\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"5\"\u003e\u003cem\u003e\u0026infin; = harm exceeds benefit regardless of N at this \u0026alpha;. DiReCT 95% CI for \u0026alpha;: [1.014, 2.541]. Tao et al. \u0026alpha;\u0026thinsp;=\u0026thinsp;0.36 is retained as conservative scenario (lower PSA bound).\u003c/em\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tfoot\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n \u003ch2\u003e4.3 Structural uncertainty analysis\u003c/h2\u003e\n \u003cp\u003eTable \u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e details the structural assumptions underlying the model. These are assumptions that PSA cannot test because it assesses parameter uncertainty within a fixed causal framework. Each row lists an assumption, its potential failure mode, the bias direction if it fails, and the type of evidence that would test it.\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eStructural uncertainties not resolved by probabilistic sensitivity analysis.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eStructural assumption\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eHow it could fail\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eBias direction if false\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eStudy type required to test\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e\u0026Delta;MRR improvement in retrieval translates to improved diagnostic accuracy (surrogate link \u0026alpha;\u0026thinsp;\u0026gt;\u0026thinsp;0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eRAG generation may compensate for retrieval quality; clinicians may not act on retrieved context; context relevance may not determine diagnosis; diagnosis-label retrieval (DiReCT) may not translate to clinician diagnostic performance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eOverstates benefit; entire economic case collapses if \u0026alpha;\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eCase-level linkage study: retrieval quality (MRR@10) measured alongside diagnostic accuracy outcomes on the same clinical encounters in a live RAG deployment. May be prospective or retrospective with logged retrieval data. Single highest-priority evidence need.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eHarm in homogeneous branch is proportional to \u0026alpha; (same surrogate multiplier)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eHarm from whitening on uniform corpora may operate through a different mechanism than benefit\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eUnderstates or overstates harm depending on direction\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eExperimental: measure \u0026Delta;ADE rate on deployments identified as using whitening on homogeneous corpora\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eNo hallucination-induced harm from higher-ranked but incorrect documents\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eA document retrieved more confidently by a whitened model may be more convincingly wrong\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eOverstates benefit; could produce net harm\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eProspective: measure hallucination and incorrect-document-surfacing rates pre/post whitening\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eFine-tuning effectiveness (\u0026Delta;MRR\u0026thinsp;=\u0026thinsp;0.35) is transportable to new deployments\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eFine-tuning on available clinical corpora may not generalise to deployment corpus\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eOverstates or understates comparative advantage of fine-tuning (supplement only)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eHead-to-head pre/post study of fine-tuned vs whitened models on deployment corpus\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eClinician adoption is independent of retrieval improvement magnitude\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eHigher retrieval quality may increase adoption non-linearly; adoption may be zero in fully automated pipelines\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eCould under- or overstate benefit depending on adoption curve shape\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eObservational: measure clinician engagement with RAG-surfaced evidence as function of retrieval quality\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eDeployment prevalence weights (p_degraded, p_heterogeneous) reflect real deployment populations\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eModels are not randomly selected; heterogeneous corpora may be over- or under-represented\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eCould over- or understate population-level impact\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eSurvey study: audit of embedding model and corpus choices in deployed clinical RAG systems\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n \u003ch2\u003e4.4 Probabilistic sensitivity analysis\u003c/h2\u003e\n \u003cp\u003eThe PSA (10,000 iterations) classified whitening as dominant (cost-saving and QALY-gaining) in 99.7% of iterations. The PSA distribution for \u0026alpha; is centered on the empirical estimate from DiReCT (Truncated N(1.11, 0.37), bounds [0.36, 2.54]), with the lower bound set at the conservative value from Tao et al. The remaining 0.3% of iterations are non-dominant, representing extreme cases of strong homogeneous-corpus harm and low surrogate link at the PSA boundary. An important caveat is that the PSA distribution for \u0026alpha; intentionally excludes zero. The possibility that \u0026alpha;\u0026thinsp;=\u0026thinsp;0\u0026mdash;indicating that retrieval improvement produces no diagnostic benefit\u0026mdash;is not sampled in PSA because it falls outside the empirically bounded distribution. The 99.7% dominance rate should be viewed as conditional: \u0026ldquo;if \u0026alpha; is substantively positive, whitening dominates with high probability.\u0026rdquo; It does not address the prior probability that \u0026alpha; is positive, which is the main structural question discussed in Section \u003cspan refid=\"Sec12\" class=\"InternalRef\"\u003e4.3\u003c/span\u003e.\u003c/p\u003e\n \u003cp\u003eThe cost-effectiveness acceptability curve (CEAC) showed over a 99% probability of cost-effectiveness above \u0026euro;5,000/QALY, reflecting the same condition. The DiReCT experiment offers empirical support for the lower bound of the \u0026alpha; distribution but does not clarify whether diagnosis-label retrieval in annotated notes improves clinician diagnostic performance in real workflows. That structural question remains unchanged by PSA.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\n \u003ch2\u003e4.5 Societal perspective and NPE scenario\u003c/h2\u003e\n \u003cp\u003eFrom a societal perspective, using NPE-derived costs (EUR 66,564 average from Norsk pasientskadeerstatning annual reports 2021\u0026ndash;2025), the minimum N* drops to 2 annual queries under base-case parameters. The societal scenario strengthens the economic case but is based on a fundamentally different cost concept: compensation payouts, including pain, suffering, and lost earnings, which are not suitable as the primary estimate from a healthcare system perspective. It is reported to demonstrate the upper limit of economic attractiveness if payers account for societal costs.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"5. Discussion","content":"\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e5.1 What the framework shows\u003c/h2\u003e \u003cp\u003eThe main finding is asymmetric: within-model parameter uncertainty is managed at low query volumes under the modeled cost structure (N* = 6\u0026ndash;43 depending on assumptions), while uncertainty about the structure of the surrogate chain remains unresolved by the model. This asymmetry itself offers insight. It shows that if a decision-maker accepts the surrogate assumption at any plausible level, whitening becomes economically viable across the modeled deployment range without further analysis. Conversely, if they do not accept it, no amount of modeling or parameter sensitivity analysis will change that.\u003c/p\u003e \u003cp\u003eThis presents a different framing than asking \u0026ldquo;is whitening cost-effective?\u0026rdquo; The answer to that question depends entirely on a single unverified assumption. The correct framing is: \u0026ldquo;if the surrogate link exists at any plausible magnitude (α\u0026thinsp;\u0026ge;\u0026thinsp;0.05), whitening is cost-effective across a broad range of modeled deployment scales; the real question is whether the link exists.\u0026rdquo;\u003c/p\u003e \u003cp\u003eThe threshold and N* analyses in Tables\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and \u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e clarify this structure and enable readers with different priors about the surrogate to draw their own conclusions. A health economist who believes α might be near zero will see the model as underpowered for adoption; someone who thinks any RAG retrieval improvement provides downstream diagnostic benefits will see the model as supporting a low-risk intervention. The paper does not settle this disagreement; it maps it.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e5.2 Relationship to VoI methodology\u003c/h2\u003e \u003cp\u003eThe framework described here is closely related to value-of-information (VoI) analysis, as it determines which uncertainty most influences the decision outcome [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The expected value of perfect information (EVPI) about the surrogate link is conceptually the current value of the net monetary benefit if the link could be confirmed; if the link is absent or negligible, this benefit drops to zero. Because the decision outcome mainly depends on whether α is practically positive rather than its exact size, this approach aligns well with expected value of sample information (EVSI): what would a case-level linkage study with sufficient causal identification add to the posterior belief about α?\u003c/p\u003e \u003cp\u003eA full EVSI analysis is beyond the scope of this paper because it would require a prior distribution over the structural question itself, not just over the parameters within the model. Instead, the framework offers a qualitative EVSI argument: a study that prospectively measures diagnostic accuracy as a function of retrieval quality in a clinical RAG deployment would either confirm the surrogate link, thereby supporting the economic case, or refute it, invalidating the model. The cost of that study, compared to the potential healthcare savings if the link holds, determines the decision horizon.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e5.3 Limitations\u003c/h2\u003e \u003cp\u003eFirst, the surrogate link α is now based on an empirical estimate from the DiReCT retrieval experiment (α\u0026thinsp;=\u0026thinsp;1.111, ClinicalBERT, PDD-level, MIMIC-IV clinical notes), with the PSA distribution bounded below by the Tao et al. transported CDSS estimate (0.36). However, the DiReCT experiment measures diagnosis-label retrieval accuracy, not clinician diagnostic performance or patient outcomes. Retrieving a document whose primary discharge diagnosis matches the query is not the same as a clinician using that document to make an accurate diagnosis. The truncated normal PSA distribution (bounds [0.36, 2.54]) reflects empirically grounded parameter uncertainty but does not resolve the structural question of whether the retrieval-to-diagnosis pathway functions in an active clinical workflow.\u003c/p\u003e \u003cp\u003eSecond, the harm branch is supported by empirical data, as whitening was harmful across all eight synthetic corpus conditions [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], though the magnitude remains uncertain. The PSA distribution (truncated normal, bounds [\u0026minus;\u0026thinsp;0.20, 0.0]) represents a range of plausible harm magnitudes but does not account for the potential of catastrophic harm from whitening amplifying noise in severely homogeneous corpora. The threshold analysis (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, row 2) indicates that harm would need to reach\u0026thinsp;\u0026minus;\u0026thinsp;0.449 to change the base-case outcome, offering a considerable margin but not guaranteeing complete certainty.\u003c/p\u003e \u003cp\u003eThird, the model does not include a hallucination-offset branch. If whitening causes a higher-ranked but incorrect document to surface more confidently, the generation stage may produce more authoritative yet wrong clinical information. This mechanism is not modeled and would operate in the same way as the harm branch, potentially amplifying it. This is the most significant unresolved bias in the model and is highlighted in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, row 3.\u003c/p\u003e \u003cp\u003eFourth, clinician adoption is regarded as a fixed parameter rather than as a function of retrieval quality and is based on pre-LLM CDSS literature. Adoption in fully automated RAG pipelines (where the clinician does not explicitly interact with retrieved documents) may differ structurally. The threshold analysis (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, row 3) confirms that adoption uncertainty is not restrictive at typical deployment scales, but the mechanism might not be well-specified.\u003c/p\u003e \u003cp\u003eFifth, the model parameters for deployment prevalence (p_degraded, p_heterogeneous) are derived from a controlled experiment with 11 model configurations and three corpora. These weights might not reflect the distribution of models and corpora in real-world clinical RAG deployments, which are not audited in the published literature.\u003c/p\u003e \u003cp\u003eSixth, the QALY loss per ADE (0.06) is entirely based on assumed severity weights (5% fatal, 15% serious, 30% moderate, 50% minor) applied to IOM mortality estimates [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. No empirical QALY estimates derived from diagnostic error patient populations are used. This means the QALY results reported in Tables\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and \u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003eS1\u003c/span\u003e are based on constructed inputs and should be interpreted accordingly. The incremental cost results (ic_BvA) depend solely on monetary costs and are unaffected by this limitation; the QALY and ICER results fully reflect it.\u003c/p\u003e \u003cp\u003eSeventh, the evidence supporting effectiveness is self-referential: the ΔMRR values, degradation prevalence (p_degraded), and corpus heterogeneity proportions (p_heterogeneous) are all derived from companion papers by the same author [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], neither of which has undergone peer review at the time of submission. This means the primary economic inputs rely on unpublished empirical work. This presents a significant limitation for a journal expecting parameter inputs to have independent empirical support. The model parameters should be considered provisional until the companion papers are independently reviewed, and the economic conclusions should be reevaluated once peer-reviewed estimates become available.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e5.4 Evidence priorities\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e ranks structural assumptions by decision relevance. The most important evidence need is a case-level linkage study that measures both retrieval quality (MRR@10 or equivalent) and diagnostic accuracy outcomes within the same clinical encounters, during the deployment of a RAG system. The study can be prospective or retrospective, as long as retrieval exposure, timing, and diagnostic outcomes are measured with enough causal clarity. A prospective before-and-after or stepped-wedge design would provide the clearest identification, but a well-designed retrospective cohort with logged retrieval data and linked outcome records could also significantly reduce the main structural uncertainty.\u003c/p\u003e \u003cp\u003eThe second-priority study is an audit of deployed clinical RAG systems to assess the prevalence of degraded embedding geometry and corpus heterogeneity. This would replace the constructed parameters p_degraded and p_heterogeneous with empirical population-level estimates, enhancing the transportability of the model\u0026rsquo;s population-level conclusions.\u003c/p\u003e \u003cp\u003eThe third priority is to assess the extent of whitening harm on structurally uniform corpora in live deployments. The current harm estimate (ΔMRR\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.05) is based on a synthetic corpus that might be more homogeneous than typical real-world collections.\u003c/p\u003e \u003c/div\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eThis early economic framework indicates that corpus-only ZCA whitening is financially viable across the modeled cost structure, provided the surrogate chain from retrieval improvement to diagnostic accuracy remains valid. The minimum query volume needed for cost-effectiveness ranges from 6 to 43 queries per year across the plausible α spectrum (from the base case to stress tests), thresholds that are low compared to the scale of typical institutional deployments. The harm component (modeled at ΔMRR\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.05 for homogeneous corpora) does not significantly alter this conclusion, given the low technology costs relative to healthcare savings.\u003c/p\u003e \u003cp\u003eThe paper\u0026rsquo;s main contribution is not to demonstrate cost-effectiveness but to map the structure of decision uncertainty. The DiReCT experiment strengthens α\u0026rsquo;s empirical foundation\u0026mdash;shifting it from an estimated value borrowed from a CDSS to a direct measurement in a clinical-note retrieval task with diagnosis-label relevance. However, it does not determine whether retrieval improvements lead to better clinician diagnostic performance. Parameter sensitivity analysis remains inconclusive within the model\u0026rsquo;s structure. Structural uncertainty is decisive: the entire economic argument depends on whether a ΔMRR improvement in a clinical RAG pipeline results in better clinician diagnostic accuracy in practice. Addressing this requires a case-level linkage study with proper causal identification. This framework outlines what such a study should measure and how its findings would influence the economic conclusion.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor contributions\u003c/h2\u003e\n\u003cp\u003eYngve Mikkelsen: Conceptualisation, Methodology, Software, Formal analysis, Investigation, writing, original draft, review \u0026amp; editing.\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe author is the Chief Medical Officer at AlgiPharma AS. AlgiPharma has no products or commercial interests related to embedding models, RAG, or clinical informatics systems. No other competing interests are declared.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003ch2\u003eData and code availability\u003c/h2\u003e\n\u003cp\u003eAll model code is available at: https://github.com/yngvemikkelsen/clinical-rag-he-model. The underlying embedding benchmark data and clinical-embedding-fix library are available at the companion repositories cited as references [6] and [15].\u003c/p\u003e\n\u003ch2\u003eGenerative AI disclosure\u003c/h2\u003e\n\u003cp\u003eClaude (Anthropic) was used as a coding assistant and for initial manuscript drafting. Grammarly was used for grammar checking. All content was critically reviewed and revised by the author, who takes full responsibility for scientific content and conclusions.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eLewis P, Perez E, Piktus A, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems. 2020;33:9459\u0026ndash;74.\u003c/li\u003e\n \u003cli\u003eGao Y, Xiong Y, Gao X. Retrieval-augmented generation for large language models: A survey. arXiv. 2023. doi: 10.48550/arXiv.2312.10997.\u003c/li\u003e\n \u003cli\u003eEthayarajh K. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. Proceedings of EMNLP-IJCNLP. 2019:55\u0026ndash;65.\u003c/li\u003e\n \u003cli\u003eLee J, Yoon W, Kim S. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36(4):1234\u0026ndash;40.\u003c/li\u003e\n \u003cli\u003eAlsentzer E, Murphy J, Boag W. Publicly available clinical BERT embeddings. Proceedings of the 2nd Clinical NLP Workshop. 2019:72\u0026ndash;8.\u003c/li\u003e\n \u003cli\u003eMikkelsen Y. Context matters more than model choice: A multi-corpus benchmark of embedding models for clinical retrieval-augmented generation. JMIR Preprints. 2026. doi: 10.2196/preprints.94241.\u003c/li\u003e\n \u003cli\u003eMikkelsen Y. Context or tuning? Layer-level analysis of embedding degradation in clinical document retrieval. doi.\u003c/li\u003e\n \u003cli\u003eSlawomirski L. The economics of diagnostic safety. Paris: OECD Publishing; 2025.\u003c/li\u003e\n \u003cli\u003eNewman-Toker DE, McDonald KM, Meltzer DO. How much diagnostic safety can we afford, and how should we decide? A health economics perspective. BMJ Quality \u0026amp;amp;amp; Safety. 2013;22(Suppl 2):ii11.\u003c/li\u003e\n \u003cli\u003eLewkowicz D, Wohlbrandt A, Boettinger E. Economic impact of clinical decision support interventions based on electronic health records. BMC Health Serv Res. 2020;20(1):871.\u003c/li\u003e\n \u003cli\u003eNational Institute for Health and Care Excellence. Evidence standards framework for digital health technologies. London: NICE; 2019.\u003c/li\u003e\n \u003cli\u003eKolasa K, Kozinski G. How to Value Digital Health Interventions? A Systematic Literature Review. Int J Environ Res Public Health. 2020;17(6).\u003c/li\u003e\n \u003cli\u003eNorwegian Medicines Agency. Guidelines for the submission of documentation for single technology assessment. 2023.\u003c/li\u003e\n \u003cli\u003eCiani O, Grigore B, Taylor RS. Development of a framework and decision tool for the evaluation of health technologies based on surrogate endpoint evidence. Health Econ. 2022;31 Suppl 1(Suppl 1):44\u0026ndash;72.\u003c/li\u003e\n \u003cli\u003eMikkelsen Y. Clinical-embedding-fix: post-hoc corrections for degenerate clinical text embeddings in RAG pipelines. SoftwareX. 2026 [submitted]. doi.\u003c/li\u003e\n \u003cli\u003eInstitute of Medicine Committee on Quality of Health Care in A. In: Kohn LT, Corrigan JM, Donaldson MS, editors. To Err is Human: Building a Safer Health System. Washington (DC): National Academies Press (US) Copyright 2000 by the National Academy of Sciences. All rights reserved.; 2000.\u003c/li\u003e\n \u003cli\u003eTao L, Zhang C, Zeng L, Zhu S, Li N, Li W, et al. Accuracy and Effects of Clinical Decision Support Systems Integrated With BMJ Best Practice-Aided Diagnosis: Interrupted Time Series Study. JMIR Med Inform. 2020;8(1):e16912.\u003c/li\u003e\n \u003cli\u003eFenwick E, Steuten L, Knies S, Ghabri S, Basu A, Murray JF, et al. Value of Information Analysis for Research Decisions-An Introduction: Report 1 of the ISPOR Value of Information Analysis Emerging Good Practices Task Force. Value Health. 2020;23(2):139\u0026ndash;50.\u003c/li\u003e\n \u003cli\u003eWang L, Yang N, Huang X, et al. Improving text embeddings with large language models. 2024.\u003c/li\u003e\n \u003cli\u003eHusereau D, Drummond M, Augustovski F, de Bekker-Grob E, Briggs AH, Carswell C, et al. Consolidated Health Economic Evaluation Reporting Standards (CHEERS) 2022 Explanation and Elaboration: A Report of the ISPOR CHEERS II Good Practices Task Force. Value Health. 2022;25(1):10\u0026ndash;31.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"YM-Medical AS","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"retrieval-augmented generation, clinical informatics, early economic evaluation, value of information, embedding geometry, ZCA whitening, decision uncertainty","lastPublishedDoi":"10.21203/rs.3.rs-9237671/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9237671/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eEmbedding geometry degradation is common in clinical retrieval-augmented generation (RAG) systems and clearly reduces retrieval accuracy. Corpus-only ZCA whitening is a no-retraining correction that boosts retrieval accuracy on diverse clinical text, but its cost-effectiveness depends on whether these improvements lead to better clinical outcomes, a connection that has not yet been empirically confirmed in RAG settings.\u003c/p\u003e\u003ch2\u003eObjective\u003c/h2\u003e \u003cp\u003eTo quantify the conditions where a low-cost retrieval-layer intervention could be economically viable, and to identify the empirical parameters whose measurement would most decrease decision uncertainty.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eAn exploratory decision model with explicit structural gating was developed from a healthcare system perspective (Norwegian reference case, 4% discount rate, 5-year horizon). Whitening effectiveness was modeled across two corpus branches: beneficial on heterogeneous corpora (base-case ΔMRR\u0026thinsp;=\u0026thinsp;+\u0026thinsp;0.221); harmful on homogeneous corpora (ΔMRR\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.05). The surrogate link from retrieval improvement to diagnostic accuracy (α) was empirically estimated from the DiReCT dataset (MIMIC-IV-Ext-DiReCT, NeurIPS 2024): 511 physician-annotated clinical notes from MIMIC-IV, with ZCA whitening applied to ClinicalBERT embeddings and measuring change in primary discharge diagnosis retrieval accuracy. The primary outputs are scale-independent: minimum annual query volume (N*) for cost-effectiveness, and outcomes per 1,000 queries.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eA DiReCT-based retrieval experiment estimated an empirical α\u0026thinsp;=\u0026thinsp;1.111 (95% CI [1.014, 2.541]; ClinicalBERT, PDD-level) in a diagnosis-label retrieval setting, replacing the transported Tao et al. CDSS estimates (0.36) as the base case; the Tao et al. estimate is maintained as the conservative scenario. The experiment used 343 MIMIC-IV clinical notes with sufficient text content (from the full DiReCT dataset of 511 annotated notes). The minimum N* for whitening to cover its \u0026euro;800 implementation cost is 6 annual queries at the base-case parameters and 18 at the conservative α\u0026thinsp;=\u0026thinsp;0.36, thresholds that are low compared to typical institutional deployment scales. Per 1,000 annual queries, whitening prevents 4.74 adverse diagnostic events (base case) or 1.53 (conservative), resulting in \u0026euro;253,008 or \u0026euro;81,983 in healthcare savings over 5 years, respectively. These estimates depend on whether improvements in diagnosis-label retrieval accuracy translate into actual clinician diagnostic performance, a structural assumption the DiReCT experiment does not itself address.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eThis framework shows that whitening appears economically plausible across the modelled cost structure. The DiReCT experiment provides an empirical α estimate in a clinical-note retrieval task with diagnosis-label relevance, substantially above the previously transported CDSS estimate (α\u0026thinsp;=\u0026thinsp;0.36), which is retained as the conservative scenario. The remaining structural uncertainty, whether diagnosis-label retrieval translates to clinician diagnostic performance, would require a case-level linkage study with adequate causal identification to resolve.\u003c/p\u003e","manuscriptTitle":"Early economic evaluation of retrieval-layer correction in clinical RAG: a decision-uncertainty framework","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-30 07:12:35","doi":"10.21203/rs.3.rs-9237671/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4f7a5dae-8cbc-4f74-a8bc-c19183f27cbd","owner":[],"postedDate":"March 30th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":65216796,"name":"Health Economics and Outcomes Research"}],"tags":[],"updatedAt":"2026-03-30T07:12:35+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-30 07:12:35","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9237671","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9237671","identity":"rs-9237671","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00