Mathematical incompleteness of the fragility index | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article Mathematical incompleteness of the fragility index Thomas F Heston This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9315771/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background The Fragility Index (FI) is intended to quantify how many outcome changes would be required to convert a statistically significant two-arm trial result into a non-significant one. For a metric defined on valid 2×2 trial tables, mathematical completeness means that a finite numeric value is obtainable for every valid input in its intended domain. This study evaluated whether the Fragility Index satisfies that property. Methods FI was analyzed as defined: baseline significance required (p < 0.05), one-way movement only, and outcome changes restricted to converting a nonevent to an event in the arm with fewer events while keeping arm size fixed. Mathematical incompleteness was assessed by determining whether valid 2×2 tables exist for which no finite FI can be obtained under these rules. Evidence is provided through formal counterexamples, complete enumeration of all valid nondegenerate 2×2 tables up to total sample size N = 60, and empirical evaluation of published two-arm trials with binary outcomes. Results Valid baseline-significant 2×2 tables exist for which FI is not attainable. A simple counterexample is {3,0,4,11}: baseline Fisher's exact p = 0.04289216, the arm with fewer events is uniquely identified, but that arm has no nonevents available for the required toggle, so no legal FI path exists. Enumeration showed that unattainable cases first appeared at N = 18 and then recurred at every larger sample size through N = 60; by N = 60, 2,390 of 20,774 evaluable baseline-significant tables were unattainable (11.5%). In an empirical dataset of published trials, 2 of 82 baseline-significant evaluable trials (2.4%) were not attainable. Conclusions The Fragility Index is mathematically incomplete. This is a structural property of the FI algorithm, not a data validity issue, confirmed by complete table enumeration and real published trial data. Biostatistics Fragility Index mathematical completeness Fisher's exact test contingency tables clinical trial methodology Introduction The Fragility Index (FI) was introduced as the minimum number of patients whose outcome status would need to change from a nonevent to an event to convert a statistically significant trial result into a non-significant one (Walsh et al., 2014 ). As originally defined, FI was calculated only for trials with statistically significant dichotomous results, and the event changes were applied to the arm with fewer events while preserving arm size. This made FI an intuitively appealing patient-level summary of how easily statistical significance could be lost. FI has since been widely used across clinical specialties because it expresses threshold stability in terms of patient outcomes rather than abstract statistical quantities, and has appeared widely throughout the medical literature, including cardiovascular, pediatric, surgical, and rheumatology specialties (Matics et al., 2017 ; Ruzbarsky, Khormaee & Daluiski, 2019 ; Khan et al., 2019 ; Misra et al., 2025 ). Prior scholarly attention has focused on interpretation and context: whether FI correlates too closely with the p-value to add independent information (Carter, McKie & Storlie, 2016 ), whether the raw count should be normalized by sample size (Ahmed, Fowler & McCredie, 2016 ), and whether the toggling direction should reflect clinical plausibility rather than algorithmic convenience (Walter, Thabane & Briel, 2020 ). These are legitimate concerns, but they are all questions about how FI should be interpreted or reported, not whether it can always be computed. The question addressed here is narrower and more fundamental: does the FI always produce a finite value on the class of valid 2×2 tables to which it is applied? We propose that mathematical completeness — the property that a metric assigns a defined value to every valid input in its domain — is a logical prerequisite for any measure proposed as a universal reporting standard. As past critiques of p-value thresholds have demonstrated, adopting rigid standards without fully understanding their functional limitations can lead to severe distortions in how scientific evidence is evaluated and reported (Wasserstein & Lazar, 2016 ). A metric that fails silently on a non-trivial subset of real trials introduces systematic gaps into the evidence base and is structurally incompatible with uniform adoption, regardless of its interpretive appeal. Whether FI satisfies this requirement has not been formally examined. This is not a mathematical abstraction — it is the same standard that applies to any routine statistical metric. For example, Fisher’s exact test is a foundational, ubiquitously applied method for evaluating 2x2 contingency tables (Choi, Blume & Dupont, 2015 ). If Fisher's exact test simply returned no value for an undetermined subset of valid tables, it would never be considered a reliable statistical tool. The Fragility Index should be held to this same standard. The practical harm is not simply that some trials get no FI value. The deeper problem is that exactly when FI fails is not clearly defined. This means any summary of FI across the literature is computed on an unknown subset of the evidence base, with trials excluded for poorly defined reasons unrelated to their raw outcomes. Much like how the rigid misuse of p-value thresholds has been shown to cause considerable distortion of the scientific process, adopting an incomplete reporting standard for FI could introduce its own systematic distortion into the evidence base (Wasserstein & Lazar, 2016 ). The question has practical urgency. Calls for the routine integration of FI into clinical trial reporting have grown substantially, with some authors also recommending that FI be in clinical guideline development (Narayan et al., 2018 ; Tignanelli & Napolitano, 2019 ). Any such mandate implicitly assumes that every valid baseline-significant trial can be assigned a finite fragility value. Whether that assumption holds has not been formally tested. Modifications to the FI have been proposed, but have not been fully standardized or recognized (Khan et al., 2020 ; Baer et al., 2021 ; Lin & Chu, 2022 ). Therefore, for this analysis, the mathematical completeness of the original published definition of the Fragility Index was evaluated (Walsh et al., 2014 ). Methods Definition of the Fragility Index For a valid 2×2 table {a, b, c, d} where a and c are events and b and d are nonevents, the FI was defined as follows: 1. The baseline table must be statistically significant by a two-sided Fisher's exact test, with p < 0.05. 2. Identify the arm with fewer events. 3. Convert one nonevent to an event in that same arm, preserving the arm total. 4. Recalculate the two-sided Fisher exact p value after each such change. 5. The FI is the minimum number of allowable changes required to make the result non-significant (p > = 0.05). 6. If no finite sequence of allowable changes exists, FI is not attainable. Tables with tied event counts were excluded from the primary analysis because the definition specifies the arm with fewer events and does not provide a tie-breaking rule. Contingency tables are referred to by standard nomenclature {a, b, c, d} where a = events in arm A, b = nonevents in arm A, c = events in arm B, and d = nonevents in arm B. Note that the FI definition does not clearly state what to do when outcome events are tied (a = c). This is the original and established definition of FI (Walsh et al., 2014 ). Criterion for Mathematical Incompleteness A metric is mathematically incomplete on a stated domain if there exists at least one valid input in that domain for which no finite output can be obtained under the metric's own rules (Rosen, 2019 ). Therefore, the existence of a single valid baseline-significant 2×2 table with unattainable FI is sufficient to establish the mathematical incompleteness of the FI. All analyses used the original definition and arm-selection rule (Walsh et al., 2014 ). The FI was calculated using the Fragility Metrics Toolkit ( http://fragilitymetrics.org/calculate_original.php ), and, for external reproducibility checks, counterexamples were also entered into the ClinCalc Fragility Index Calculator ( https://clincalc.com/Stats/FragilityIndex.aspx ), which also applies the Walsh-style iterative Fisher exact test. Formal Counterexample Analysis Constructed tables were examined to determine whether valid baseline-significant examples exist for which the FI procedure cannot produce a finite value. For each counterexample, the baseline Fisher's exact p-value was computed, and every allowable step was evaluated. Complete Enumeration To determine whether non-attainability is isolated or recurrent, all valid nondegenerate 2×2 tables were enumerated for total sample sizes N = 2 through N = 60. For each N, every integer-valued table {a,b,c,d} satisfying a + b + c + d = N was generated. A table was retained for analysis if it met all four of the following criteria: each arm contained at least one subject (a + b > 0 and c + d > 0); each outcome column was present (a + c > 0 and b + d > 0); the baseline two-sided Fisher exact p-value was less than 0.05; and event counts were not tied (a ≠ c). FI was then evaluated on all retained tables. For each N, the following were recorded: total number of valid tables, number of baseline-significant tables, number with attainable FI, number with unattainable FI, and number excluded because event counts were tied. This enumeration was not intended as a prevalence estimate for published trials. Its purpose was structural: to determine whether unattainability recurs across the valid table space, and to characterize the mechanisms by which it arises. The R code to produce this enumeration, as well as the output, has been deposited in the Zenodo repository (DOI: 10.5281/zenodo.18916363 ) Empirical Dataset An empirical convenience dataset of published two-arm trials with binary outcomes was analyzed under the same definition. Only baseline-significant, non-tied tables were counted as evaluable FI cases. Statistical analysis was performed using IBM SPSS Statistics, version 31.0. Claude (Anthropic) was used for language editing, formatting assistance, and code verification; the author reviewed, verified, and is fully responsible for all content. Results Mathematical Proof Theorem The Fragility Index is mathematically incomplete. Proof A metric is mathematically incomplete on its domain if there exists at least one valid input for which the algorithm produces no finite value. Consider the 2×2 table {3, 0, 4, 11}. The baseline two-sided Fisher's exact test yields p = 0.0429, placing it within the FI domain. Arm A has fewer events than Arm B (3 < 4), so all toggles must occur in Arm A. However, Arm A has zero nonevents (b = 0), so no legal move exists. The algorithm cannot begin, and no finite FI value can be obtained. This valid, baseline-significant table demonstrates that the Fragility Index is mathematically incomplete. A second counterexample, {9,1,10,90}, demonstrates that non-attainability extends beyond zero-nonevent cases. The two-sided Fisher's exact p-value is < 0.0001, placing it firmly within the FI domain. Arm A has fewer events (9 < 10), so all changes must occur in arm A. Only one legal change is possible, because arm A has a single nonevent. After that change, yielding {10,0,10,90}, the result remains highly significant (p < 0.0001). No further changes are possible. FI is not attainable, and therefore mathematically incomplete. A third counterexample, {9,35,8,8}, isolates the algorithm's arm-selection behavior without relying on cells with value 0 or 1. The two-sided Fisher's exact p-value is 0.0487, placing it within the FI domain. Arm A has 9 events and 35 nonevents, and arm B has 8 events and 8 nonevents. Arm B has fewer events in absolute terms (8 < 9), so all FI changes must occur in arm B. After each allowable toggle, arm B's event rate increases from 50% toward 100%, widening the disparity between the arms. Each successive toggle moves the two-sided Fisher exact p-value further from the non-significance boundary rather than toward it. All 8 available nonevents in arm B are exhausted, and the p-value does not reach 0.05. FI is not attainable, again demonstrating mathematical incompleteness. Enumeration Results Complete enumeration of all valid nondegenerate 2×2 tables up to total sample size N = 60 showed that unattainable FI cases are recurrent rather than isolated. No unattainable cases were found for N < 18. The first unattainable cases appeared at N = 18, where 2 of 316 evaluable baseline-significant tables were unattainable. Thereafter, unattainable cases recurred at every larger sample size examined; representative results are shown in Table 1 . Among evaluable tables (attainable plus unattainable, excluding ties), FI was not attainable in 28,982 out of 296,192 tables (9.8%). The unattainability rate peaked at 11.5% at N = 60. Representative results are shown in Table 1 ; complete enumeration results for all N = 2–60 are available in the Zenodo repository. Table 1 Selected results from complete enumeration by total sample size. N Total significant FI attainable FI unattainable Ties excluded Unattainability rate 18 324 314 2 8 0.6% 20 460 442 8 10 1.8% 25 1060 998 36 26 3.5% 30 2010 1850 112 48 5.7% 40 5404 4856 432 116 8.2% 50 11578 10220 1142 216 10.1% 60 21112 18384 2390 338 11.5% Legend: For each total sample size shown, the table reports the number of baseline-significant nondegenerate 2×2 tables, the number with attainable and unattainable Fragility Index (FI), the number excluded because event counts were tied, and the resulting unattainability rate calculated as unattainable/(attainable + unattainable). To determine whether non-attainability is an artifact of zero-cell tables — which Walsh's single-arm, single-direction algorithm cannot escape by construction — the minimum cell value was recorded for each unattainable table. If zero-cell tables were the sole mechanism, all unattainable cases would have a minimum cell value of zero. In the enumerated tables (N = 2–60), the minimum cell value ranged from 0 to 8, as shown in Table 2 . Table 2 Minimum cell count of tables with an unattainable FI. Minimum Cell Value Count Percent 0 12,278 42.36% 1 7178 24.77% 2 4284 14.78% 3 2602 8.98% 4 1492 5.15% 5 748 2.58% 6 286 0.99% 7 102 0.35% 8 12 0.04% Total 28,982 100% Legend: This table shows, across all enumerated unattainable cases from N = 2–60, the minimum cell value present in each table and its frequency, to assess whether FI non-attainability is limited to zero-cell tables or also occurs when all cells are nonzero. The minimum cell value of zero accounted for 42.4% of unattainable cases, consistent with the expected dead-end in which the toggling arm is exhausted before significance is lost. The remaining 57.6% of unattainable cases had no zero cell. In these tables, the toggling arm contained at least one nonevent at baseline, yet FI remained unattainable because successive toggles moved the p-value away from the significance boundary rather than toward it, exhausting the arm before the threshold was crossed. Non-attainability is therefore not reducible to a zero-cell artifact: the majority of cases arise from the directional behavior of the Fisher exact p-value under the Walsh toggling path, and would not be resolved by continuity corrections or exclusion of zero-cell tables. Empirical Dataset In the empirical dataset, of 143 total trials, 82 met the FI evaluation criterion of baseline significance with non-tied event counts. Of these, 80 had attainable FI, and 2 had unattainable FI, yielding an empirical non-attainability rate of 2.4% (2/82). The two unattainable cases were {910, 1262, 1028, 1646} (N = 4,846) and {1594, 638, 282, 58} (N = 2,572), confirming that non-attainability is not restricted to small tables. Mechanisms of Non-Attainability Across the full enumeration through N = 60, a total of 28,982 unattainable table instances were identified (cumulative across all evaluated sample sizes). Two distinct mechanisms accounted for all cases. In 12,278 instances (42%), the arm selected for FI toggling already contained zero nonevents at baseline. No allowable move existed, and the algorithm could not begin. In 16,704 instances (58%), at least one legal toggle was available, but the permitted FI path moved the Fisher exact p-value farther from the non-significance boundary (0.05) rather than closer to it. The available toggle room was exhausted before the threshold was reached. Both mechanisms reflect the same underlying constraint: the FI arm-selection rule targets the arm with fewer absolute events, regardless of relative event rates or available toggle room. Within the full enumerated range (N = 2–60), allocation imbalance showed a modest association with non-attainability (Spearman r = 0.324), whereas total sample size showed only a negligible association (Spearman r = 0.015). Discussion The Fragility Index is mathematically incomplete. The key result is formal and does not depend on prevalence estimation: because at least one valid baseline-significant 2×2 table yields no finite FI, the metric is incomplete on its domain. The theorem establishes that non-attainability is possible; the enumeration shows it is recurrent; and the empirical dataset confirms it occurs in published trials. Together, these three layers of evidence demonstrate that the incompleteness is not a theoretical artifact but a structural property. Although the existence of even a single valid table where FI is not attainable proves the theorem, the complete enumeration further strengthens the theorem by showing that non-attainability is not confined to a single unusual example. Unattainable cases first appear at N = 18 and recur at every larger sample size examined through N = 60. The mechanistic analysis adds a further dimension: in the majority of unattainable cases across the enumerated range, at least one legal toggle was available, but the permitted FI path moved Fisher's exact p-value farther from the non-significance boundary rather than closer to it. The arm-selection rule — targeting fewer absolute events — does not account for relative event rates or available toggle room. The empirical result provides real-world confirmation. In this dataset, 2 of 82 baseline-significant evaluable trials had unattainable FI. That proportion should not be overinterpreted as a population prevalence estimate, as it is a convenience sample, but it confirms that non-attainability is not a purely theoretical concern. A central implication follows. Calls for inclusion in routine statistical reporting rest on the premise that FI is computable for every statistically significant two-arm trial. That assumption is false. Valid baseline-significant trials exist for which the algorithm yields no finite result. The phrase "not attainable" is therefore not a marginal implementation detail; it reflects a genuine failure in analysis of valid inputs to numeric outputs. This incompleteness is also distinct from standard statistical uncertainty. Fisher's exact test remains defined for these tables. The tables themselves are valid. The failure arises because the FI algorithm imposes a constrained discrete path through table space and, for some inputs, that path either does not exist or does not reach the significance boundary regardless of the number of steps taken. The practical stakes of this finding are direct. Proposals for routine integration of FI into clinical trial reporting — including explicit recommendations that FI be included in guideline development — presuppose that every valid trial can receive a fragility value (Tignanelli & Napolitano, 2019 ). A mandatory reporting standard built on an incomplete metric would either silently exclude a subset of trials or require ad hoc handling of non-attainable cases with no standardized solution, introducing inconsistency into the very literature the standard is meant to organize (Moher et al., 2010 ). Active research into updated and alternative fragility metrics reflects recognition of the structural limitations of the original FI definition. The Lin and Chu R package, for example, allows event status modifications in both treatment arms rather than restricting toggling to the arm with fewer events, avoiding one source of non-attainability (Lin & Chu, 2022 ). The Global Fragility Index goes further, permitting cell-to-cell reallocation in any direction across the entire table, thereby avoiding independent arm-level constraints altogether (Heston, 2025 ). Whatever direction the field pursues, the core requirement is the same: a fragility metric used as a universal reporting standard must be mathematically complete. The present findings provide the formal basis for that transition. Limitations This study examined only 2×2 tables analyzed with a two-sided Fisher's exact test at alpha = 0.05. This scope is appropriate because the FI was originally defined under exactly these conditions, so the incompleteness proof operates within the metric's own stated domain. Tied-event tables were excluded because the FI definition does not provide a tie-breaking rule; this exclusion is a limitation of the metric itself rather than of the analysis. The empirical dataset was not a systematic review and the 2.4% non-attainability rate should not be interpreted as a population prevalence estimate; however, the two empirical cases occurred in large trials (N = 4,846 and N = 2,572), confirming that non-attainability is not an artifact of small-sample table geometry. The enumeration extended through N = 60 only, but this is sufficient: the theorem is established by counterexample alone, and the enumeration serves to demonstrate recurrence rather than to determine a prevalence ceiling. Conclusion The Fragility Index is mathematically incomplete. Valid baseline-significant 2×2 tables exist for which no finite FI can be computed under the metric's own algorithmic rules. Complete enumeration shows that these cases recur across the valid table space, and empirical analysis confirms their occurrence in published trials. The incompleteness is a formal property of the FI algorithm itself. Declarations Data Availability: All data and enumeration code are archived at Zenodo (https://doi.org/10.5281/zenodo.18916363). Funding: No specific funding received for this methodological work Conflicts of Interest: No competing interests declared Ethics Statement: This study includes a secondary analysis of aggregate statistical results reported in previously published medical literature. No individual-level data, identifiable private information, or identifiable biospecimens were obtained, accessed, or analyzed. Under 45 CFR 46, this project does not constitute human subjects research, and institutional review board oversight was not required. References Ahmed W, Fowler RA, McCredie VA (2016) Does sample size matter when interpreting the fragility index? Crit Care Med 44:e1142–e1143. 10.1097/CCM.0000000000001976 Baer BR, Gaudino M, Charlson M, Fremes SE, Wells MT (2021) Fragility indices for only sufficiently likely modifications. Proc Natl Acad Sci USA 118:e2105254118. 10.1073/pnas.2105254118 Carter RE, McKie PM, Storlie CB (2016) The Fragility Index: a P-value in sheep’s clothing? European Heart Journal:ehw495. 10.1093/eurheartj/ehw495 Choi L, Blume JD, Dupont WD (2015) Elucidating the Foundations of Statistical Inference with 2 x 2 Tables. PLoS ONE 10:e0121263. 10.1371/journal.pone.0121263 Heston TF (2025) The Global Fragility Index: A Path-Independent Measure of Statistical Fragility. SSRN Electron J 5709162. 10.2139/ssrn.5709162 Khan MS, Fonarow GC, Friede T, Lateef N, Khan SU, Anker SD, Harrell FE, Butler J (2020) Application of the reverse fragility index to statistically nonsignificant randomized clinical trial results. JAMA Netw open 3:e2012469. 10.1001/jamanetworkopen.2020.12469 Khan MS, Ochani RK, Shaikh A, Usman MS, Yamani N, Khan SU, Murad MH, Mandrola J, Doukky R, Krasuski RA (2019) Fragility Index in Cardiovascular Randomized Controlled Trials. Circ Cardiovasc Qual Outcomes 12:e005755. 10.1161/CIRCOUTCOMES.119.005755 Lin L, Chu H (2022) Assessing and visualizing fragility of clinical results with binary outcomes in R using the fragility package. PLoS ONE 17:e0268754. 10.1371/journal.pone.0268754 Matics TJ, Khan N, Jani P, Kane JM (2017) The Fragility Index in a Cohort of Pediatric Randomized Controlled Trials. J Clin Med 6:79. 10.3390/jcm6080079 Misra DP, Mukhtyar CB, Chandwar K, Putman M, Walsh M (2025) The fragility of randomized controlled trials in large vessel vasculitis. Autoimmun rev 24:103917. 10.1016/j.autrev.2025.103917 Moher D, Hopewell S, Schulz KF, Montori V, Gøtzsche PC, Devereaux PJ, Elbourne D, Egger M, Altman DG (2010) CONSORT 2010 explanation and elaboration: updated guidelines for reporting parallel group randomised trials. BMJ (Clinical Res Ed) 340:c869. 10.1136/bmj.c869 Narayan VM, Gandhi S, Chrouser K, Evaniew N, Dahm P (2018) The fragility of statistically significant findings from randomised controlled trials in the urological literature. BJU Int 122:160–166. 10.1111/bju.14210 Rosen KH (2019) Discrete mathematics and its applications. McGraw-Hill, New York, NY Ruzbarsky JJ, Khormaee S, Daluiski A (2019) The Fragility Index in Hand Surgery Randomized Controlled Trials. The Journal of Hand Surgery 44:698.e1-698.e7. 10.1016/j.jhsa.2018.10.005 Tignanelli CJ, Napolitano LM (2019) The Fragility Index in Randomized Clinical Trials as a Means of Optimizing Patient Care. JAMA Surg 154:74–79. 10.1001/jamasurg.2018.4318 Walsh M, Srinathan SK, McAuley DF, Mrkobrada M, Levine O, Ribic C, Molnar AO, Dattani ND, Burke A, Guyatt G, Thabane L, Walter SD, Pogue J, Devereaux PJ (2014) The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index. J Clin Epidemiol 67:622–628. 10.1016/j.jclinepi.2013.10.019 Walter SD, Thabane L, Briel M (2020) The fragility of trial results involves more than statistical significance alone. J Clin Epidemiol 124:34–41. 10.1016/j.jclinepi.2020.02.011 Wasserstein R, Lazar N (2016) The ASA Statement on p -Values: Context, Process, and Purpose. Am Stat 70:129–133. 10.1080/00031305.2016.1154108 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9315771","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Method Article","associatedPublications":[],"authors":[{"id":617307220,"identity":"84e4726d-20f9-41c2-9db0-61aa83ad94bb","order_by":0,"name":"Thomas F Heston","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAqklEQVRIiWNgGAWjYDACCQglByZ5iNHBA9JygIHBmHQtiQ1Ea7GXbj78+UPFnfQNNxIYH7xtI8YWmWNpEgfOPMsFamE2nEuUFokcM4aDbYdzN9xOYJPmJU5L/ucPQC3pBrcT2H8TqSWHQQKoJQGohY2ZOC030swkzpw5bDjz/sNmyTnniNDCPiP58YeKisPyfGcOH/zwpowILUiAsYE09aNgFIyCUTAKcAMA+0U5crEgZPIAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-5655-2512","institution":"University of Washington","correspondingAuthor":true,"prefix":"","firstName":"Thomas","middleName":"F","lastName":"Heston","suffix":""}],"badges":[],"createdAt":"2026-04-03 19:18:38","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9315771/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9315771/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106403305,"identity":"9111a090-cbbe-4898-8877-d147291c3ea5","added_by":"auto","created_at":"2026-04-08 09:14:02","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":495937,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9315771/v1/7043f4c0-c3a3-44ea-8fc6-17bc948c3b37.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eMathematical incompleteness of the fragility index\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe Fragility Index (FI) was introduced as the minimum number of patients whose outcome status would need to change from a nonevent to an event to convert a statistically significant trial result into a non-significant one (Walsh et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). As originally defined, FI was calculated only for trials with statistically significant dichotomous results, and the event changes were applied to the arm with fewer events while preserving arm size. This made FI an intuitively appealing patient-level summary of how easily statistical significance could be lost.\u003c/p\u003e \u003cp\u003eFI has since been widely used across clinical specialties because it expresses threshold stability in terms of patient outcomes rather than abstract statistical quantities, and has appeared widely throughout the medical literature, including cardiovascular, pediatric, surgical, and rheumatology specialties (Matics et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Ruzbarsky, Khormaee \u0026amp; Daluiski, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Khan et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Misra et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e \u003cp\u003ePrior scholarly attention has focused on interpretation and context: whether FI correlates too closely with the p-value to add independent information (Carter, McKie \u0026amp; Storlie, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), whether the raw count should be normalized by sample size (Ahmed, Fowler \u0026amp; McCredie, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), and whether the toggling direction should reflect clinical plausibility rather than algorithmic convenience (Walter, Thabane \u0026amp; Briel, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). These are legitimate concerns, but they are all questions about how FI should be interpreted or reported, not whether it can always be computed.\u003c/p\u003e \u003cp\u003eThe question addressed here is narrower and more fundamental: does the FI always produce a finite value on the class of valid 2\u0026times;2 tables to which it is applied? We propose that mathematical completeness \u0026mdash; the property that a metric assigns a defined value to every valid input in its domain \u0026mdash; is a logical prerequisite for any measure proposed as a universal reporting standard. As past critiques of p-value thresholds have demonstrated, adopting rigid standards without fully understanding their functional limitations can lead to severe distortions in how scientific evidence is evaluated and reported (Wasserstein \u0026amp; Lazar, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). A metric that fails silently on a non-trivial subset of real trials introduces systematic gaps into the evidence base and is structurally incompatible with uniform adoption, regardless of its interpretive appeal. Whether FI satisfies this requirement has not been formally examined.\u003c/p\u003e \u003cp\u003eThis is not a mathematical abstraction \u0026mdash; it is the same standard that applies to any routine statistical metric. For example, Fisher\u0026rsquo;s exact test is a foundational, ubiquitously applied method for evaluating 2x2 contingency tables (Choi, Blume \u0026amp; Dupont, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). If Fisher's exact test simply returned no value for an undetermined subset of valid tables, it would never be considered a reliable statistical tool. The Fragility Index should be held to this same standard.\u003c/p\u003e \u003cp\u003eThe practical harm is not simply that some trials get no FI value. The deeper problem is that exactly when FI fails is not clearly defined. This means any summary of FI across the literature is computed on an unknown subset of the evidence base, with trials excluded for poorly defined reasons unrelated to their raw outcomes. Much like how the rigid misuse of p-value thresholds has been shown to cause considerable distortion of the scientific process, adopting an incomplete reporting standard for FI could introduce its own systematic distortion into the evidence base (Wasserstein \u0026amp; Lazar, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2016\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe question has practical urgency. Calls for the routine integration of FI into clinical trial reporting have grown substantially, with some authors also recommending that FI be in clinical guideline development (Narayan et al., \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Tignanelli \u0026amp; Napolitano, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Any such mandate implicitly assumes that every valid baseline-significant trial can be assigned a finite fragility value. Whether that assumption holds has not been formally tested.\u003c/p\u003e \u003cp\u003eModifications to the FI have been proposed, but have not been fully standardized or recognized (Khan et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Baer et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Lin \u0026amp; Chu, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Therefore, for this analysis, the mathematical completeness of the original published definition of the Fragility Index was evaluated (Walsh et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2014\u003c/span\u003e).\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDefinition of the Fragility Index\u003c/h2\u003e \u003cp\u003eFor a valid 2\u0026times;2 table {a, b, c, d} where a and c are events and b and d are nonevents, the FI was defined as follows:\u003c/p\u003e \u003cp\u003e1. The baseline table must be statistically significant by a two-sided Fisher's exact test, with p\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003cp\u003e2. Identify the arm with fewer events.\u003c/p\u003e \u003cp\u003e3. Convert one nonevent to an event in that same arm, preserving the arm total.\u003c/p\u003e \u003cp\u003e4. Recalculate the two-sided Fisher exact p value after each such change.\u003c/p\u003e \u003cp\u003e5. The FI is the minimum number of allowable changes required to make the result non-significant (p\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;0.05).\u003c/p\u003e \u003cp\u003e6. If no finite sequence of allowable changes exists, FI is not attainable.\u003c/p\u003e \u003cp\u003eTables with tied event counts were excluded from the primary analysis because the definition specifies the arm with fewer events and does not provide a tie-breaking rule. Contingency tables are referred to by standard nomenclature {a, b, c, d} where a\u0026thinsp;=\u0026thinsp;events in arm A, b\u0026thinsp;=\u0026thinsp;nonevents in arm A, c\u0026thinsp;=\u0026thinsp;events in arm B, and d\u0026thinsp;=\u0026thinsp;nonevents in arm B. Note that the FI definition does not clearly state what to do when outcome events are tied (a\u0026thinsp;=\u0026thinsp;c). This is the original and established definition of FI (Walsh et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2014\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eCriterion for Mathematical Incompleteness\u003c/h3\u003e\n\u003cp\u003eA metric is mathematically incomplete on a stated domain if there exists at least one valid input in that domain for which no finite output can be obtained under the metric's own rules (Rosen, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Therefore, the existence of a single valid baseline-significant 2\u0026times;2 table with unattainable FI is sufficient to establish the mathematical incompleteness of the FI. All analyses used the original definition and arm-selection rule (Walsh et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). The FI was calculated using the Fragility Metrics Toolkit (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://fragilitymetrics.org/calculate_original.php\u003c/span\u003e\u003cspan address=\"http://fragilitymetrics.org/calculate_original.php\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and, for external reproducibility checks, counterexamples were also entered into the ClinCalc Fragility Index Calculator (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://clincalc.com/Stats/FragilityIndex.aspx\u003c/span\u003e\u003cspan address=\"https://clincalc.com/Stats/FragilityIndex.aspx\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), which also applies the Walsh-style iterative Fisher exact test.\u003c/p\u003e\n\u003ch3\u003eFormal Counterexample Analysis\u003c/h3\u003e\n\u003cp\u003eConstructed tables were examined to determine whether valid baseline-significant examples exist for which the FI procedure cannot produce a finite value. For each counterexample, the baseline Fisher's exact p-value was computed, and every allowable step was evaluated.\u003c/p\u003e\n\u003ch3\u003eComplete Enumeration\u003c/h3\u003e\n\u003cp\u003eTo determine whether non-attainability is isolated or recurrent, all valid nondegenerate 2\u0026times;2 tables were enumerated for total sample sizes N\u0026thinsp;=\u0026thinsp;2 through N\u0026thinsp;=\u0026thinsp;60. For each N, every integer-valued table {a,b,c,d} satisfying a\u0026thinsp;+\u0026thinsp;b + c\u0026thinsp;+\u0026thinsp;d = N was generated. A table was retained for analysis if it met all four of the following criteria: each arm contained at least one subject (a\u0026thinsp;+\u0026thinsp;b \u0026gt; 0 and c\u0026thinsp;+\u0026thinsp;d \u0026gt; 0); each outcome column was present (a\u0026thinsp;+\u0026thinsp;c \u0026gt; 0 and b\u0026thinsp;+\u0026thinsp;d \u0026gt; 0); the baseline two-sided Fisher exact p-value was less than 0.05; and event counts were not tied (a\u0026thinsp;\u0026ne;\u0026thinsp;c). FI was then evaluated on all retained tables. For each N, the following were recorded: total number of valid tables, number of baseline-significant tables, number with attainable FI, number with unattainable FI, and number excluded because event counts were tied.\u003c/p\u003e \u003cp\u003eThis enumeration was not intended as a prevalence estimate for published trials. Its purpose was structural: to determine whether unattainability recurs across the valid table space, and to characterize the mechanisms by which it arises. The R code to produce this enumeration, as well as the output, has been deposited in the Zenodo repository (DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5281/zenodo.18916363\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.18916363\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e)\u003c/p\u003e\n\u003ch3\u003eEmpirical Dataset\u003c/h3\u003e\n\u003cp\u003eAn empirical convenience dataset of published two-arm trials with binary outcomes was analyzed under the same definition. Only baseline-significant, non-tied tables were counted as evaluable FI cases. Statistical analysis was performed using IBM SPSS Statistics, version 31.0.\u003c/p\u003e \u003cp\u003eClaude (Anthropic) was used for language editing, formatting assistance, and code verification; the author reviewed, verified, and is fully responsible for all content.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eMathematical Proof\u003c/h2\u003e \u003cp\u003e \u003cstrong\u003eTheorem\u003c/strong\u003e \u003cp\u003eThe Fragility Index is mathematically incomplete.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eProof\u003c/strong\u003e \u003cp\u003eA metric is mathematically incomplete on its domain if there exists at least one valid input for which the algorithm produces no finite value.\u003c/p\u003e \u003c/p\u003e \u003cp\u003eConsider the 2\u0026times;2 table {3, 0, 4, 11}. The baseline two-sided Fisher's exact test yields p\u0026thinsp;=\u0026thinsp;0.0429, placing it within the FI domain. Arm A has fewer events than Arm B (3\u0026thinsp;\u0026lt;\u0026thinsp;4), so all toggles must occur in Arm A. However, Arm A has zero nonevents (b\u0026thinsp;=\u0026thinsp;0), so no legal move exists. The algorithm cannot begin, and no finite FI value can be obtained. This valid, baseline-significant table demonstrates that the Fragility Index is mathematically incomplete.\u003c/p\u003e \u003cp\u003eA second counterexample, {9,1,10,90}, demonstrates that non-attainability extends beyond zero-nonevent cases. The two-sided Fisher's exact p-value is \u0026lt;\u0026thinsp;0.0001, placing it firmly within the FI domain. Arm A has fewer events (9\u0026thinsp;\u0026lt;\u0026thinsp;10), so all changes must occur in arm A. Only one legal change is possible, because arm A has a single nonevent. After that change, yielding {10,0,10,90}, the result remains highly significant (p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001). No further changes are possible. FI is not attainable, and therefore mathematically incomplete.\u003c/p\u003e \u003cp\u003eA third counterexample, {9,35,8,8}, isolates the algorithm's arm-selection behavior without relying on cells with value 0 or 1. The two-sided Fisher's exact p-value is 0.0487, placing it within the FI domain. Arm A has 9 events and 35 nonevents, and arm B has 8 events and 8 nonevents. Arm B has fewer events in absolute terms (8\u0026thinsp;\u0026lt;\u0026thinsp;9), so all FI changes must occur in arm B. After each allowable toggle, arm B's event rate increases from 50% toward 100%, widening the disparity between the arms. Each successive toggle moves the two-sided Fisher exact p-value further from the non-significance boundary rather than toward it. All 8 available nonevents in arm B are exhausted, and the p-value does not reach 0.05. FI is not attainable, again demonstrating mathematical incompleteness.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eEnumeration Results\u003c/h3\u003e\n\u003cp\u003eComplete enumeration of all valid nondegenerate 2\u0026times;2 tables up to total sample size N\u0026thinsp;=\u0026thinsp;60 showed that unattainable FI cases are recurrent rather than isolated.\u003c/p\u003e \u003cp\u003eNo unattainable cases were found for N\u0026thinsp;\u0026lt;\u0026thinsp;18. The first unattainable cases appeared at N\u0026thinsp;=\u0026thinsp;18, where 2 of 316 evaluable baseline-significant tables were unattainable. Thereafter, unattainable cases recurred at every larger sample size examined; representative results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Among evaluable tables (attainable plus unattainable, excluding ties), FI was not attainable in 28,982 out of 296,192 tables (9.8%). The unattainability rate peaked at 11.5% at N\u0026thinsp;=\u0026thinsp;60. Representative results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e; complete enumeration results for all N\u0026thinsp;=\u0026thinsp;2\u0026ndash;60 are available in the Zenodo repository.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSelected results from complete enumeration by total sample size.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTotal significant\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFI attainable\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFI unattainable\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTies excluded\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUnattainability rate\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e324\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e314\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.6%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e460\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e442\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.8%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1060\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e998\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e3.5%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2010\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1850\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e112\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e5.7%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5404\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4856\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e432\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e116\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e8.2%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e11578\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10220\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1142\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e216\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e10.1%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e21112\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e18384\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2390\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e338\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e11.5%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eLegend: For each total sample size shown, the table reports the number of baseline-significant nondegenerate 2\u0026times;2 tables, the number with attainable and unattainable Fragility Index (FI), the number excluded because event counts were tied, and the resulting unattainability rate calculated as unattainable/(attainable\u0026thinsp;+\u0026thinsp;unattainable).\u003c/p\u003e \u003cp\u003eTo determine whether non-attainability is an artifact of zero-cell tables \u0026mdash; which Walsh's single-arm, single-direction algorithm cannot escape by construction \u0026mdash; the minimum cell value was recorded for each unattainable table. If zero-cell tables were the sole mechanism, all unattainable cases would have a minimum cell value of zero. In the enumerated tables (N\u0026thinsp;=\u0026thinsp;2\u0026ndash;60), the minimum cell value ranged from 0 to 8, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMinimum cell count of tables with an unattainable FI.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMinimum Cell Value\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCount\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePercent\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12,278\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e42.36%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7178\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e24.77%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e4284\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14.78%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2602\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8.98%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1492\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.15%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e748\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2.58%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e286\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.99%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e102\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.35%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.04%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e28,982\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eLegend: This table shows, across all enumerated unattainable cases from N\u0026thinsp;=\u0026thinsp;2\u0026ndash;60, the minimum cell value present in each table and its frequency, to assess whether FI non-attainability is limited to zero-cell tables or also occurs when all cells are nonzero.\u003c/p\u003e \u003cp\u003eThe minimum cell value of zero accounted for 42.4% of unattainable cases, consistent with the expected dead-end in which the toggling arm is exhausted before significance is lost. The remaining 57.6% of unattainable cases had no zero cell. In these tables, the toggling arm contained at least one nonevent at baseline, yet FI remained unattainable because successive toggles moved the p-value away from the significance boundary rather than toward it, exhausting the arm before the threshold was crossed. Non-attainability is therefore not reducible to a zero-cell artifact: the majority of cases arise from the directional behavior of the Fisher exact p-value under the Walsh toggling path, and would not be resolved by continuity corrections or exclusion of zero-cell tables.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eEmpirical Dataset\u003c/h2\u003e \u003cp\u003eIn the empirical dataset, of 143 total trials, 82 met the FI evaluation criterion of baseline significance with non-tied event counts. Of these, 80 had attainable FI, and 2 had unattainable FI, yielding an empirical non-attainability rate of 2.4% (2/82). The two unattainable cases were {910, 1262, 1028, 1646} (N\u0026thinsp;=\u0026thinsp;4,846) and {1594, 638, 282, 58} (N\u0026thinsp;=\u0026thinsp;2,572), confirming that non-attainability is not restricted to small tables.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eMechanisms of Non-Attainability\u003c/h2\u003e \u003cp\u003eAcross the full enumeration through N\u0026thinsp;=\u0026thinsp;60, a total of 28,982 unattainable table instances were identified (cumulative across all evaluated sample sizes). Two distinct mechanisms accounted for all cases.\u003c/p\u003e \u003cp\u003eIn 12,278 instances (42%), the arm selected for FI toggling already contained zero nonevents at baseline. No allowable move existed, and the algorithm could not begin.\u003c/p\u003e \u003cp\u003eIn 16,704 instances (58%), at least one legal toggle was available, but the permitted FI path moved the Fisher exact p-value farther from the non-significance boundary (0.05) rather than closer to it. The available toggle room was exhausted before the threshold was reached.\u003c/p\u003e \u003cp\u003eBoth mechanisms reflect the same underlying constraint: the FI arm-selection rule targets the arm with fewer absolute events, regardless of relative event rates or available toggle room.\u003c/p\u003e \u003cp\u003eWithin the full enumerated range (N\u0026thinsp;=\u0026thinsp;2\u0026ndash;60), allocation imbalance showed a modest association with non-attainability (Spearman r\u0026thinsp;=\u0026thinsp;0.324), whereas total sample size showed only a negligible association (Spearman r\u0026thinsp;=\u0026thinsp;0.015).\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe Fragility Index is mathematically incomplete. The key result is formal and does not depend on prevalence estimation: because at least one valid baseline-significant 2\u0026times;2 table yields no finite FI, the metric is incomplete on its domain.\u003c/p\u003e \u003cp\u003eThe theorem establishes that non-attainability is possible; the enumeration shows it is recurrent; and the empirical dataset confirms it occurs in published trials. Together, these three layers of evidence demonstrate that the incompleteness is not a theoretical artifact but a structural property.\u003c/p\u003e \u003cp\u003eAlthough the existence of even a single valid table where FI is not attainable proves the theorem, the complete enumeration further strengthens the theorem by showing that non-attainability is not confined to a single unusual example. Unattainable cases first appear at N\u0026thinsp;=\u0026thinsp;18 and recur at every larger sample size examined through N\u0026thinsp;=\u0026thinsp;60. The mechanistic analysis adds a further dimension: in the majority of unattainable cases across the enumerated range, at least one legal toggle was available, but the permitted FI path moved Fisher's exact p-value farther from the non-significance boundary rather than closer to it. The arm-selection rule \u0026mdash; targeting fewer absolute events \u0026mdash; does not account for relative event rates or available toggle room.\u003c/p\u003e \u003cp\u003eThe empirical result provides real-world confirmation. In this dataset, 2 of 82 baseline-significant evaluable trials had unattainable FI. That proportion should not be overinterpreted as a population prevalence estimate, as it is a convenience sample, but it confirms that non-attainability is not a purely theoretical concern.\u003c/p\u003e \u003cp\u003eA central implication follows. Calls for inclusion in routine statistical reporting rest on the premise that FI is computable for every statistically significant two-arm trial. That assumption is false. Valid baseline-significant trials exist for which the algorithm yields no finite result. The phrase \"not attainable\" is therefore not a marginal implementation detail; it reflects a genuine failure in analysis of valid inputs to numeric outputs.\u003c/p\u003e \u003cp\u003eThis incompleteness is also distinct from standard statistical uncertainty. Fisher's exact test remains defined for these tables. The tables themselves are valid. The failure arises because the FI algorithm imposes a constrained discrete path through table space and, for some inputs, that path either does not exist or does not reach the significance boundary regardless of the number of steps taken.\u003c/p\u003e \u003cp\u003eThe practical stakes of this finding are direct. Proposals for routine integration of FI into clinical trial reporting \u0026mdash; including explicit recommendations that FI be included in guideline development \u0026mdash; presuppose that every valid trial can receive a fragility value (Tignanelli \u0026amp; Napolitano, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). A mandatory reporting standard built on an incomplete metric would either silently exclude a subset of trials or require ad hoc handling of non-attainable cases with no standardized solution, introducing inconsistency into the very literature the standard is meant to organize (Moher et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2010\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eActive research into updated and alternative fragility metrics reflects recognition of the structural limitations of the original FI definition. The Lin and Chu R package, for example, allows event status modifications in both treatment arms rather than restricting toggling to the arm with fewer events, avoiding one source of non-attainability (Lin \u0026amp; Chu, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). The Global Fragility Index goes further, permitting cell-to-cell reallocation in any direction across the entire table, thereby avoiding independent arm-level constraints altogether (Heston, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Whatever direction the field pursues, the core requirement is the same: a fragility metric used as a universal reporting standard must be mathematically complete. The present findings provide the formal basis for that transition.\u003c/p\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eThis study examined only 2\u0026times;2 tables analyzed with a two-sided Fisher's exact test at alpha\u0026thinsp;=\u0026thinsp;0.05. This scope is appropriate because the FI was originally defined under exactly these conditions, so the incompleteness proof operates within the metric's own stated domain. Tied-event tables were excluded because the FI definition does not provide a tie-breaking rule; this exclusion is a limitation of the metric itself rather than of the analysis. The empirical dataset was not a systematic review and the 2.4% non-attainability rate should not be interpreted as a population prevalence estimate; however, the two empirical cases occurred in large trials (N\u0026thinsp;=\u0026thinsp;4,846 and N\u0026thinsp;=\u0026thinsp;2,572), confirming that non-attainability is not an artifact of small-sample table geometry. The enumeration extended through N\u0026thinsp;=\u0026thinsp;60 only, but this is sufficient: the theorem is established by counterexample alone, and the enumeration serves to demonstrate recurrence rather than to determine a prevalence ceiling.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThe Fragility Index is mathematically incomplete. Valid baseline-significant 2\u0026times;2 tables exist for which no finite FI can be computed under the metric's own algorithmic rules. Complete enumeration shows that these cases recur across the valid table space, and empirical analysis confirms their occurrence in published trials. The incompleteness is a formal property of the FI algorithm itself.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability:\u003c/strong\u003e All data and enumeration code are archived at Zenodo (https://doi.org/10.5281/zenodo.18916363).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e No specific funding received for this methodological work\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest:\u003c/strong\u003e No competing interests declared\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics Statement:\u003c/strong\u003e This study includes a secondary analysis of aggregate statistical results reported in previously published medical literature. No individual-level data, identifiable private information, or identifiable biospecimens were obtained, accessed, or analyzed. Under 45 CFR 46, this project does not constitute human subjects research, and institutional review board oversight was not required.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAhmed W, Fowler RA, McCredie VA (2016) Does sample size matter when interpreting the fragility index? Crit Care Med 44:e1142\u0026ndash;e1143. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/CCM.0000000000001976\u003c/span\u003e\u003cspan address=\"10.1097/CCM.0000000000001976\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaer BR, Gaudino M, Charlson M, Fremes SE, Wells MT (2021) Fragility indices for only sufficiently likely modifications. Proc Natl Acad Sci USA 118:e2105254118. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1073/pnas.2105254118\u003c/span\u003e\u003cspan address=\"10.1073/pnas.2105254118\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarter RE, McKie PM, Storlie CB (2016) The Fragility Index: a P-value in sheep\u0026rsquo;s clothing? European Heart Journal:ehw495. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/eurheartj/ehw495\u003c/span\u003e\u003cspan address=\"10.1093/eurheartj/ehw495\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChoi L, Blume JD, Dupont WD (2015) Elucidating the Foundations of Statistical Inference with 2 x 2 Tables. PLoS ONE 10:e0121263. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0121263\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0121263\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeston TF (2025) The Global Fragility Index: A Path-Independent Measure of Statistical Fragility. SSRN Electron J 5709162. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2139/ssrn.5709162\u003c/span\u003e\u003cspan address=\"10.2139/ssrn.5709162\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan MS, Fonarow GC, Friede T, Lateef N, Khan SU, Anker SD, Harrell FE, Butler J (2020) Application of the reverse fragility index to statistically nonsignificant randomized clinical trial results. JAMA Netw open 3:e2012469. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamanetworkopen.2020.12469\u003c/span\u003e\u003cspan address=\"10.1001/jamanetworkopen.2020.12469\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan MS, Ochani RK, Shaikh A, Usman MS, Yamani N, Khan SU, Murad MH, Mandrola J, Doukky R, Krasuski RA (2019) Fragility Index in Cardiovascular Randomized Controlled Trials. Circ Cardiovasc Qual Outcomes 12:e005755. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/CIRCOUTCOMES.119.005755\u003c/span\u003e\u003cspan address=\"10.1161/CIRCOUTCOMES.119.005755\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin L, Chu H (2022) Assessing and visualizing fragility of clinical results with binary outcomes in R using the fragility package. PLoS ONE 17:e0268754. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0268754\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0268754\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMatics TJ, Khan N, Jani P, Kane JM (2017) The Fragility Index in a Cohort of Pediatric Randomized Controlled Trials. J Clin Med 6:79. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/jcm6080079\u003c/span\u003e\u003cspan address=\"10.3390/jcm6080079\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMisra DP, Mukhtyar CB, Chandwar K, Putman M, Walsh M (2025) The fragility of randomized controlled trials in large vessel vasculitis. Autoimmun rev 24:103917. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.autrev.2025.103917\u003c/span\u003e\u003cspan address=\"10.1016/j.autrev.2025.103917\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoher D, Hopewell S, Schulz KF, Montori V, G\u0026oslash;tzsche PC, Devereaux PJ, Elbourne D, Egger M, Altman DG (2010) CONSORT 2010 explanation and elaboration: updated guidelines for reporting parallel group randomised trials. BMJ (Clinical Res Ed) 340:c869. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/bmj.c869\u003c/span\u003e\u003cspan address=\"10.1136/bmj.c869\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNarayan VM, Gandhi S, Chrouser K, Evaniew N, Dahm P (2018) The fragility of statistically significant findings from randomised controlled trials in the urological literature. BJU Int 122:160\u0026ndash;166. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/bju.14210\u003c/span\u003e\u003cspan address=\"10.1111/bju.14210\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosen KH (2019) Discrete mathematics and its applications. McGraw-Hill, New York, NY\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRuzbarsky JJ, Khormaee S, Daluiski A (2019) The Fragility Index in Hand Surgery Randomized Controlled Trials. \u003cem\u003eThe Journal of Hand Surgery\u003c/em\u003e 44:698.e1-698.e7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jhsa.2018.10.005\u003c/span\u003e\u003cspan address=\"10.1016/j.jhsa.2018.10.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTignanelli CJ, Napolitano LM (2019) The Fragility Index in Randomized Clinical Trials as a Means of Optimizing Patient Care. JAMA Surg 154:74\u0026ndash;79. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamasurg.2018.4318\u003c/span\u003e\u003cspan address=\"10.1001/jamasurg.2018.4318\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWalsh M, Srinathan SK, McAuley DF, Mrkobrada M, Levine O, Ribic C, Molnar AO, Dattani ND, Burke A, Guyatt G, Thabane L, Walter SD, Pogue J, Devereaux PJ (2014) The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index. J Clin Epidemiol 67:622\u0026ndash;628. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jclinepi.2013.10.019\u003c/span\u003e\u003cspan address=\"10.1016/j.jclinepi.2013.10.019\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWalter SD, Thabane L, Briel M (2020) The fragility of trial results involves more than statistical significance alone. J Clin Epidemiol 124:34\u0026ndash;41. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jclinepi.2020.02.011\u003c/span\u003e\u003cspan address=\"10.1016/j.jclinepi.2020.02.011\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWasserstein R, Lazar N (2016) The ASA Statement on p -Values: Context, Process, and Purpose. Am Stat 70:129\u0026ndash;133. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1080/00031305.2016.1154108\u003c/span\u003e\u003cspan address=\"10.1080/00031305.2016.1154108\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"University of Washington","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Fragility Index, mathematical completeness, Fisher's exact test, contingency tables, clinical trial methodology","lastPublishedDoi":"10.21203/rs.3.rs-9315771/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9315771/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe Fragility Index (FI) is intended to quantify how many outcome changes would be required to convert a statistically significant two-arm trial result into a non-significant one. For a metric defined on valid 2\u0026times;2 trial tables, mathematical completeness means that a finite numeric value is obtainable for every valid input in its intended domain. This study evaluated whether the Fragility Index satisfies that property.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eFI was analyzed as defined: baseline significance required (p\u0026thinsp;\u0026lt;\u0026thinsp;0.05), one-way movement only, and outcome changes restricted to converting a nonevent to an event in the arm with fewer events while keeping arm size fixed. Mathematical incompleteness was assessed by determining whether valid 2\u0026times;2 tables exist for which no finite FI can be obtained under these rules. Evidence is provided through formal counterexamples, complete enumeration of all valid nondegenerate 2\u0026times;2 tables up to total sample size N\u0026thinsp;=\u0026thinsp;60, and empirical evaluation of published two-arm trials with binary outcomes.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eValid baseline-significant 2\u0026times;2 tables exist for which FI is not attainable. A simple counterexample is {3,0,4,11}: baseline Fisher's exact p\u0026thinsp;=\u0026thinsp;0.04289216, the arm with fewer events is uniquely identified, but that arm has no nonevents available for the required toggle, so no legal FI path exists. Enumeration showed that unattainable cases first appeared at N\u0026thinsp;=\u0026thinsp;18 and then recurred at every larger sample size through N\u0026thinsp;=\u0026thinsp;60; by N\u0026thinsp;=\u0026thinsp;60, 2,390 of 20,774 evaluable baseline-significant tables were unattainable (11.5%). In an empirical dataset of published trials, 2 of 82 baseline-significant evaluable trials (2.4%) were not attainable.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eThe Fragility Index is mathematically incomplete. This is a structural property of the FI algorithm, not a data validity issue, confirmed by complete table enumeration and real published trial data.\u003c/p\u003e","manuscriptTitle":"Mathematical incompleteness of the fragility index","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-06 06:35:04","doi":"10.21203/rs.3.rs-9315771/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e871e90c-d592-4ce8-a365-dd2c3a1d4d08","owner":[],"postedDate":"April 6th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":65695691,"name":"Biostatistics"}],"tags":[],"updatedAt":"2026-04-06T06:35:04+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-06 06:35:04","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9315771","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9315771","identity":"rs-9315771","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.