Dataset
We use a rich dataset of all the claims made in the contributory regime of the Colombian health system between 2017 and 2019. Our dataset contains the costs of the claims of the individual sex, age, localization, and 255 dummy variables that take the value of one if the individual has the health condition denoted by the dummy and zero otherwise.
In Colombia, health insurers are reimbursed through two types of risk-based payments: a prospective risk adjustment formula, which determines the annual capitation payments they receive ex ante, and a concurrent payment mechanism that compensates for exceptionally high-cost events such as high-cost diseases or cancer treatments [ 11 ]. The present study focuses exclusively on the prospective risk adjustment formula, as this is the primary instrument used to allocate resources across insurers and the mechanism most directly linked to incentives for risk selection. Colombia applies a prospective risk adjustment formula, in which only information available before the coverage period is used to predict future expenditures. This differs from concurrent formulas, used in some other countries, that incorporate diagnostic and expenditure data from the same coverage period. While the prospective design better reflects the information constraints insurers face when setting premiums, it may capture morbidity less precisely than concurrent systems. In this study, we do not construct year-specific panels. Instead, we focus on a single experience period. Specifically, we define \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${y}_{i}$$\end{document} as the average monthly cost incurred by individual \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$i$$\end{document} during 2019. To assess the adequacy of payments for high-cost individuals, we also compute an underpayment metric based on 2019 spending, which allows us to identify undercompensation within that year.
This dataset contains health conditions classified from all records and populations in the contributory regime [ 24 ]. 4 Its value lies in the richness of the available information: for each individual, we observe age, type of municipality, and diagnosed health conditions based on historical data from 2017–2019. However, the dataset is cross-sectional rather than longitudinal, which prevents the construction of a true panel. Consequently, we assess model fit by comparing predicted payments to actual costs observed in 2019. Demographic and geographic adjusters (age, sex, and location) are measured in 2019, while health conditions are classified using diagnoses from 2017. For this classification, we rely on the 255 diagnosis groups developed by the U.S. Centers for Medicare & Medicaid Services, as adapted to the Colombian context by Mejía, Duncan, and Ahmed [ 24 ]. Our empirical strategy therefore reflects the incentives and predictive challenges of a prospective formula, though the methodology could equally be applied in settings where concurrent models are used. Figure 2 shows the distribution of average monthly costs (in euros) by age and sex. Fig. 2 Average healthcare expenditure per capita by age and sex in Colombia. The graph shows the monthly healthcare expenses in euros, conditional on age and sex. Data from Mejía et al. [ 24 ]. These are not the official coefficients from the Colombian risk adjustment formula, but estimates based on our sample. Official coefficients published by the Ministry of Health follow a similar age and sex pattern, with minor differences in magnitudes due to data coverage
Average healthcare expenditure per capita by age and sex in Colombia. The graph shows the monthly healthcare expenses in euros, conditional on age and sex. Data from Mejía et al. [ 24 ]. These are not the official coefficients from the Colombian risk adjustment formula, but estimates based on our sample. Official coefficients published by the Ministry of Health follow a similar age and sex pattern, with minor differences in magnitudes due to data coverage
The graph presents the cost heterogeneity that is already considered in the Colombian risk adjustment formula: sex and age, and it shows a distribution that is consistent with the life cycle of the costs for both genders. However, as discussed, the risk adjustment formula in Colombia does not account for heterogeneity within disease categories, grouping diverse patients under broad classifications. For example, enrollees with neoplasms have mean treatment costs nearly ten times higher than the population average, whereas those classified under “symptoms” incur costs about 1.7 times the average.
A limitation of our study is that the administrative dataset we analyze covers approximately 10 million enrollees, while the total insured population in Colombia in 2016 was about 46 million [ 11 ]. This discrepancy arises for two main reasons. First, the subsidized regime, which accounts for about half of the insured population, is not included in our dataset. Second, even within the contributory regime, we restrict the analysis to the subset of affiliates with records that were used and found consistent for the period in Mejía et al. [ 24 ]. While this scope limits representativeness of the entire Colombian system, the contributory regime is the core setting where risk adjustment formulas are formulated and represents a substantial share of national health spending. Moreover, our objective is not to produce population level prevalence estimates, but to illustrate how the proposed methodology can be operationalized in a large-scale administrative database. Nonetheless, this gap underscores the importance of future work applying the framework to the subsidized regime and other populations once comprehensive data become available.
Results
This section presents results for different possibilities for choosing the parameters. In one group of estimations we include sex, age, and type of municipality of residence, which are currently used in the Colombian risk adjustment formula 5 and we consider the same parameter of penalization: \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s},$$\end{document} to all the variables, The second set considers the existence of different values of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}$$\end{document} and, in consequence, different \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} according to the prevalence. Finally, following expert recommendations, our third and preferred specification defines \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}$$\end{document} using the procedure described in Sect. 3.
For every group of estimations, we have three scenarios: setting the \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} with coordinate descent, 6 setting the \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} manually to 20.000 and 40.000 (Table 2 ). Table 2 Model Specifications Coordinate descent Lambda = 20,000 Lambda = 40,000 Equal \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} for every beta Model 1 Model 2 Model 3 Different \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} according to disease prevalence Model 4 Model 5 Model 6 Different \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} according to the experts advice Model 7 Model 8 Model 9 The table presents the combinations of the penalization parameter ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} ) values and the corresponding variability of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} across the available variables
Model Specifications
The table presents the combinations of the penalization parameter ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} ) values and the corresponding variability of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} across the available variables
We fit models 4–9, considering different \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} parameters for each variable and forcing the model to keep the coefficients used today in the Colombian risk adjustment model. Technically, this is implemented by setting the penalty factor for these variables to zero in the LASSO optimization problem. In other words, while all other candidate predictors are subject to coefficient shrinkage and possible exclusion, sex, age, and location are exempt from penalization and therefore always retained.
Figure 3 shows the number of variables selected in each model, including the scenario where no penalization parameter ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} ) is applied ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} = 0). In the Colombian official formula, only age–sex groups and location are included as risk adjusters. This is equivalent to treating all other candidate variables as if they were penalized with an infinite penalty, forcing their coefficients to zero. By contrast, our specification with \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda =0$$\end{document} allows all variables to enter without penalization, which is not the institutional practice but serves as a data-driven benchmark. The figure presents results for models built for 255 disease groups. Notably, when \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} is determined using coordinate descent (models 1, 4, and 7), we effectively remove 14 variables. 7 Coordinate descent algorithm iteratively updates one coefficient at a time while holding the others fixed, cycling repeatedly until convergence. This is the standard optimization method for LASSO because it efficiently exploits the separability of the penalty: by updating one coefficient at a time while holding the others fixed, it converges quickly even in high-dimensional settings [ 25 ]. Given our large dataset and repeated model specifications, coordinate descent provided a reliable and computationally feasible implementation. Fig. 3 Number of diagnostic variables retained in the models. The graph displays the number of diagnostic variables retained for different values of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} . Nine conditions did not appear in our sample. We chose to retain these conditions in the analysis for consistency to preserve comparability across specifications and with future applications of the methodology
Number of diagnostic variables retained in the models. The graph displays the number of diagnostic variables retained for different values of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} . Nine conditions did not appear in our sample. We chose to retain these conditions in the analysis for consistency to preserve comparability across specifications and with future applications of the methodology
Nine variables with no observations in our dataset are removed: Cancer of ovary (CCS40), Acute posthemorrhagic anemia (CCS58), Parkinson’s disease (CCS81), Multiple sclerosis (CCS82), Essential hypertension (CCS98), Hyperplasia of prostate (CCS162), Prolonged pregnancy (CCS182), Syncope (CCS241), and Nausea and vomiting (CCS246). Additionally, the penalization process excludes another 5 variables: Gout and other crystal arthropathies (CCS55), Anxiety disorders (CCS70), Hemorrhoids (CCS122), Female infertility (CCS172), and Hemolytic jaundice and perinatal jaundice (CCS220). For the specifications with penalization levels of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda =\mathrm{20,000}$$\end{document} and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda =\mathrm{40,000}$$\end{document} , the resulting lists of retained health condition variables are presented in the Appendix . The lower value ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\mathrm{20,000}$$\end{document} ) represents a moderate penalization level, which still allows many predictors to enter the model, while the higher value ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\mathrm{40,000}$$\end{document} ) represents a more stringent penalization, shrinking several coefficients to zero.
Even with a high penalization parameter, the models that set the \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} constant or positively depending on the prevalence of the health conditions, the LASSO regression still keeps variables as Residual codes, unclassified all E codes . This occurs because, despite their limited clinical interpretability, these categories capture non-negligible predictive signals in the data, often reflecting heterogeneous patient groups or coding practices that correlate with high expenditures. We do not interpret these codes as meaningful predictors per se, but rather as statistical placeholders that absorb variation not otherwise explained by more specific diagnostic groups.
Importantly, when expert judgment is incorporated through the multicriteria framework, these residual categories are assigned higher penalizations and excluded, since their lack of interpretability and susceptibility to manipulation make them undesirable for implementation. This illustrates a key contribution of our approach: combining statistical selection with expert oversight to avoid reliance on vague or residual predictors.
It is interesting that when we use medical expertise, the number of variables kept is higher. When we see Table 5in Appendix , it is clear that with this kind of estimation, a higher number of diagnoses related to neoplasms are kept in the final model. This is important because neoplasms are among the most persistent and high-cost conditions in the Colombian health system, generating predictable expenditure patterns over multiple years. Their inclusion improves the model’s ability to capture the spending of patients with severe chronic diseases, thereby reducing systematic underpayment for this group.
The literature shows that R 2 can be problematic for assessing insurers’ incentives for risk selection [ 26 ]. In particular, R 2 is difficult to interpret in this context because there is no benchmark; it is nonlinearly related to predictable profits and losses; and a high R 2 does not necessarily indicate better performance than a low one. To address these limitations, van Veen et al. [ 23 ] propose evaluating the difference between the average equalization-based predicted expenses and the average actual expenses for relevant non-random groups. This difference can be interpreted as over or undercompensation and as predictable profit or loss.
Accordingly, we complement R 2 with additional measures of risk selection incentives, such as the undercompensation of the upper decile of expenditure in the validation dataset, the mean absolute error (MAE), and the area under the curve (AUC). Unfortunately, our dataset for 2017–2019 is cross-sectional, which prevents us from implementing the ideal Colombian benchmark: a prospective specification where all the variables are defined using information from a previous period. Given this limitation, we evaluate model fit training the model on a 90% sample and assessing fit on the remaining 10% validation sample fit (as in [ 14 ]). 8
We use statistical measures such as \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${R}^{2}$$\end{document} , MAE, and UAC to evaluate predictive performance. If the model persistently underpredicts expenditures for patients with multiple chronic conditions, insurers may have incentives to avoid enrolling such individuals or to reduce the quality of care they provide. Conversely, persistent overprediction for relatively healthy groups may encourage cream-skimming practices. Yet these indicators provide only indirect evidence of the strength of risk selection incentives, what ultimately matters is how systematic prediction errors shape insurers’ economic behavior.
Regarding model fit, we estimate regressions using three specifications: (i) the variables currently included in the Colombian risk adjustment formula, (ii) the full set of available variables, and (iii) our proposed models. The resulting fit measures are presented in Table 3 . Table 3 compares the performance of alternative risk adjustment models that vary in their penalization structure (constant, relative to prevalence, or relative to expert panel recommendations) and the λ parameter used in LASSO estimation (coordinate descent, fixed at 20,000, or fixed at 40,000). The best results are obtained by the full model and by the coordinate descent specifications (Models 1, 4, and 7), achieving an R 2 of 8.51%, an AUC of 0.89, the lowest mean absolute error, and reducing residual underpayment to 43%. Models with λ fixed at 20,000 (Models 2, 5, and 8) deliver mixed performance. Model 8, for instance, achieves underpayment of 50% but at the cost of a lower AUC (0.84). In contrast, larger λ values (Models 3, 6, and 9) reduce model complexity but perform worse in model fit, leaving underpayment between 57 and 62%. Table 3 Model fit statistics for alternative risk adjustment specifications Model R2 Underpayment top decil MAE AUC Number of health condition variables All variables 8.51% 43% 94071 0.89 255 Model 1, 4, and 7 8.51% 43% 94070 0.89 241 Model 8 7.65% 50% 95700 0.84 64 Model 5 7.30% 52% 98356 0.85 38 Model 2 7.06% 51% 97965 0.86 34 Model 9 6.49% 57% 100989 0.81 26 Model 6 5.53% 62% 104850 0.81 12 Model 3 5.21% 61% 103805 0.82 11 Variables in Colombia 0.59% 82% 118877 0.7 0 The table reports the coefficient of determination (R 2 ), the percentage of underpayment for the top expenditure decile in the validation sample, the mean absolute error (MAE, in Colombian pesos), and the area under the ROC curve (AUC) for nine alternative models, as well as for the variables currently used in the Colombian risk adjustment formula
Model fit statistics for alternative risk adjustment specifications
The table reports the coefficient of determination (R 2 ), the percentage of underpayment for the top expenditure decile in the validation sample, the mean absolute error (MAE, in Colombian pesos), and the area under the ROC curve (AUC) for nine alternative models, as well as for the variables currently used in the Colombian risk adjustment formula
The restricted model using only variables currently available in Colombia illustrates the challenge regulators face: while it is simple and easy to implement, it explains less than 1% of spending variation, has limited discriminatory capacity (AUC = 0.70), and leaves 82% of underpayment uncorrected for the top decile. Adding variables indiscriminately can increase administrative complexity and create perverse incentives, but excessive parsimony exacerbates incentives for risk selection. Notably, model 9 has a tenfold increase in fit measured as R 2 while utilizing variables with low supervision costs and reduced upcoding risks. This model incorporates only 26 variables, compared to the best-case scenario (including all variables), which achieves a 15-fold increase in R 2 but comes with the previously mentioned drawbacks.
Figure 4 presents the mean squared error (MSE) of each risk adjustment model across population groups defined by diagnosis prevalence. The largest errors occur among individuals with conditions affecting fewer enrollees. Here, the model restricted to variables available in Colombia performs worst, with MSE exceeding one million, while the best-performing approaches are those with λ estimated via coordinate descent and certain relative-prevalence penalizations, which substantially reduce prediction error for these small groups. Fig. 4 Mean squared error of predicted versus actual amounts across prevalence categories for Colombia and nine model specifications. The y-axis shows mean squared error divided by 10 6 , with positive values indicating overcompensation and negative values indicating undercompensation. Each line represents a different model, as indicated in the legend
Mean squared error of predicted versus actual amounts across prevalence categories for Colombia and nine model specifications. The y-axis shows mean squared error divided by 10 6 , with positive values indicating overcompensation and negative values indicating undercompensation. Each line represents a different model, as indicated in the legend
For more prevalent conditions, MSE values drop sharply across all models, but differences in penalization approach still matter: larger λ values (Models 3, 6, and 9) produce consistently higher errors. In the most prevalent categories and among those without a diagnosis, MSE is comparatively low and differences between models narrow.
The results indicate that the models with more variables diminish the incentives for risk selection. Models 1, 4, 7, and with all variables reduce the under-compensation but don’t address the costs of including variables in the risk adjustment model. Model 9 reduces the under-compensation of the Colombian model including the variables kept after penalizing for the costs of the variables with experts’ opinion in the LASSO regression model. The ranking remains very similar to Table 3 .
In some specifications, residual diagnostic codes were retained. These codes capture uncategorized or unspecific conditions that, while not clinically informative, still explain a portion of expenditure variation. Excluding them would artificially improve model fit by discarding unexplained spending, so we kept them in certain models to reflect the costs that insurers face in practice. However, when expert criteria are incorporated, residual codes are systematically dropped, since they provide little clinical meaning and can be subject to manipulation. This contrast highlights an important contribution of the multicriteria framework: it allows the model to balance predictive performance with expert judgment on the policy relevance and potential adverse incentives of including certain variables.
Taken together, the results from both the model fit analysis and the prevalence-specific error patterns point toward the models incorporating expert opinion (Models 7–9) as the most balanced option for policy implementation. While these specifications do not achieve the absolute lowest residual underpayment or MSE, their performance loss relative to the best-fitting models is modest, and they avoid the sharp reductions in cost-control incentives that can arise when penalization heavily favors high-prevalence or highly predictive variables. By weighting variable inclusion according to expert assessments, these models maintain strong predictive capacity across prevalence groups while constraining the risk of overcompensation for easily upcoded conditions. This combination of adequate fit, fairness in payment distribution, and preserved efficiency incentives makes the expert opinion-based approach a pragmatic choice for regulators seeking to refine risk adjustment without undermining the financial discipline of health insurers.
We carry out a stability analysis of variable selection for four representative models: Model 5 (λ = 20,000, relative penalization by prevalence), Model 6 (λ = 20,000, relative penalization by expert opinion), Model 8 (λ = 40,000, relative penalization by prevalence), and Model 9 (λ = 40,000, relative penalization by expert opinion). The aim is to measure the consistency of selected variables when the models are re-estimated on different subsamples of the same population.
We apply a bootstrap approach by first drawing a random subsample comprising 70% of the total enrollee population. From this subsample, we generate ten bootstrap samples with replacement. For each bootstrap sample, we re-estimate the corresponding LASSO model, record the selected variables, and calculate the frequency with which each variable is retained across the ten runs. This procedure quantifies the stability of variable selection by showing how consistently predictors are chosen when the data vary.
Figure 5 shows the cumulative percentage of variables (y-axis) plotted against the maximum number of bootstrap samples in which they appear (x-axis). Importantly, a point such as 10% at x = 5 means that only 10% of variables are selected in five or fewer samples, while the remaining 90% appear in six or more samples. Lower percentages on the left side of the graph indicate greater stability, because most variables are consistently selected across many bootstrap replications. Fig. 5 Cumulative percentage of variables (y-axis) plotted against the maximum number of bootstrap samples in which they appear (x-axis) for four model specifications (Models 5, 6, 8, and 9)
Cumulative percentage of variables (y-axis) plotted against the maximum number of bootstrap samples in which they appear (x-axis) for four model specifications (Models 5, 6, 8, and 9)
By this metric, Model 8 stands out as the least stable specification: nearly 20% of its variables appear in five or fewer samples. In contrast, Model 6 exhibits the highest stability, with a very flat curve up to the midpoint, indicating that the vast majority of its variables are retained in six or more samples. Models 5 and 9 display intermediate but still high stability. Model 5 retains a substantial majority of its variables in six or more samples, though slightly fewer than Model 6. Model 9’s curve is close to that of Model 6.
Overall, the bootstrap stability analysis strengthens the case for models that incorporate expert opinion–based penalization. Both models with expert input achieve high levels of stability: Model 6 delivers the highest observed stability, while Model 9 attains nearly the same level, preserving expert oversight in variable selection. These results suggest that our proposal maintain most of the predictive accuracy of the best-fitting models, ensuring stability in implementation, and safeguarding incentives for cost control.
Choosing
As there are costs related to including variables in the risk adjustment model, the optimal number of variables included in the risk adjustment method could differ from the total number of available variables. In particular, in the presence of high supervision costs, for example, derived from upcoding possibilities, excluding some variables from the model is optimal. Therefore, the problem to solve under this framework is related to variable selection. In this context, "model" refers to the specific set of variables chosen to be included in the risk adjustment formula. The model determines how accurately the risk adjustment method can predict an individual’s healthcare expenses based on available data. Variable selection involves deciding which variables to include to balance the trade-off between prediction accuracy and the potential costs of including more variables, such as increased complexity, higher supervision costs, and the risk of upcoding. We use the term supervision costs to refer to the additional administrative and monitoring resources required to prevent or detect manipulation in coding practices. These include, for example, auditing activities, verification of diagnostic validity, and the broader oversight needed to ensure compliance with coding standards. While variable selection could also imply determining the size of the coefficients for each variable, the scope of this study is limited to analyzing the problem of selecting which variables to include in the model.
The most popular approaches in linear regression to choose variables for a model are the stepwise selection and regularization models. Stepwise selection is a method of fitting regression models where the choice of predictive variables is made through a procedure that considers, at each step, adding or removing a variable from the set of explanatory variables [ 20 ]. This type of procedure is made with forward stepwise selection and backward stepwise selection. With forward stepwise selection, the starting point is a model with no explanatory variables, and then predictors are added one at a time based on which one improves the model's fit. The best model is selected according to some criteria, such as the Akaike Information Criterion (AIC), 1 the Bayesian Information Criterion (BIC), 2 or adjusted \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${R}^{2}$$\end{document} [ 21 ]. In backward stepwise selection, the process starts with the total set of variables. It removes one at a time based on similar criteria than forward and then selects the best model in terms of AIC, BIC, or adjusted \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${R}^{2}$$\end{document} [ 21 ].
When selecting variables for a model, we face a combinatorial challenge. The total number of possible models that can be created from k variables is \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${2}^{k}$$\end{document} . For instance, with k = 255, there are \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${2}^{255}$$\end{document} possible combinations, making it computationally infeasible to evaluate all models individually. Stepwise selection methods, which add or remove one variable at a time, only explore a small subset of these models and may miss the optimal combination. In contrast, regularization methods like LASSO (Least Absolute Shrinkage and Selection Operator) and Ridge regression use all available information more effectively.
LASSO regression includes a penalty for the absolute size of the coefficients in its loss function. It can shrink some coefficients exactly to zero and effectively exclude those variables from the model. This automatic variable selection is particularly useful when our goal is to identify the most relevant variables while considering a cost parameter \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} , which reflects the cost of including each variable in the model. Ridge regression shrinks but never eliminates variables, and Elastic Net mixes both but adds complexity. LASSO thus aligns best with the goal of discouraging certain risk adjusters by zeroing them out. Thus, for our purposes, where the focus is on choosing a specific set of variables and removing others, LASSO regression is more suitable due to its ability to perform automatic variable selection and simplify the model. We use LASSO because it performs both shrinkage and variable selection, which is central to our approach.
It is possible to set \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} with the data using different methods. In particular, \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} can be chosen endogenously to improve the model's prediction. Also, the researchers can set the penalization parameter \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} exogenously. Note that if the policymaker defines that the risk adjustment contains k variables, \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} can be fixed to get the same number of variables. We explore the implications of both endogenous and exogenous settings of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} in the context of risk adjustment. Specifically, when setting \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} exogenously, we experiment with different values, including scenarios where a unique \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda$$\end{document} is assigned to each variable in the model.
The costs of having a variable are not necessarily the same for all the variables. For example, the costs related to having a variable of insured's previous hospitalizations are different from having a variable in the risk adjustment, which reflects if the person has reported having hypertension. Despite limitations in information on cardinal cost measures for each variable, 3 our analysis allows us to estimate and compare the relative costs of variables within a given dataset for risk adjustment purposes. We run the LASSO regression with different penalization parameters per every health condition variable, as shown in the following loss function: 2 \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$L\left(\beta \right)=\sum\nolimits_{i=1}^{n}({y}_{i}-{{x}{\prime}}_{i}\beta {)}^{2}+ \sum\nolimits_{s=2}^{k}{\lambda }_{s}|{\beta }_{s}|,$$\end{document} where \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\lambda }_{s}$$\end{document} = \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\lambda {\tau }_{s}$$\end{document}
In our implementation, the LASSO estimator minimizes the sum of squared residuals plus a penalization term proportional to the absolute value of the coefficients, weighted by the penalty parameter λ. The sum of squared residuals reflects the model’s predictive error across enrollees, while the penalization term captures the potential negative consequences of including additional risk adjusters, such as administrative burden, upcoding, or reduced incentives for prevention [ 1 , 3 ]. It is important to clarify that the residual sum of squares is not assumed to be a direct measure of selection incentives. Rather, in the absence of an explicit behavioral model of insurer selection, we treat predictive error as an indirect proxy for the extent to which relevant cost variation remains unaccounted for by the risk adjustment formula, variation that could, in turn, create incentives for selection [ 22 , 23 ].
For example, adding a variable such as prior hospitalization status may improve statistical fit by capturing substantial variation in expenditures. However, it may simultaneously weaken cost-control incentives if insurers can influence coding intensity or service provision in ways that increase the likelihood of hospitalizations being recorded. In this case, the improvement in predictive R 2 does not necessarily translate into reduced selection incentives, since residual variation remains an imperfect proxy for the economic gains insurers can achieve through risk selection. This illustrates why evaluating model performance requires balancing statistical accuracy with the incentive properties of candidate adjusters.
Our approach does not derive from a fully specified economic model but rather implements one component of an integrated framework: variable selection under penalty constraints. The LASSO method minimizes the residual sum of squares subject to penalization weights that reflect potential gaming risks identified by a panel of experts. In this way, the method illustrates how statistical optimization can incorporate qualitative evaluations, while recognizing that it does not capture all incentive effects that a comprehensive economic model would address.
We implement two alternative approaches to define the vector \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}$$\end{document} . In both approaches, individual weights are given by \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}=1+{r}_{s},$$\end{document} where \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}_{s}$$\end{document} is the rank of health condition s according to one of the two ranking measures introduced below.
In the first scenario, supervision costs are assumed to be proportional to the prevalence of each health condition. Let \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}$$\end{document} denote the prevalence rank of condition, with \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}$$\end{document} = k (total number of conditions) for the most prevalent condition and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}$$\end{document} = 1 for the least prevalent. We define \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}_{s}=\frac{{p}_{s}}{k}.$$\end{document}
In this scenario, we operationalize gameability through condition prevalence, employing it as a practical proxy. Nonetheless, prevalence provides only a partial measure and does not encompass all dimensions of manipulability. For instance, diabetes is highly prevalent but relatively easy to verify. Yet in administrative datasets, high prevalence often reduces the regulator’s ability to monitor and detect gaming effectively. We rank conditions by prevalence rather than using a continuous index. Ranking provides a simple and robust way to incorporate relative gameability into penalization, and it avoids undue sensitivity to extreme prevalence differences (e.g., 3% vs. 75%). Nonetheless, future work could explore indexed measures that preserve absolute magnitudes.
Our second and preferred approach involves setting the vector of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}$$\end{document} based on medical experts' advice. To translate the qualitative concerns about diagnostic precision, upcoding, and manipulation risks into the penalization parameters used in our LASSO framework, we convened a panel of physicians with direct experience with the Colombian risk adjustment system and coding practices in the providers, health insurers, and in the regulator institution. We gathered a group of experts knowledgeable about the nuances of the Colombian healthcare system and familiar with the common challenges of diagnosing and reporting health conditions.
The panel followed a structured evaluation process. First, they assess the 255 health conditions by groups as shown in the next figure (Fig. 1 ). Fig. 1 Processes used by the experts’ panel to assess diagnosis groups (CCS categories) for inclusion in the risk adjustment model. The panel comprised specialists with experience in provider coding, health insurance operations, and risk adjustment regulation. Each expert reviewed CCS groups from Infectious and parasitic diseases (e.g., CCS1, CCS2) to Symptoms, signs, and ill-defined conditions (e.g., CCS253, CCS254), to evaluate their relevance and susceptibility to manipulation
Processes used by the experts’ panel to assess diagnosis groups (CCS categories) for inclusion in the risk adjustment model. The panel comprised specialists with experience in provider coding, health insurance operations, and risk adjustment regulation. Each expert reviewed CCS groups from Infectious and parasitic diseases (e.g., CCS1, CCS2) to Symptoms, signs, and ill-defined conditions (e.g., CCS253, CCS254), to evaluate their relevance and susceptibility to manipulation
The panel independently rated each health condition on three attributes using a 10 point scale: diagnostic precision ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${d}_{s}$$\end{document} , where 1 = poorly defined or vague diagnoses and 10 = precise, clinically validated diagnoses), likelihood of upcoding or manipulation ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${u}_{s}$$\end{document} , where 1 = very low and 10 = very high), and predictability of disease occurrence ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${q}_{s}$$\end{document} , where 1 = very high and 10 = random). Ratings were averaged across experts to obtain \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\overline{{d }_{s}}, \overline{{u }_{s}}$$\end{document} and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\overline{{q }_{s}}$$\end{document} . The composite penalty weight for condition s was then calculated as the simple mean of these three scores, \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\omega }_{s}=\frac{\overline{{d }_{s}}+ \overline{{u }_{s}}+ \overline{{q }_{s}} }{3} ,$$\end{document} with \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\omega }_{s} \in \left[0, 10\right]$$\end{document} .
In addition to the risk of upcoding, both diagnostic precision and the predictability of disease occurrence are closely related to adverse incentives. Diagnoses with low precision give providers greater discretion in coding, which increases the potential for manipulation. By contrast, conditions with low predictability do not allow insurers to anticipate future costs and therefore do not create strong risk-selection incentives. For this reason, such conditions are less useful for risk adjustment and are penalized more heavily in the variable selection process.
We then sort the health conditions by their composite score and assign each condition a rank. This rank is used to calculate the specific penalization parameter where \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s} \in \left[1, 2\right]$$\end{document} as \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}=1+ {r}_{s} ,$$\end{document} where \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}_{s}=\frac{ {p}_{s} }{k}$$\end{document} is the rank of condition s (with the highest score assigned \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}=k$$\end{document} ), and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$k$$\end{document} is the total number of conditions. An example is provided below (Table 1 ). Table 1 Example calculation of the penalization parameter \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}$$\end{document} based on health condition scores. Scores are ranked from lowest to highest Score Rank \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\boldsymbol{r}}}_{{\boldsymbol{s}}}$$\end{document} \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\boldsymbol{\tau}}}_{{\boldsymbol{s}}}$$\end{document} 1.5 1 1/255 1 + (1/255) 2.0 2 2/255 1 + (2/255) 3.1 3 3/255 1 + (3/255) 5.0 4 4/255 1 + (4/255) … … … … 9.9 254 254/255 1 + (254/255) 10.0 255 255/255 1 + (255/255) The relative rank \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}_{s}$$\end{document} is computed as \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}/k$$\end{document} , where \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}$$\end{document} is the rank of the condition and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$k=255$$\end{document} is the total number of conditions. The penalization parameter is given by \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}=1+ {r}_{s}$$\end{document}
Example calculation of the penalization parameter \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}$$\end{document} based on health condition scores. Scores are ranked from lowest to highest
The relative rank \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}_{s}$$\end{document} is computed as \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}/k$$\end{document} , where \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${p}_{s}$$\end{document} is the rank of the condition and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$k=255$$\end{document} is the total number of conditions. The penalization parameter is given by \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\tau }_{s}=1+ {r}_{s}$$\end{document}
Vague residual diagnostic codes (e.g., "unclassified all E codes") consistently received the highest penalties, while neoplasms and chronic conditions like congestive heart failure were consistently prioritized as reliable adjusters.
The expert panel primarily focused on evaluating each potential risk adjuster for its vulnerability to upcoding and other forms of gaming, rather than directly quantifying selection incentives. In our framework, selection incentives are captured indirectly through the residual sum of squares in the LASSO objective function, which reflects the degree of unexplained variation in spending after risk adjustment. Nevertheless, the panel considered extreme cases in which selection incentives were deemed negligible, such as diagnoses linked to accidents or acute events unrelated to underlying disease patterns. Within a prospective risk adjustment model such as Colombia’s, these variables add little predictive value for future costs, since they largely reflect unpredictable shocks rather than underlying morbidity. For this reason, the multicriteria framework and expert panel agreed to assign higher penalizations to such variables. This was a structured way of ensuring that the model focuses on conditions that persist over time and better capture systematic differences in health risks across populations.
A possible limitation of using an expert panel is that members may carry specialty-related biases when assessing the gameability of conditions; for example, being more lenient toward diagnoses within their own clinical domain. While our exercise provides a structured way to incorporate clinical judgment, this potential bias highlights the need for caution in implementation. In practice, regulators would need to complement expert input with independent validation mechanisms to ensure consistent and unbiased identification of gameable conditions.
Conclusion
Our study highlights the importance of accounting for heterogeneity in the costs associated with different variables in risk adjustment models. By incorporating variable-specific penalization parameters, we address variations in supervision, upcoding practice, and disincentives to disease prevention, which can improve model accuracy while preserving its practical relevance. This approach enhances the understanding of how cost differences can be integrated to improve the performance of risk adjustment formulas, particularly within the context of the Colombian healthcare system.
Our research contributes to the health economics literature by providing empirical insights into assessing this trade-off in Colombia. Our empirical contribution involves proposing a methodology with penalized regressions, specifically the LASSO Regression, to statistically select variables dealing with potential adverse incentives and address the trade-off in a real world context using a comprehensive dataset from the Colombian health system.
We also apply a multi-criteria analysis to refine the selection process. Our findings indicate that involving multidisciplinary teams and diverse evaluation criteria is highly beneficial. In contrast, certain risk-adjustment methods generate gaming or manipulation costs that can outweigh the gains from improved model fit.
One limitation of our approach is that, while the penalization term explicitly accounts for the potential adverse incentives and downsides of including additional variables, the model’s treatment of selection incentives is indirect [ 8 ]. Specifically, the residual sum of squares captures unexplained variation in spending, which could be related to incentives for risk selection, but this link is not formally modeled. In reality, selection incentives depend on both the structure of the payment system and insurer behavior in response to residual under- or over-compensation for specific groups [ 22 , 27 ]. A more explicit incorporation of selection incentives would require either behavioral data on insurer enrollment strategies or a structural model of enrollee risk sorting under alternative payment formulas [ 3 , 28 ]. While this is beyond the scope of the present analysis, it represents a promising avenue for future research and could strengthen the empirical basis for regulatory decisions on risk adjuster inclusion.
The present study illustrates how expert opinion and statistical modeling can be combined to address some of the trade-offs inherent in risk adjuster selection. However, this is not a flawless or complete integrated approach. Moving toward a full integration would require building and calibrating an economic model that explicitly represents how risk adjustment parameters influence insurer behavior, enrollee selection, and overall welfare similar in spirit to the ex-ante evaluation framework proposed by van Kleef et al. [ 8 ]. Such a model could then be linked with the empirical variable selection method used here to enable optimization based on both predictive accuracy and modeled incentive effects.
Another limitation of this study is that model performance is evaluated using 2019 expenditure and sociodemographic data in combination with diagnostic information from 2017. This design shares similarities with concurrent payment models, since sociodemographics and expenditures are measured in the same year, but it departs from them by relying on lagged diagnoses. As a result, the reported fit is not directly comparable to either fully concurrent or fully prospective approaches, but it provides useful insights into how models perform when only earlier diagnostic information is available.
While our measure of underpayment in the top expenditure decile is informative, it is based on 2019 spending combined with diagnostic information from 2017. This design provides a more stable classification of persistently high-cost patients and better reflects insurers’ ability to anticipate expensive individuals. Nonetheless, it remains a proxy, as insurers cannot perfectly observe future expenditures; ideally, prior-year spending deciles would be preferable.
The relative distance between penalization scores can influence which variables are retained or excluded by the LASSO. This implies that careful calibration of penalty weights is essential, since overstating or understating the risks associated with one variable may distort the selection process. For policymakers, this highlights the need for transparent criteria and sensitivity analyses when operationalizing penalties.
Our selection procedure primarily treats predictive accuracy and gameability as substitutes, while allowing for potential complementarities through higher penalties on expert-flagged, high-gameability variables. We do not impose hard exclusion thresholds; instead, penalization allows the data to determine whether such variables are down-weighted or shrunk out. In settings where the incentive risk is judged to be very high by regulators or payers, a separate policy rule, as excluding specific codes by regulation or audit criteria, may be appropriate.
Our analysis implicitly assumes that regulators can detect and address gaming before it permanently distorts the underlying data. This reflects the Colombian context, where monitoring systems exist but are constrained resource. If instead data corruption occurs undetected, then highly gameable variables should not only be penalized but may need to be excluded entirely. Future research could extend our framework by incorporating thresholds that differentiate between conditions that can be safely monitored and those that cannot.
While our empirical application uses Colombian data, the proposed methodology is not specific to this institutional setting. The combination of penalized regression with expert judgment can be transferred to other health insurance markets, even where the availability or quality of diagnostic and expenditure data differs. The precise set of candidate adjusters may vary, but the general framework offers a replicable strategy for improving risk adjustment formulas in diverse institutional settings. In this sense, the Colombian case illustrates the method’s applicability beyond one country, highlighting its transferability to a wide range of health systems.
Literature
Over the past three decades, risk adjustment and health plan payment have become a focal point of research due to their critical policy implications [ 9 ]. This burgeoning field encompasses a rich body of research, including theoretical explorations of the link between risk adjustment and economic objectives in plan payments [ 11 ]. Empirical studies delve into classification systems, estimation methods, evaluations of existing systems, and simulation methods to assess the functionality of these payment models. Our work builds upon this foundation by introducing a novel empirical methodology for selecting variables within a risk adjustment system. This approach addresses the challenge of achieving optimal statistical fit while balancing predictive accuracy against the costs and potential adverse incentives of adding more variables.
Prior research has explored the application of constrained regressions and machine learning methods in health plan payments. For instance, Zink and Rose [ 12 ] propose penalized regressions for risk adjustment, focusing on a penalty that minimizes the net compensation disparity between under-compensated groups. In contrast, our work utilizes LASSO regression for variable selection within the risk adjustment framework. LASSO regression employs a distinct penalty function that directly shrinks the coefficients of the model toward zero, as expressed in the following equation: 1 \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\underset{\beta }{\mathit{min}}L\left(\beta \right)=\sum\nolimits_{i=1}^{n}({y}_{i}-{{x}{\prime}}_{i}\beta {)}^{2}+ \lambda |\beta |$$\end{document}
The loss function contains the squared errors between predicted and actual spending and simultaneously penalizes the absolute value of the coefficients with a weight λ. This approach discourages overfitting in the model, potentially leading to the exclusion of less impactful variables. This focus on variable selection allows us to improve model fit and potentially reduce the complexity, cost associated with the risk adjustment formula and potential adverse incentives.
Rose [ 13 ] compares the performance of various machine learning techniques, including regression, penalized regression, decision trees, neural networks, and a super learner, for estimating risk adjustment formulas. Using factors like age, sex, geography, and diagnoses as explanatory variables, the study reveals that variable selection can significantly reduce the number needed to achieve a desired model fit. While Rose [ 13 ] highlights the benefits of variable selection in reducing overfitting, we aim to minimize costs associated with incorporating numerous variables into risk adjustment models.
The work of McGuire et al. [ 9 ] presents the most closely related research to our study. Their analysis suggests that efforts to improve fit through more risk adjusters have reached a point of diminishing returns in some countries and sectors. While McGuire et al. [ 9 ] focus on minimizing the number of variables while maintaining a specific level of fit, our approach offers alternative flexibility as we do not impose a pre-determined fit level.
Our work explicitly addresses the trade-off of having more variables in the Colombian risk adjustment formula, considering that costs related to having a variable are not necessarily the same for all the variables. Variables related to specific conditions may be more costly due to factors like supervision requirements and the potential for upcoding. We apply differential costs per variable to account for this heterogeneity, penalizing those costlier variables more heavily under various scenarios. Our paper contributes to exploring how accounting for these cost differences can enhance risk adjustment, thereby improving both the model's fit and applicability.
Andriola et al. [ 14 ] developed a machine learning algorithm that utilizes Diagnostic Item (DXI) categories and Diagnostic Cost Group (DCG) methods to automate the creation of predictive models for risk adjustment. Their method enhances the risk adjustment formula by incorporating a score of Appropriateness to Include (ATI), which reduces reliance on vague and manipulable DXIs. However, their algorithm relies on p-values for variable selection and only examines the specified model, which may result in the omission of optimal variable combinations because it is unfeasible to evaluate all possible combinations of variables. In contrast, our approach employs LASSO (Least Absolute Shrinkage and Selection Operator), allowing for a more effective utilization of all available information in variable selection and model development because it evaluates all variables and selects variables under the information of all.
Risk adjustment formulas worldwide face a crucial challenge: balancing the need for accurate risk stratification with the potential drawbacks of having too many variables. In countries like Colombia, limitations in risk adjusters (sex, age, municipality) create incentives for insurer risk selection [ 15 – 17 ]. Studies like Riascos et al. [ 17 ] demonstrate significant fit improvements (R-squared from 1.45% to 13.53%) with the inclusion of additional variables. Conversely, countries like Israel prioritize limiting variables to address similar concerns. With its morbidity measure based on 80 HCCs [ 18 ], Germany exemplifies the persistent debates regarding the inclusion criteria for these conditions [ 19 ]. In this context, our paper proposes a statistical methodology to address this challenge. Our approach allows us to consider the potential drawbacks associated with each variable. This framework can inform data-driven decisions about risk adjustment formulas, balancing fit improvement with cost-effectiveness.
Introduction
The primary tool to face the problem of risk selection by health insurers is the use of risk adjustment methods [ 1 – 3 ]. The insurer's payment for an enrollee is adjusted by characteristics related to the possibility of experiencing differential costs [ 4 – 6 ]. Countries using risk adjustment include more variables in their formulas to reflect more risk heterogeneity of the insurer's enrollees. However, this approach may create tension, as reducing risk selection incentives through risk adjustment can also weaken incentives for cost control [ 1 ]. These costs include increased gaming, supervision, the potential for poor data quality, upcoding practices, and reduced incentives for disease prevention [ 7 , 8 ]. Risk adjustment formulas have traditionally prioritized comprehensiveness, but concerns about the downsides of including too many variables persist [ 9 ].
In current practice, modifications to risk adjustment formulas are often made incrementally, with the primary objective of improving predictive accuracy by adding variables that capture more variation in health spending. For example, the Netherlands has repeatedly expanded its morbidity adjusters to better predict costs for chronic disease groups, while Germany has added diagnosis-based indicators to strengthen payments for high-need populations. However, the regulator must compare well documented gains in fit with only hypothesized risks of upcoding and cost escalation [ 9 ]. Previous work has emphasized the need for an integrated framework to guide the design and evaluation of risk adjustment formulas. In particular, van Kleef et al. [ 8 ] describe such a framework. Our contribution is complementary: rather than proposing an entirely new conceptual cycle, we operationalize this idea empirically by combining statistical performance measures with expert-based assessments of potential adverse incentives.
An integrated methodology for selecting risk adjusters requires that the regulator, acting on behalf of society, systematically identifies both the potential positive and negative effects of each candidate variable and assigns explicit weights to these effects. Crucially, identifying these effects is not solely a statistical exercise; ideally, it would require an economic model that links the incentives generated by the payment formula to insurer behavior and, ultimately, to welfare outcomes. This paper does not aim to develop a complete economic model of risk adjustment. Instead, it shows the concept of an integrated methodology that balances predictive accuracy with potential adverse incentives, and it uses an empirical analysis with Colombian administrative records as an illustrative example of how such a methodology could be operationalized. While the empirical analysis is based on Colombian administrative records, the framework illustrates a general approach that could be adapted to other contexts. The focus of this paper, however, is not to provide a definitive universal solution, but to show how such a methodology can be implemented and assessed in a specific health system.
More variables can improve the accuracy of risk scores, but they can also lead to higher costs. To address this trade-off practically, we introduce a method that incorporates the costs of having more variables. We utilize penalized regressions, a statistical technique that imposes a penalty on the coefficients' size to prevent overfitting and to find the optimal number of variables while considering these costs. This approach allows us to select variables based on the specific methodology and cost parameters.
We use LASSO (Least Absolute Shrinkage and Selection Operator) regression, a type of penalized regression, to directly incorporate variable costs into the risk adjustment formula. LASSO is a type of regression analysis that aims to minimize the difference between predicted and actual outcomes, like ordinary least squares regression, and includes a penalty to encourage simpler models. As the strength of this penalty increases, this method shrinks the coefficients of less important variables, potentially reducing them to exactly zero. In this way, LASSO effectively performs variable selection, excluding variables that could distort insurers’ behavior. The resulting model retains the set of variables most relevant for balancing predictive accuracy with policy objectives.
We use data from the Colombian health system, which operates under a managed competition model with near-universal coverage. The government finances health insurance through regulated per capita payments (capitation) to insurers, who are required to provide the same benefit package to all enrollees [ 10 ]. While universal coverage eliminates free riding, insurers face persistent incentives for adverse selection and risk selection, as transfers cannot be risk-rated. Risk adjustment is therefore a cornerstone of the Colombian system, ensuring that transfers reflect risk differences while discouraging selective enrollment practices. Our analysis uses a comprehensive dataset covering all claims and health conditions for the contributory regime population, allowing us to test which variables should be retained under different penalization scenarios.
Our empirical strategy reveals the selection of desirable variables for risk adjustment related to some high-cost and chronic health conditions. However, when the penalization parameters are not calibrated considering medical experts' advice, the statistical procedure cannot eliminate undesirable variables, such as the ones that capture residual or unclassified diagnosis codes. When we calibrate the penalization parameters for each variable based on expert advice, the remaining variables in the model are those most suitable for risk adjustment.
This study contributes to the health economics literature by addressing the trade-off between prediction accuracy and cost containment incentives in the risk adjustment formula. Our empirical contribution lies in proposing a methodology incorporating this trade-off into a statistical procedure for selecting the risk adjustment model in Colombia. McGuire et al. [ 9 ] have an approach related to the objective of this paper. In their work, they start from the model with all the variables, and then they get rid of the variables, keeping the fit levels of the model. Our approach departs from traditional methods by prioritizing cost considerations alongside model fit. We incorporate a unique penalty parameter for each coefficient within the LASSO regression framework. This innovation is crucial because it enables our analysis to select variables while explicitly accounting for the heterogeneity in potential adverse incentives associated with each variable. In essence, our method considers the impact of a variable on model accuracy and its heterogeneous cost implications for the healthcare system.
We treat predictive accuracy and gameability as substitutes: improvements in one dimension typically come at the expense of the other. Nonetheless, in certain contexts, these objectives may operate as complements. For instance, a variable that is both highly predictive and easily manipulated can simultaneously enhance model fit while generating strong incentives for gaming. We recognize this as an important caveat and therefore prioritize higher penalties for expert-flagged, high-gameability variables, allowing the procedure to shrink or, when warranted by the data, exclude them. Regulatory oversight and systematic validation provide additional safeguards. We therefore handle complementarities via penalty up-weighting and robustness checks, while viewing any hard exclusion rules as policy decisions that can be layered on top of our framework.
The costs associated with including variables can vary significantly in risk adjustment models. The cost of incorporating a variable related to the insured's previous hospitalization differs from including a variable indicating whether an individual has reported hypertension. Although information on the exact cardinal costs for each variable is often limited, our study presents a novel approach to consider the relative costs of variables within a dataset for risk adjustment purposes in Colombia. By applying LASSO regression with varying penalization parameters for each health condition, we tailor the model to reflect the differing costs of variables, thereby improving risk adjustment accuracy and avoiding potential adverse incentives. We approximate the risk of adverse incentives using proxies such as prevalence and expert judgment. As a consequence, our penalization cannot be directly interpreted in monetary terms, but only as relative weights. If exact cardinal costs were available, the trade-off between predictive accuracy and potential adverse incentives could be expressed in a common cost–benefit metric, allowing policymakers to decide, for example, whether one percentage point improvement in R 2 is worth a given monetary increase in manipulation risk.
Our empirical exercise shows how a data-driven penalization framework can incorporate expert judgment on potential downsides of variables while preserving predictive accuracy. A fully integrated methodology, however, would also require an explicit economic model of how payment incentives influence insurer behavior and the resulting welfare effects. By framing the analysis as an empirical example, we aim to clarify both the potential and the limits of this approach.
The rest of the document is organized as follows: the second section presents a literature review; the third presents our methodology and proposal; the fourth section presents the data; the fifth section contains the results, followed by concluding remarks.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.