A lexicographic optimisation approach to promote more recent features on longitudinal decision-tree-based classifiers: Applications to the English Longitudinal Study of Ageing
preprint
OA: closed
Abstract
Supervised machine learning algorithms rarely cope directly with the temporal information inherent to longitudinal datasets, which have multiple measurements of the same feature across several time points and are often generated by large health studies. In this paper we report on experiments which adapt the feature-selection function of decision tree-based classifiers to consider the temporal information in longitudinal datasets, using a lexicographic optimisation approach. This approach gives higher priority to the usual objective of maximising the information gain ratio, and it favours the selection of features more recently measured as a lower priority objective. Hence, when selecting between features with equivalent information gain ratio, priority is given to more recent measurements of biomedical features in our datasets. We performed experiments with 20 longitudinal datasets created from a human ageing study, and concluded from our results that this lexicographic optimisation approach is worthwhile and computationally inexpensive. In addition to a slight improve in predictive accuracy, the changed feature-selection function promotes models based on information that is more directly related to the subject's current biomedical situation and, thus, more actionable.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00