Identification of Key Influencers for Secondary Distribution of HIV Self-Testing among Chinese MSM: A Machine Learning Approach

preprint OA: closed
📄 Open PDF View at publisher

Abstract

Background HIV self-testing (HIVST) has been rapidly scaled up and additional strategies further expand testing uptake. Secondary distribution has people (indexes) apply for multiple kits and pass these kits to people (alters) in their social networks. However, identifying key influencers is difficult. This study aimed to develop an innovative ensemble machine learning approach to identify key influencers among Chinese men who have sex with men (MSM) for HIVST secondary distribution. Method We defined three types of key influencers: 1) key distributors who can distribute more kits; 2) key promoters who can contribute to finding first-time testing alters; 3) key detectors who can help to find positive alters. Four machine learning models (logistic regression, support vector machine, decision tree, random forest) were trained to identify key influencers. An ensemble learning algorithm was adopted to combine these four models. Simulation experiments were run to validate our approach. Results 309 indexes distributed kits to 269 alters. Our approach outperformed human identification (self-reported scales cut-off), exceeding by an average accuracy of 11·0%, could distribute 18·2% (95%CI: 9·9%-26·5%) more kits, find 13·6% (95%CI: 1·9%-25·3%) more first-time testing alters and 12·0% (95%CI: -14·7%-38·7%) more positive-testing alters. Our approach could also increase simulated intervention efficiency by 17·7% (95%CI: -3·5%-38·8%) than human identification. Conclusion We built machine learning models to identify key influencers among Chinese MSM who were more likely to engage in HIVST secondary distribution. Key Findings (can also be found in Figure.2-Infographic) Our proposed ensemble machine learning approach outperformed human identification (self-reported scales cut-off) in accuracy & F 1 by classification metrics and in intervention efficiency by simulation experiments. Our model could also distribute more kits, find more first-time/positive-testing alters than human identification.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-07-29T07:08:50.896093+00:00