Implementation of Machine Learning–Based Risk Targeting for HIV Care in Tanzania
Abstract
Background: Machine learning (ML) models have been developed to identify people living with HIV (PLHIV) at risk of disengagement from care and poor treatment outcomes. In Tanzania, the Rudi Kundini, Pamoja Kundini (RKPK) Objective 2 trial tested whether enhanced adherence counseling (EAC) with financial incentives improves viral suppression among PLHIV identified as high-risk using an ML algorithm. We evaluated the clinic-based implementation of this ML-guided strategy by comparing characteristics of individuals enrolled in the trial with those identified as high-risk but not enrolled.
Methods: An ML model was developed using electronic medical record data from HIV care clinics in Tanzania to predict disengagement from care, high viral load, and death. Individual risk scores were generated and a predefined threshold classified individuals as high-risk. High-risk individuals were approached during routine clinic visits for recruitment into the RKPK Objective 2 randomized trial. We compared characteristics of trial-enrolled participants with ML-flagged individuals who were not enrolled. Distributions of ML risk score, age, and time on antiretroviral therapy (ART) were compared using kernel density plots and Wilcoxon rank-sum tests; gender distributions were compared using chi-square tests.
Results: Across four ML risk list rounds, 3,087 individuals were identified as high-risk. Of these, 692 were enrolled and 2,395 were not enrolled. Compared with enrolled participants, ML-flagged individuals not enrolled had higher predicted risk scores (median 0.60 vs 0.57, p<0.001) and shorter time on ART (median 52.0 vs 62.2 months, p<0.001). The proportion of men was higher among unenrolled individuals (29.7% vs 23.6%, p=0.002), while age distributions were similar (median 32 vs 34 years, p=0.096).
Conclusions: The ML algorithm identified individuals at elevated risk of poor HIV outcomes, but those at the highest predicted risk were less likely to be reached and enrolled. Individuals most likely to disengage from care may therefore also be the hardest to reach through routine clinic-based recruitment. The large proportion of ML-flagged individuals not enrolled highlights the need for continual updates to outreach lists and improved engagement strategies, including timely EMR data use and proactive follow-up. Differences between enrolled and ML-flagged populations suggest direct trial estimates may not be externally valid. Formal transportability analyses, such as inverse-odds-of-sampling weighting, are a logical next step to estimate population-relevant effects.

