Integrating EHR and Survey Data via Machine Learning to Dissect the Rare Genetic Architecture of Alcohol Use Disorder: Findings from All of Us Biobank
Abstract
Background: Alcohol use disorder (AUD) etiology involves complex social (SDOH) and genetic factors. Traditional studies overlook the phenotypic spectrum of AUD with simplified definitions. We used the All of Us (AoU) biobank to integrate multi-modal data and identify genetic variations for AUD via a machine learning (ML)-derived risk score.
Methods: We identified 19k AUD cases and 190k controls in AoU. A CatBoost ML model was deployed to generate continuous individual-level liability scores as AUD risk scores across 7k features from EHR, SDOH, prescription and survey data. This score was then used as the primary trait in genetic analysis, including meta-analyzed GWAS and rare variant (RV) burden analysis for constrained geneset (pLI>0.9) and individual genes.
Results: The ML model achieved high validity (AUC = 0.94), separating cases (score=2.8±1.77) from controls (score= -0.29±1.02). ML model highlights top features of self-reported alcohol use, smoking, sex/gender, specific SDOH, and substance prescriptions, captured across EHR and survey data. Genetic analysis (N = 126,730) successfully recapitulated the known ADH1B variant (rs1229984; β=0.10, P=2.5×10-20). RV geneset analysis revealed increased AUD risk with rare deleterious variants (β ~0.05, P < 2.2 x10-6). Gene-based analysis identified more than 20 significant risk genes (P < 2.5×10-6) including known AUD genes such as POMC. Interestingly, it also discovered several novel genes such as CHRNA4, previously reported in other substance use traits, and are critical in brain neuronal and cognitive functions. These genes may explain a cross-trait addiction vulnerability.
Conclusion: Transforming binary diagnoses into an ML-based liability risk spectrum enhances the power to detect rare genetic contributors. This “deep phenotyping” reveals novel biological drivers hidden in traditional frameworks.
