Integrating wastewater-based epidemiology, socio-behavioral, and demographic data to enhance predictive modeling of SARS-CoV-2 spike risk in Virginia
Abstract
Integrating wastewater-based epidemiology, socio-behavioral, and demographic data to enhance predictive modeling of SARS-CoV-2 spike risk in Virginia
Background: Wastewater-based epidemiology (WBE) is an increasingly important tool for monitoring SARS-CoV-2 transmission at the population level. However, models relying solely on viral load trends may inadequately capture spatial heterogeneity in population vulnerability. We evaluated whether integrating longitudinal WBE trends with baseline socio-behavioral and demographic indicators over a multi-year window improves predictive accuracy of SARS-CoV-2 spike risk.
Methods: We constructed a longitudinal panel dataset comprising 10,632 monthly observations across 46 distinct wastewater treatment plant catchments in Virginia over a 5-year period from January 2021 through December 2025. Predictor matrices merged structural demographics from the American Community Survey (ACS) with health indicators from the Behavioral Risk Factor Surveillance System (BRFSS). Spike risk was defined as a binary threshold event (³1.5 x rise above the 2-week trailing mean). To prevent mathematical artifacts from traditional zero-imputation or complete-case deletion, missing values in population covariates were preserved natively within sparsity-aware splitting tree architectures. Predictive discrimination was evaluated on an independent, stratified hold-out test set using multinomial logistic regression, random forests, and extreme gradient boosting (XGBoost). Spatial clustering was assessed via Moran’s I; structural drivers were evaluated using a Generalized Linear Mixed Model (GLMM) with random intercepts for location, and temporal forecasting utilized log-transformed seasonal frameworks.
Results: Models integrating WBE, ACS, and BRFSS indicators outperformed models using environmental tracking alone. Accounting for panel clustering revealed honest, non-overfitted predictive power, with the native-NA XGBoost model demonstrating the highest discrimination (Test AUC = 0.6504), closely followed by random forest (AUC = 0.6448) and classification trees (AUC = 0.6253), while multinomial regression performed as a random baseline (Balanced Accuracy = 0.5042). Optimizing the classification cutoff to 0.312 via Youden’s J Index successfully mitigated naïve zero-spike estimation bias. In GLMM fixed-effects testing, poor physical health (PHLTH) was identified as the strongest socio-behavioral engine (OR = 1.198, p < 0.001), alongside healthcare access constraints (OR = 0.916, p = 0.008). Spatial analysis confirmed significant geographic viral clustering (I = 0.55, p < 0.001), while log-transformed time-series forecasting generated a consensus 3-month upward trajectory across the Commonwealth (ARIMA Month 3: 3.73, Prophet Month 3: 3.61).
Conclusions: Integrating environmental wastewater surveillance with native sparse socio-behavioral and census-derived structural datasets yields an honest, reproducible predictive signal that significantly improves public health risk stratification. Hybrid forecasting structures control for hidden geographic confounding and provide robust 3-month predictive windows to guide localized, targeted intervention infrastructure before clinical surges manifest.
Keywords: wastewater-based epidemiology; SARS-CoV-2; BRFSS; predictive modeling; health
