Formatted Title
Predicting Closure for Thousands of Sites Using Machine Learning on Public Data
Background/Objectives
Predicting time to close is a major challenge for contaminated sites. Many efforts to do so rely on modeling of either attenuation rates or environmental fate to forecast when site objectives might be met. While necessary, these analyses are highly invasive, can take years, and have an environmental footprint that should not be ignored. Furthermore, they do not consider contextual site attributes that may also be important in predicting closures. These could include resource constraints on regulators, political leanings of communities, parcel zoning, consultant turnover, and many more.
This work takes an alternative – and perhaps complementary – approach to predicting time to close for sites. We mined publicly available data, and posed these questions:
-
How accurately can time to close be predicted using only public data?
-
How early in the site life cycle can a useful prediction be made?
-
What are the major drivers affecting time to close, and which of these are under-utilized as a conventional dataset?
Approach/Activities
Datasets were aggregated from public sources, including state-level contaminated site databases, EPA’s EJScreen, and the National Land Cover Dataset (NLCD), to name a few. More than 50,000 sites were analyzed. Feature engineering required a combination of tabular summarization, natural language processing, and geospatial analysis. Hundreds of initial features were calculated for each site, and subsequent statistical and domain-informed feature selection was performed. Machine learning was conducted as both a binary classification and regression problem. Multiple target variables were explored, including time to close and probability of closure within a given timeframe. Feature importance analysis was used to quantify the contribution of each feature to the prediction.
Results/Lessons Learned
Greater than 75% accuracy was achieved in predicting time to close based on only what was known in a site’s first year, suggesting strong predictive power may exist early in a site’s life cycle. Feature importance analysis revealed that many factors driving time to closure cannot be captured from only physical data gathered from sampling. Some of the biggest drivers included community demographics, impacted environmental media, and the nature of sampling activity early in a project. Most surprisingly, one of the strongest predictors was related to environmental justice (EJ) and reflected the socioeconomic condition of a site’s region.
This study suggests that a big-data approach to understanding contaminated site closure timelines is a valuable complement to invasive, physical closure models. Early prediction of closure timelines may be a useful resource for responsible parties, who can use this information to adjust their activities, set reserves, or minimize the environmental footprint required to describe a site. Regulators and government agencies may find these results useful to help guide resource allocation and policy.