Formatted Title
What Does It All Mean? Recognizing the Uncertainty in the Data We Collect
Background/Objectives
The real world is heterogeneous and dynamic but, historically, because of a sparsity of data and investigative limitations, we have had to make significant simplifying assumptions. Consequently, the starting conceptual site model (CSM) that guided our data collection and analyses treated our sites as homogeneous and static. While we have made significant advances over the last couple of decades in our data collection methods via high resolution site characterization (HRSC) tools that provide more data than ever before. This new abundance of data, though, is a bit of a double-edged sword—as the details embedded in these data have the capacity to improve resource management and process understanding, we can also be lulled into thinking that this more data-rich environment means that we have more certainty. However, there is always going to be inherent uncertainty relative to our data accuracy/precision, scaling, and timing. The obvious question then is: How do we separate the wheat from the chaff? Or maybe more importantly, can we even identify the chaff?
Approach/Activities
As groundwater remediation professionals, we rely on data to inform our decisions—primarily focusing on the mass that moves. And, in a mass flux context, the fundamental pieces of data that are required are associated with the aquifer hydraulics (e.g., hydraulic conductivity, water-level data/hydraulic gradient) and the subsurface distribution of contaminant mass (e.g., contaminant concentration data). Ideally, our data is both precise and accurate, but the current state of technology is not there yet. As such, we examine and illustrate not only the potentially significant impacts that minor changes of these parameter values have on the dynamic and heterogeneous groundwater systems present at our sites (i.e., why this is important) but also the uncertainties associated with measurements of these properties (e.g., uncertainties associated with hydraulic conductivity estimates, automated and manual water level measurements and subsequent hydraulic gradient calculations, concentration measurement/laboratory errors, and, ultimately, mass flux estimates). We present numerous examples culled from our own experience and from the literature to demonstrate and highlight the importance of placing the data that we collect at our sites into the proper context of uncertainty.
Results/Lessons Learned
In this work, we outline some of the key limitations we face relative to site characterization data, and the potential implications that these limitations have on the data that we collect and how that can ultimately affect our remedies. As groundwater professionals, we need to recognize and accept that there is uncertainty all around us. What do we do with this knowledge? We always need to ask ourselves and answer some basic questions, such as: Does the data make sense? Are the conclusions supported by the data? Have we considered alternative interpretations of the data? Have we estimated the variability (uncertainty) in our assessment and the risks associated with the variability (uncertainty)? As remedial goals decrease (i.e., as with per- and polyfluoroalkyl substances) these questions become increasingly important as the risks from the uncertainties become amplified as remediation costs becomes intrinsically connected to treatment volumes rather than concentrations as remedial targets become exceedingly smaller. We also need to challenge ourselves not to solely rely on technology to derive/filter our interpretations and recognize that one of the most important tools we have at our disposal is our brain. Big data analytics/AI/machine learning are not substitutes for thoughtful review and dialogue within technical teams, professional judgement, and the practical experience of ourselves and others.