Formatted Title
How AI is Helping Improve Data Transparency and Regulatory Compliance for Environmental Sites
Background/Objectives
A big need in the environmental space is finding a way to capture environmental data to ensure regulatory compliance. For many sites, a vast and rich dataset exists in text form, and mining this text presents the opportunity to build or enhance our knowledge about a site. We have encountered multiple use cases where this need becomes clear:
-
Document Review. Upon site inheritance or during major project milestones, it is often necessary to “comb” through the sources of information for a site and extract both high- and low-level details.
-
Data Migration. Project teams may need to build a database of site information, but the data only exist in PDFs, forcing the task of transcribing analytical data, boring logs, and water levels en masse.
-
Regulatory Contextualization. Understanding the regulatory framework in which a site operates is paramount to a compliant, efficient, and sustainable site. The regulatory framework for a site is another dataset that is likewise encoded as text.
Mining historical text by conventional means is a monumental and error-prone task. We are utilizing the latest developments in artificial intelligence, such as generative AI, computer vision (CV) and natural language processing (NLP) to extract the value from text.
Approach/Activities
We have developed various custom AI models, tools, processes, and capabilities to help our clients turn text from documents into data assets. We will present a few successful use cases where we have developed and implemented custom AI to extract the value from text data:
-
Making content easier to find using a hybrid computer vision and natural language approach, which dynamically categorizes content down to the page level.
-
Reverse-engineering datasets from figures to eliminate transcription of information such as boring logs.
-
Extracting site details to populate site and portfolio summaries based on site-specific objectives.
-
Creating a chatbot for professionals to contextualize their site data against a site’s regulatory framework.
-
Improving the performance of predictive models that forecast closure timelines by using generative AI as an advanced feature extractor.
Results/Lessons Learned
AI/ML document mining capabilities are adding efficiencies and building new knowledge. These capabilities provide environmental professionals unprecedented access to text data. We are combining this data with information aggregated from dozens of other sources to provide our sites a data lake, or one-stop shop for information. Lastly, we use these capabilities to build features for machine learning models that can forecast site closures and costs.