Integrating Geochemical Maturity Analysis and Machine Learning for Geothermal Prospect Identification
2026 SPE Conference at Oman Petroleum and Energy Show, OPES 2026, Muscat, Umman, 18 - 20 Mayıs 2026, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Doi Numarası: 10.2118/232357-ms
- Basıldığı Şehir: Muscat
- Basıldığı Ülke: Umman
- Orta Doğu Teknik Üniversitesi Adresli: Evet
Özet
This study investigates the use of machine learning to characterize geothermal water samples and infer potential geothermal reservoir sources in Colorado. Traditional geochemical interpretation is labor-intensive and often limited by subjective bias and incomplete data. This work applies exploratory data analysis, decision trees and K-means clustering to classify water samples by geochemical maturity and identify reservoir groupings, suggesting improved consistency and computational efficiency. A publicly available data set of 402 water samples containing temperature, depth, conductivity, pH, total dissolved solids, and major ion concentrations (Ca, Mg, Na, K, SiO2) was analyzed. Of the 402 samples, 196 contained complete geochemical information. After identifying missing and inconsistent values, a pre-processing workflow that consists of data cleaning, filtering, and feature-based subsetting was applied to the available data. Maturity levels were determined using the Giggenbach Na-K-Mg ternary diagram, generating subsets of immature, full and partial equilibrium samples. A decision-tree based classification model was trained that can classify water samples as either immature (IM) or partially/fully in equilibrium (EQ). A comprehensive exploratory data analysis, consisting of box plots and correlation matrices including density distributions and cross plots, was performed for data quality control. K-means clustering was applied to scaled and transformed geochemical and spatial parameters. Optimal cluster count was determined using the Elbow method as well as silhouette width analysis, and clusters were validated by mapping spatial distributions and comparing results with known geothermal prospects. The subset of full and partial equilibrium samples (72 samples) yielded the most reliable clustering results, as immature samples distorted reservoir-level patterns and resulted in lower cluster coherence and higher cluster separation. Testing different numbers of clusters revealed that 4 clusters offered the clearest separation, minimal overlap, and geologically meaningful grouping. Mapping the final clusters showed coherent spatial groupings that align with known geothermal trends in Colorado. Comparison with Colorado's heat flow maps and known geological descriptions of existing geothermal resources confirmed good agreement, with the cluster boundaries approximating established geothermal reservoir locations. Overall, the results demonstrate that K-means clustering, when combined with geochemical maturity screening, can provide a rapid and data-driven method to infer geothermal reservoir zones. This study introduces a streamlined machine-learning workflow for geothermal reservoir screening that integrates maturity classification, EDA-driven quality control, and unsupervised clustering validated against spatial patterns. The approach reduces dependence on manual geochemical interpretation and offers an objective, real-data driven scalable method for identifying reservoir signatures. It demonstrates how machine learning can enhance early-stage geothermal exploration by identifying reservoir-related hydrochemical domains from surface water chemistry data.