Assessment of Utility Statutory Records' Reliability
Léon olde Scholtenhuis
Organisations involved || Heijmans, Alliander, KPN, Kadaster BAO
Candidate || Ilya Gorbunov
Project Type || EngD
Period || Maart 2026 – 2028
While KLIC, the Dutch registry of underground utility records, is a pioneering system, it has a major problem: a significant portion of records are not accurate. This makes KLIC data hard to trust, since it's not always clear whether a given record is accurate. However, patterns in the data exist. For example, it is widely acknowledged that a telecom line in the city is much less reliable than a gas line in the field. Other patterns are known only to a handful of professionals, while others are perhaps yet to be discovered, unnoticed because they are too complex to pick up. All of this knowledge is undocumented, resting within the minds of experienced people, and their judgment on whether a given record is reliable has also never been tested. Given this, we ask: can assessing KLIC data become a data-driven process, rather than one resting on craft-based intuition alone?
This project aims to create a system that predicts how accurate a given KLIC record is, based on the thousands of trial trench records that exist in the industry. By comparing historical KLIC and trial trench data using machine learning, we identify which attributes — such as age, network owner, region, and utility type — influence accuracy, and use these patterns to predict how reliable a KLIC record is. Once trained on trial trench data, the model's future application would be to produce KLIC reliability scores for any record.
This project is split into three phases. The first and most crucial is gathering the data needed for modelling. Right now, the actual locations drawn in trial trench records are not explicitly linked to the KLIC records they correspond to — linking the two is a tedious process. Ilya has developed a web application that makes this linking process as easy as possible, enabling a crowd-sourced approach to gathering training data for the model
Secondly, once enough data has been gathered, the statistical modelling process can begin. The priority here is to understand how promising the model is: how well it performs, and whether that performance scales with more data. Other pressing questions include: since the linking process is itself imperfect, how do we account for variation and messiness in the data? And how do we calculate the confidence of the model's predictions?
The third and final phase concerns how the resulting model is further improved and operationalized. Figure 2 shows a proof of concept of how the model could eventually work in practice: users upload KLIC data and receive the model's estimations, along with further information on the records and the surrounding area. Beyond the interface itself, operationalization raises further questions: how will the different actors collaborate on this? How will a steady stream of new data be sustained? And how will the model change the research phase as it's currently practiced?
Overall, the three phases form a dynamic process that keeps iterating for as long as the model is being developed. The goal of the EngD is mainly to get started: to begin data collection, modelling, and working out how such a model will function in practice.