Optimization of sampling methodology for seismic exposure modelling: establishing homogeneous zones using machine learning
Autor
Rivera Álvarez, Juliana Esther
Fecha
2024Resumen
Loss estimation in seismic events requires an exposure model that captures the structural characteristics and spatial distributions of buildings and populations. This information is derived from existing data and virtual inspection conducted over zones with similar construction practices, referred as homogeneous zones. This study explores the feasibility of employing machine learning techniques to automate the creation of homogeneous zones (block clustering) in order to reduce time requirements and ensure replicability. It proposes a methodology that integrates socio-economic and cadastral data correlation of census blocks with location information. The approach begins with missing data estimation using a random forest algorithm. Subsequently, geolocation data enhances the dataset, computing distances to designated reference points on cartographic sheets. Principal Component Analysis (PCA) reduces data dimensionality, and the resulting variables inform a spatial density-based clustering model (DBSCAN) to determine the quantity of homogeneous zones. Finally, a Partitioning Around Medoids (PAM) model assigns each block to a homogeneous zone. This methodology optimally segments cities into homogeneous zones primarily considering economic aspects due to their high correlation with various construction practices.
