APP下载

Aquifer hydraulic conductivity prediction via coupling model of MCMC-ANN

2021-04-16GUIChunleiWANGZhenxingMARongZUOXuefeng

地下水科学与工程(英文版) 2021年1期

GUI Chun-lei,WANG Zhen-xing,MA Rong,ZUO Xue-feng

The Institute of Hydrogeology and Environmental Geology,Key Laboratory of Groundwater Sciences and Engineering,MNR,Shijiazhuang 050061,China.

Abstract:Grain-size distribution data,as a substitute for measuring hydraulic conductivity (K),has often been used to get K value indirectly.With grain-size distribution data of 150 sets of samples being input data,this study combined the Artificial Neural Network technology (ANN)and Markov Chain Monte Carlo method (MCMC),which replaced the Monte Carlo method(MC) of Generalized Likelihood Uncertainty Estimation (GLUE),to establish the GLUE-ANN model for hydraulic conductivity prediction and uncertainty analysis.By means of applying the GLUE-ANN model to a typical piedmont region and central region of North China Plain,and being compared with actually measured values of hydraulic conductivity,the relative error ranges are between 1.55% and 23.53% and between 14.08% and 27.22% respectively,the accuracy of which can meet the requirements of groundwater resources assessment.The global best parameter gained through posterior distribution test indicates that the GLUEANN model,which has satisfying sampling efficiency and optimization capability,is able to reasonably reflect the uncertainty of hydrogeological parameters.Furthermore,the influence of stochastic observation error (SOE) in grain-size analysis upon prediction of hydraulic conductivity was discussed,and it is believed that the influence can not be neglected.

Keywords:Grain-size distribution;Hydraulic conductivity;ANN;GLUE;MCMC;Stochastic Observation Error (SOE)

Introduction

The southwest part of the North China Plain(NCP) has been widely covered by Cenozoic unconsolidated sediments.For each local area of the plain,Quantization of hydraulic conductivity (K)at different scales can provide scientific support for groundwater exploitation and pollutant transport model-based groundwater resources assessment&pollution remediation.At present,in addition to laboratory tests,there are more numerical simulation inversions to be used to obtain hydraulic conductivity (Smiles and Youngs,1963;Namunuet al.1989;JI Rui-liet al.2016).Despite some merits,those methods have such demerits as high costs and limited applications due to regional scales(Koekkoek and Booltink,1999;DONG Pei,2010;Alfaro Sotoet al.2017).

Since the grain-size distribution is the most basic property of sediments and is easily available,researchers both at home and abroad attach much importance to the relationship between the grainsize distribution and hydraulic conductivity to predict the numerical value of hydraulic conductivity (Mahmoudet al.1993;Salarashayeri and Siosemarde,2012;FAN Gui-shenget al.2012).As a matter of fact,much attention has been paid to developing empirical or semi-empirical formulas on the basis of grain-size distribution data to provide reliable predication of the K value since the end of 19thcentury.The Hazen formula (David and Asce,2003) is the earliest model to predict K value that connects the grain size distribution with hydraulic conductivity through empirical coefficient.From then on,researchers have developed more empirical prediction formulas by selecting a certain effective grain size as their parameters,which can be applied to different kinds of samples (Russell,1989;Justine,2007).Compared with other methods of determining the K value,grain-size data has often been used to indirectly obtain the K value for it has been one of the most economical approaches to get K value,and moreover,it does not have any dependence on hydrogeological conditions of study areas during the computing process (Awad and Bassam,2001).However,there are certain risks to use these empirical formulas to predict K value.For instance,for the well-konwn KC equation,the constant CK-Chas been proved not to be a constant but a function between porosity and fractal dimension,which changes positively with porosity increase (XU Penget al.2011),thus enlarging the uncertainty of K value prediction of the equation.Furthermore,the most important limitation of these empirical and semi-empirical formulas is that all of them use only one or several parameters of grainsize distribution data,which may omit or neglect information that is included in complete grainsize distribution data system and correlated with hydraulic conductivity prediction.

In recent years,researchers have taken high interest in developing models that can be contrasted with grain-size distribution relationship to calibrate the K value.Artificial neural net (ANN),which serves as a widely used computational tool in a wide range of research areas,has played a satisfying role in the inversion calculation of the K value.Some researchers have employed ANN technology to establish hydraulic property models of the rock mass or soil,including the K value as a function of grain-size distribution.LI Shou-juet al.(2002) constructed a numerical method of K value recognition based on ANN technology and waterhead observation data of rock seepage field as well as a priori information of pumping tests,which has empirically been proved to be able to enhance the accuracy of water-head prediction.Nakhaei(2005) used 8 cumulative grain-size fractions for predicting log-transformed hydraulic conductivity for loamy sand with ANN technology,indicating that individual modeling of different soil types is superior to joint modeling.Hasanet al.(2006)used grain-size distribution,bulk density,and three different kinds of porosity as input parameters to build ANN model and multi-linear regression model for calculating K value of vadose zone soil,indicating that compared with multiple-linear regression model,ANN model produced more accurate results.TANG Xiao-songet al.(2007)used coarse-grained soil of Three Gorges Reservoir area as samples and got K values of different graduation soil through seepage experiments.Then,they employed the powerful nonlinear and dynamic processing capacity of ANN to predict K values.Trough a comparison between the predicted K values and the ones from the seepage experiments,the result indicated that it was feasible to use ANN to predict K of coarse-grained soil.Isiket al.(2012) used an ANN models which had three input parameters and one output parameter to predict the KL value of coarse-grained soils with different calculation methods of ANN,suggesting that those different calculation methods of ANN had almost the same prediction capability.

Nowadays most of applications of ANN for predicting K values are confined within the soil research area;there are few ANN applications,which take contents of clay,silt and sand as input variables of grain-size distribution,used for regional water resources assessment.Furthermore,in the framework of stochastic modeling and risk assessment,the quantification of uncertainty variables related to these predictions is of equal importance,which seems to be rarely mentioned and noticed.Therefore,this study combines the general likelihood uncertainty estimation (GLUE)with ANN technology to establish a holistic model for the aquifer hydraulic conductivity inversion and related uncertainty analysis,which takes complete grain-size distribution data of samples as input parameters of ANN model to predict hydraulic conductivity.As a long-term used nonlinear modeling approach,it is not uncommon to apply the ANN model to deduce hydrogeological parameters;however,it is not so common to couple ANN with GLUE as an overall model for the K value prediction and its uncertainty analysis.This paper tries to make an empirical study in an area of North-China Plain by coupling ANN model with GLUE.In addition,besides common uncertainty analysis of parameter,in recent years researchers has paid close attention to the influence of stochastic observation error on simulation results (WANG Donget al.2009;LU Le and WU Ji-chun,2010),the study also adequately discussed the influence.

1 Model construction

The model is established through the coupling between ANN and GLUE.The details are illustrated as follows.

1.1 ANN technology

As a deeply-applied computing tool in a wide range of research areas,the ANN can be regarded as a form of nonlinear regression,and multi-layer feedforward network has been proved to possess the property of universal approximation (Haykin,2004).Some particular ANN models (one type of their structures is shown in Fig.1) are capable of better predicting hydraulic conductivity (Erzinet al.2009;Park,2011;Daset al.2012).

Fig.1 Neural network architecture

1.2 Bayesian inference

The phenomenon of equifinality for different parameters (Keith,2006) leads to very large uncertainty for optimal selection of parameters of hydrogeological models.When the degree of uncertainty needs to be expressed,probability and probability distribution is the best language (MAO Shi-song,1999).As a method of probability analysis,Bayesian method currently is less applied in hydrogeological parameter identification process (LU Leet al.2008).The Bayesian method performs its statistic inference based on the general information,sample information and a priori information,and its density function is shown in Equation (1).

π(θ) is a prior probability that is the K value probability distribution obtained from rock and soil sample test with definite particle size component;

π(θ|x) is a posterior probability that is the proba-bility that a rock and soil sample with an uncertain particle size component has a certain K value.

1.3 MCMC-ANN coupling model

Markov chain Monte Carlo method (MCMC)was introduced into parameter uncertainty research to estimate the Bayesian distribution sampling of parameters in the 1990s (Smith and Robert,1993).Compared with Monte Carlo method,MCMC’s sampling efficiency is enhanced dramatically and the calculation work is remarkably reduced (GONG Guang-lu and QIAN Min-ping,2003).The single component adaptive metropolis algorithm(SCAM) (Haarioet al.2005) is adopted here to replace Monte Carlo method of traditional GLUE in the paper.SCAM counts the parameter set as a multidimensional vector in which each component represents a parameter.The predicting model of hydraulic conductivity constructed in the study is a coupling model between three-layered ANN and improved GLUE.An ANN model sample set is generated by considering the following random variations of the model parameters:(i) the variation range of each grain-size composition in the given study area;(ii) the quantity variation of the model hidden nods;(iii) initial values of the network weight and bias.The specific steps of model construction is as follows:(1) according to borehole data of the study area and related literature review,the range of parameter values of the model is determined;(2) a likelihood function needs to be selected,and the model efficiency coefficient R2is adopted here:

K0iis the actually measured hydraulic conductivity;Kpiis predicted hydraulic conductivity;is the mean value of actually measured hydraulic conductivity series;Nrepresents the length of the measured series;(3) according to a priori distribution of parameters,the initial sample X0of ANN model is randomly generated;(4) for theithcomponent of thetthsample,new sample componentziis generated by using one-dimensional normal distributionN(Xit-1’Vit),and the meaning and calculation method of parameters can be seen in the above-mentioned 24thliterature review;(5)new candidate sample component ziis accepted by a given probability α:

(6) steps 4 and 5 are repeated until all components of thetthsample is regenerated;(7) steps 4 to 6 are repeated until sufficient samples are obtained,and those samples need to be adjusted to reach the prediction accuracy set in advance;(8) selected likelihood function values are taken as measuring standard,and the samples that are not up to the standard are abandoned,and then the contrastive scatter plot between prediction values and measured values needs to be drawn;(9)it is required to define upper and lower bound of predicted K value and renew likelihood function value.Based on the defined valve value and the ranking order of likelihood function values,the ANN model uncertainty of given confidence level should be estimated.

2 Application example

The constructed GLUE-ANN coupling model has been applied to predict the K value of a study area in the NCP.Details are illustrated as follows:

2.1 Introduction to the study area

The study area is located in the southeast of Shijiazhuang city,southwest of Hebei province.It is situated in the northern latitude of 37°45′-37°51′ and the eastern longitude 114°34′-114°40′with an area of 126.78 km2,which is part of the inclined piedmont plain region where east Taihang Mount meets the NCP.The area belongs to the alluvial-proluvial-fan groundwater system of Hutuo River and Huaisha River of Ziya River basin,in which the groundwater mainly lies in the pore of quaternary unconsolidated rock layers.The aquifers below the area transit from single-layer,bi-layer to multi-layer structures.In the horizontal direction,single-layer thickness of the aquifers gradually thickens from west to east,its grain size becomes coarser,and the number of layers increases with the water yield property getting stronger.In the vertical direction,the grain size in the upper and lower layers is finer with relatively thinner aquifers,compared with coarser grain size and thicker aquifers in the middle layers.

Fig.2 Study area and borehole distribution

2.2 Data measurement

Six boreholes (zk05,zk06,zk07,zk11,zk14 and zk23) were selected in this study,and their distribution is shown in Fig.2.High-recovery and low-disturbance boosting core sampler was used to take 0.4 m undisturbed soil samples every 2 m,and then DZS70 constant head permeameter and TST55 variable head permeameter were applied to test the K value of sandy soil samples and clay soil ones.During tests,deionized water was injected to the bottom of the samples under a constant pressure and K value was determined by flow measurements in accordance with the Darcy’s law.For high clay contents,the measurement could last for 2 to 3 weeks until steady flow was achieved.In the case of more silty or sandy samples,a 1 m high water column was employed with the testing time not less than one day and the accuracy could reach about 10%.Meanwhile,samples used for grain size analysis were taken from each undisturbed soil sample.There were altogether 351 soil samples for grain size analysis,including 170 sets used for obtain contrastive data between grain size composition and actually measured values of hydraulic conductivity.

2.3 Generation of ANN model samples

By taking the model efficiency coefficient R2as the likelihood function,the SCAM method was adopted here to replace the Monte Carlo to take samples.12 000 ANN model samples were taken which contained actually measured values of grain size data.The a priori distribution of each parameter was tested by the Bayesian assumption,assuming that parameters within their value range subject to uniform distribution,as shown in Table 1.

Table 1 Parameters in the ANN model

2.4 Model training and uncertainty analysis

150 sets of data sets from borehole zk05,zk06,zk07,zk11,zk14 and zk23 were applied to train the models.For each K value of the 150 samples,it was assumed that the ensemble of model samples consisted of 1 000 ANN models.The number of models was set to 1 000 so that the estimated distribution can keep balance to achieve reasonable convergence and balanced calculation load.All of the analyses and calculations were processed by using relevant functions of neural network BP toolkit in the Matlab environment.Whole grain size fractions,initial weights,bias data,and the number of node in hidden layers were used as input parameters of the network and output data was taken as the samples’ hydraulic conductivity.Standardization of input is capable of speeding up neural network training and reducing the probability of being blocked in the local optimization process,so the input data of network was normalized actually-measured values of grain size fractions.In order to ensure that output data could range from 0 to 1,all the transfer functions used in hidden layers and output layers were the logsig functions.

Levenberg-Marquart rule was adopted to train the network and the maximal training step was set to 2 500.Error values were obtained according to Equation (4):

Where:E value is 10-14m/s.In the formula:E is the total error between actually measured and output values;pl are the grain size components;tkare actually measured values;ykare output values.Based on related literature review and the testing results,the minimum of hydraulic conductivity can be 1.12*10-9m/s,and it was used as minimum target value,and after squaring it,the error accuracy reached 10-18m/s.Therefore it is reasonable to value the total error E as 10~14 m/s.For minimizing the error,initial weight and bias needed to be reasonably selected.Evendistributed decimals between -1 and 1 are generally chosen.The number of node in hidden layers has larger random,and its optimal number depends on complexity of problem,which is usually determined by trial and error.Generally speaking,the number does not exceed twice as much as the number of input nodes.

After training the accuracy of the model reached the requirement.The value of likelihood function R2was 0.7.After those ANN model samples below the value were abandoned,5 520 ANN model samples were finally obtained.When likelihood function value was 0.82,the scatter diagram of output values vs.measured values is shown in the Fig.3.

Fig.3 The output versus measured value scatter plots (R2=0.82)

The critical value of model efficiency coefficient took 0.91 when uncertainty of model parameter was analyzed,and zero was assigned to the likelihood values of parameter sets below the critical value.The likelihood values of parameter sets above the critical value were normalized,and sorted by size of the likelihood values,the uncertainty range of models that are below the 90% confidence level needs to be calculated.

2.5 Application

It’s required to use the before-mentioned model to predict the hydraulic conductivities of 20 samples from the borehole zk11 of the same study area and the comparison results of predicted and measured values are shown in Table 2.

Table 2 Comparison between the ANN model’s output and measured values

In Table 2,K values were measured by DZS70 constant head permeameter;grain-size composition contents were calculated based on the grain-size testing report made by the monitoring center of groundwater &mineral water &environment,Ministry of Land and Resources,China.

The test data in the Table 2 was obtained from 20 sedimentary samples of aquifer and aquitard at various depths in the borehole zk11,which includes various lithologies from coarse sand,fine sand to silt clay.The predicted results showed that for silt and fine sand samples (in zk11-01,zk11-02,zk11-05,zk11-07 and zk11-19,etc.),ANN model has higher calculation accuracy with relative error from 1.55% to 5.58%.Most of the order of magnitude of predicted hydraulic conductivity values of these samples are 10-9m/s.With the content of clay in samples rising,the predicted orders of magnitude drop a little,which indicates the model has higher sensitivity to clay content.For coarse sand and medium coarse sand samples(in zk11-06,zk11-09,zk11-12 and zk11-13,etc.),the relative errors of predicted values are between 7.07% and 10.19%,and the prediction accuracy is a little lower than that of silt samples.Among all the samples,the predicted and measured values of zk11-06 and zk11-17 are the highest,which shows that the BP neural network can better reflect the rule that the higher sorting capability of the grain is,the larger their hydraulic conductivity is.For silt clay and some silt samples that have higher clay content (zk11-03,zk11-10,zk11-14 and zk11-15,etc.),the relative errors of predicted values of ANN model are between 9.56% and 23.53%.The relative errors and variation range are relatively larger,which indicates that the model needs to be further improved.

In order to validate the model’s applicability to variant sedimentary environment,the established model above was used to predict K values of 7 sedimentary samples in a scientific borehole of central region of NCP,and the comparison results of predicted and measured values are shown in Table 3.

Table 3 The comparison between the ANN model’s output and measured values of samples from middle region of NCP

In Table 3,K values were measured by DZS70 constant head permeameter;grain-size composition contents were calculated based on the grain-size testing report made by the monitoring center of groundwater &mineral water &environment,Ministry of Land and Resources,China.

In Table 3,seven sedimentary samples of borehole S18 that were obtained from aquifer and aquitard in different depth are contained,these samples consist of coarse sand,medium sand and fine sandet al.The data in the table shows that the relative errors of prediction results of GLUE-ANN model are between 14.08% and 27.22%,the prediction accuracy also conform to basic requirements of groundwater resources evaluation.

Posterior distributions of some models’ parameters that were obtained by using improved GLUE are shown in Fig.4,which reflect value range probability of parameters in the whole domain of definition.As the figure shows,the high probability areas of parameters’ posterior distribution are discontinuous,and the areas of global optimum are easily determined from the diagram.For instance,the posterior distribution of hidden node number appears to reach relatively high probability value when the number is around 14,which corresponds to the number of hidden node of hidden layers used in this study when the model efficiency coefficient is at higher levels.For most of the models,posterior distributions of their parameters have higher searching performance because of the application of SCAM method.The posterior distribution of D5,however,appears to be flat,indicating that the model still needs to be improved to enhance its sensitivity to some parameters.

Fig.4 Posterior probability distribution of GLUE-ANN model’s parameters

When R2is values as 0.91,uncertainty range of ANN model coefficient at the 90% confidence level is shown in Fig.5,which includes the actually-measured values of hydraulic conductivity and GLUE-ANN model’s prediction values.As shown in the figure,all of the prediction values are within two orders of magnitude of corresponding actually-measured values,and furthermore,most of the differences are within one order of magnitude.When K values are between 10-7m/s and 10-8m/s,the model prediction values have less deviation,demonstrating that the prediction and uncertainty analysis made by the constructed model are satisfying.

3 Stochastic observation error and uncertainty

In grain-size test of sedimentary sample,instrument’s measuring error and inexpert command of measuring skill can all cause observation error.Because the observation error is stochastic,using the data with stochastic observation error(SOE) to simulate aquifer parameter can lead to uncertainty of simulation results.

In order to be compared with uncertainty of simulation results with SOE,the uncertainty of a sedimentary sample for reference is analyzed firstly.The contents of medium sand and fine sand of the assumed reference sample are respectively 22.5% and 35.10%.After using MCMC to sample the reference sample,GLUE-ANN model is applied to predict and evaluate parameter uncertainty.When the number of samples was beyond 45 000,the series of reference sample converge to posterior distribution.As shown in Fig.6.

Fig.5 Actually-measured K values versus the GLUE-ANN model prediction and uncertainty estimation

Fig.6 The posterior distribution of refence sample

For inspecting the influence of SOE in grainsize analysis upon uncertainty of model output,the relative error that mean is 0 and variance is 0.001 is artificially added to the observation data of medium sand and fine sand of reference samples.Simulation and uncertainty analysis is performed using the observation data with SOE thereby to explore the effect of SOE on uncertainty of simulation results.When the number of samples exceeds 30 000,the samples converge to parameter’s posterior distribution as shown in Fig.7.

Fig.7 The posterior distribution of reference samples with SOE

As shown in Fig.7,compared to reference sample’s posterior distribution (Fig.6),the posterior distributions of the two parameters after adding SOE to them appear more scattered.It is evident that SOE has significantly increased the uncertainty of posterior parameters and consequently increased uncertainty of simulation results.

4 Conclusions

An improved GLUE-ANN model has been established in this study by coupling ANN technology and MCMC method,which replaces the Monte Carlo algorithm in traditional GLUE method.Compared with the Monte Carlo method,SCAM possesses superior searching performance for posterior distribution of model parameters,ensuring that the MCMC method can both search the global scale of posterior distribution and locate the extent of optimum values of parameters.Resultsof this study show that the improved GLUE-ANN model is capable of effectively predicting the K value and analyzing uncertainty of the parameters.

The advantages of the GLUE-ANN model created in the paper lies in taking fully into account the influence of different grain sizes and their content on permeability and the availability and universality of grain-size data in hydrogeological borehole data.The input parameters of the model adopted a complete grain size component consisting of clay,silt and gravel that can reflect the real soil structure in the study area.Therefore,the K values predicted by the model are closer to the value in natural state.If larger quantities of typical samples are available,predictions from the established model may become more rational.The study has not taken organic matter,carbonate content,soil porosity and its density as input parameters,and if these factors were considered,the prediction accuracy could be further enhanced.This reflects that there are many influencing factors on the hydraulic conductivity calculation leading to the uncertainty of predictions,which needs to be further studied.

Overall,the prediction results made by the GLUE-ANN model are relatively closer to actuallymeasured values with satisfying uncertainty estimation,which can serve as a precise way of predicting the hydraulic conductivity.It has been proved that it is capable of accurately predicting the hydraulic conductivity in the study area and other similar alluvial-proluvial plain regions and providing fundamental data for the improvement of solute transport model.

By way of adding SOE to observation data,the influence of SOE on simulation results is markedly reflected,which indicates that the uncertainty of simulation results caused by SOE can not be neglected.

Acknowledgements

This study was supported by Key Laboratory of Groundwater Sciences and Engineering,Ministry of Natural Resources (MNR) and the China Geological Survey project (No.DD20190252).


登录APP查看全文