APP下载

Forest fire smoke recognition based on convolutional neural network

2021-10-22XiaofangSunLipingSunYinglaiHuang

Journal of Forestry Research 2021年5期

Xiaofang Sun·Liping Sun·Yinglai Huang

Abstract Traditional fire smoke detection methods mostly rely on manual algorithm extraction and sensor detection;however,these methods are slow and expensive to achieve discrimination.We proposed an improved convolutional neural network (CNN) to achieve fast analysis.The improved CNN can be used to liberate manpower.The network does not require complicated manual feature extraction to identify forest fire smoke.First,to alleviate the computational pressure and speed up the discrimination efficiency,kernel principal component analysis was performed on the experimental data set.To improve the robustness of the CNN and to avoid overf itting,optimization strategies were applied in multi-convolution kernels and batch normalization to improve loss functions.The experimental analysis shows that the CNN proposed in this study can learn the feature information automatically for smoke images in the early stages of fire automatically with a high recognition rate.As a result,the improved CNN enriches the theory of smoke discrimination in the early stages of a forest fire.

Keywords Forest fire smoke·Convolutional neural network·Image classification·Kernel principal component analysis

Introduction

China is a large forestry country with several natural forest resources and planted forest plantations with management areas at the forefront of the world.However,forest fires are extremely destructive and are accompanied by many complex physical phenomena.In the early stages of a forest fire,there is smoke and fire plumes.Smoke spreads faster than flames,and combustion is often restricted to solid combustibles.Early stages of forest fires have the characteristics of slow burning and are difficult to detect.This often leads to many major fires;and smoke detection has greater value than flame detection.If the smoke can be spotted quickly and the fire can be put out,this can save considerable forest resources as well as lives and properties.As a result,smoke is the basis of forest fire detection.The sooner the smoke is detected,the earlier the fire is detected.Therefore,the identification of forest fire smoke is a topic that deserves special attention.

Numerous studies have been carried out on the detection of fire smoke.Presently,sensors and manual image feature extraction are commonly used.A large number of sensors can be placed to sample particles of combustion products.If sensors receive enough particles,they can analyze the changes in the humidity and temperature to implement fire detection.Manual extraction of smoke features is based on different algorithms.Yuan (2011) used local binary patterns(LBP) and local binary pattern variance (LBPV) methods to extract the LBPV feature vector of the images based on the texture features of the smoke to achieve fire smoke detection.Prema et al.(2016) proposed an image processing method to detect smoke using multiple features.They performed spatiotemporal and dynamic texture analysis on candidate smoke areas,while performing classifying and discriminating with support vector machines (SVMs) (Prema et al.2016).Shafar et al.(2017) first performed color analysis based on the HSV color space,and then applied a wavelet transformation to extract the smoke features to detect the smoke.Matlani and Shrivastava (2019) applied the Haar feature,Bhattacharya distance method,and the Gabor wavelet method to extract features of smoke images.However,the discrimination based on the sensor method is lengthy and expensive.The manual feature extraction algorithm is very complicated and depends on feature extraction before using detection technology.Only when feature extraction is reasonable,the recognition can be accurate.However,extracting effective features often requires professional knowledge and experience.This is mainly due to the large differences in the color and shape of the smoke.Therefore,it is necessary to develop a rapid,low-cost smoke detection method with a strong ability to extract features.

The forest image has the characteristics of being able to visually show the wide field environment in the wild,and it can achieve the tasks of rapid warning and accurate identification with an excellent recognition method.Convolutional neural networks (CNNs) have been widely used in the fields of image,sound,and natural language processing in recent years (Ding et al.2016;Ozer et al.2018;Poria et al.2016),and they have shown excellent performance.For example,Cirean et al.(2012) used CNNs to achieve efficient feature learning and to classify photos with remote sensing scenes.In 2012,researchers applied CNNs to ImageNet.This contained a large database in the field of image recognition for NIPS2012 and this achieved good results(Krizhevsky et al.2012).Ignatov et al.(2018) used CNNs to implement successfully continuous real-time activity classification in order to analyze human daily activities.Based on the above research,it is possible to prove the ability of CNNs to self-learn feature information and have a good classification performance.However,there are few studies that have applied CNNs to smoke image detection.

This study intended to apply a CNN for forest fire smoke image detection.Unlike the traditional manual method,the CNN could accomplish feature extraction and smoke identification in the early stages of a fire at the same time,and it could avoid the complexity and blindness of artificial feature extraction.The CNN has been further optimized through optimization strategies such as multi-convolution kernels,batch normalization,and custom loss functions.

Materials and methods

Experimental data sets

The scene photos taken by the ignition experiment and the relevant pictures of mobile phones in the internet.There are 1350 pictures in total and as shown in Fig.1,and the ratio of the training set to the testing set is 9:1.The selection of the training set and the testing set pictures is based on the principle of random seed.The size of the original image can be too large and this can limit the memory;hence,the size of the image calculated in the study is 90 pixels × 120 pixels.

Fig.1 Forest pictures in different scene

Image preprocessing

Image preprocessing is a necessary step before doing image classification tasks.In this study,the experimental data set was normalized;the transformation formula is described in Eq.1.In addition,the normalization operation was calculated pixel-by-pixel.Image normalization can prevent the effect of affi ne transformation and speed up the gradient descent to find the optimal solution.The reason for this isthat neural network training generally uses smaller weight values for fitting.When the value of the training data is a large integer value,this may slow down the model training process.

Kernel principal component analysis

The experimental data set is colored and relatively large in size.If it is directly used as the input of the experimental model,the number of calculations is large;therefore,data compression was performed in this study.Kernel principal component analysis is a non-linear extension of the principal component analysis.It has a better performance for nonlinear data and it can better match the nonlinear characteristics of the CNN.As a result,the data compression method for this study is kernel principal component analysis.The basic principle is to first perform nonlinear transformation on samples composed of time-domain and frequency-domain features.They are then mapped to a high-dimensional feature space and principal component analysis is applied to complete the feature extraction in the high-dimensional feature space (Hussein 2018;Shiokawa et al.2018).After several experiments,the radial basis kernel function was finally selected,the parameter gamma was 10,and the number of principal components was 256.

Convolutional neural network

A CNN cleverly reduces the number of parameters involved in the calculation through mechanisms such as weight sharing,and it achieves the effect that a fully connected neural network cannot achieve (Rocco et al.2019).In comparison to traditional machine learning algorithms and feature extraction algorithms,there is no need to manually extract features.Feature extraction and abstraction are automatically accomplished in model training;therefore,key feature information in the data can be learned better.

The basic structure of a CNN consists of five parts:an input layer,several convolution layers,a pooling layer,a fully connected layer,and an output layer (Fig.2).Details can be further divided into filters,step sizes,convolution operations,and pooling operations.CNN models construct a learning network structure through a combination of convolution layers and pooling layers.

Fig.2 Structure of a classic CNN

Convolution layer

The convolution layer performs feature extraction on the input image information and perform local convolution operations on the input signal through convolution kernel sliding.The network parameters are reduced through local connection and weight sharing mechanisms,thus,reducing the network complexity (Anwar et al.2018).The general feature extraction formula of the convolution layer is shown in Eq.2:

where,lis the number of convolution layers,f(·) the activation function,kthe convolution kernel,andbis the off set value.

Pooling layer

The main purpose of the pooling layer is to perform the dimensionality reduction operation.This minimizes the size of the array while maintaining the original features to reduce the experimental parameters (Deng et al.2018).Adding a pooling layer can increase the calculation speed,enhance the adaptability of the network to changes in the image size,and effectively prevent overf itting problems,thus,resulting in a robust system.The calculation formula for the subsampling layer is shown in Eq.3:

wheredown(·) is the subsampling function,ωis the weight value,bis the off set value,andf(·) is the activation function.

Fully connected layer

Several fully connected layers are usually designed behind the convolution layers and pooling layers to synthesize the features extracted by the convolution and the pooling layers.The fully connected layer will further extract these features to obtain other valuable information and classify these features.

Activation function

The activation function plays an essential role in the implementation of the CNN.It is a non-linear function that enhances the non-linear data transformation capabilities of the neural network model (Pandya et al.2018).It activates a part of the neurons in the neural network and transfers the information to the network structure of the next layer.The commonly used activation functions are the Sigmoid,Tanh,and Relu functions.The formulas for these three are shown in Eqs.4 − 6.

Experimental model

Model structure

(1) Input layer:The input variables are feature variables after pixel normalization and kernel principal component analysis;

(2) First convolution layer:there are 16 feature maps,the filter size is 7 × 7,and the step size of the convolution kernel is one;

(3) Pooling layer:when considering the maximum pooling method,the filter size is 2 × 2,the step size is two,and the number of feature maps after pooling is the number of feature maps,which corresponds to the previous convolution layer;

(4) Second convolution layer:there are 32 feature maps,the step size of the convolution kernel is one,and the filter size is 5 × 5;

(5) Pooling layer:when considering the maximum pooling method,the filter size is (2,2) and the step size is two;

(6) Third convolution layer:there are 64 feature maps,the step size of the convolution kernel is one,and the filter size is 3 × 3;

(7) First fully connected layer:f alttens the output variable of the third convolution layer as the input variable of this layer;

(8) Second fully connected layer:the neuron node is set to 512;

(9) Third fully connected layer:the neuron node is set to 64;

(10) Output layer:The output variable of the third fully connected layer is calculated by the Softmax function to complete the two classification tasks.

The CNN established in this research includes the input layer,three convolution layers,two pooling layers,three fully connected layers.The specific settings are:

Improved convolutional neural network model

The input variables of the CNN established in this study are the characteristic variables after normalization and the kernel principal component analysis.

In order to fully extract the differential feature information for the different types of training samples,this study increased the number of neurons in the convolution layer.This is under the premise that the number of fixed convolution layers is three and the different convolution kernel sizes are investigated in terms of the impact on the classification accuracy of the model training and test samples.To find the most suitable convolution kernel size,an experiment used three multi-channel convolution kernel techniques of 3 × 3 and 5 × 5,3 × 3 and 7 × 7,and 5 × 5 and 7 × 7,respectively,on the second convolution layer.

The setting of the learning rate has a great influence on the learning efficiency of a model.A large learning rate can improve the training speed;however,the learning accuracy is not insufficient.A small learning rate can improve the learning accuracy,but the learning rate is too slow.This study used a decaying learning rate strategy instead of a fixed numerical learning rate,with an initial learning rate of 0.01 and a decay index of 0.8.

In order to avoid overf itting,the Dropout mechanism was introduced after each convolutional layer.The retention rate parameter keep_prob was set to 0.8.As a result,20% of the neurons were randomly selected to be discarded and did not participate in the calculation.The EarlyStopping strategy was introduced to stop training while the model reached stability to improve the generalization ability of the CNN.The initial number of iterations was set to 50.

Batch normalization is introduced,which ensures that the output of each forward propagation is on the same distribution as the maximum,thus,eliminating the distribution differences between the layers.In addition,batch normalization can improve the generalization ability of the model.

Model training

For the CNN,the training process is divided into two stages:forward propagation and back propagation.The forward propagation stage occurs where the data is transmitted from a low level to a high level.Back propagation stage occurs where the error is transmitted from a high level to a low level when the results obtained by the forward propagation do not match the expectations.The loss function is the process of calculating the error;hence,it is very important to define the loss function.Cross entropy is a commonly used loss function in deep learning.This section introduces the L2 regularization strategy and regularization can avoid overf itting.Using a custom loss function,the calculation principle is the sum of the loss function and the parameter L2 regularization.The calculation formula of the optimized loss function is shown in Eq.7:

where,y iis the true category,a iis the predicted category value,and ω is the model weight.

Results and analysis

Selection of activation function

Choosing different activation functions may bring different performances to the model.The variables output by the Sigmoid function fall in [0,1],and the variables output by the Tanh function fall in [-1,1].The Relu function has advantages in terms of signal response and the prosperity data has better sparsity.Here,the activation function refers to the entire neural network model and it is shown in Table 1.The experimental results are only the average of three replications.

The CNN using the Sigmoid activation function performs better in smoke recognition in the early stages of forest fires(Table 1).There will be clouds and a cloudy sky in the image with or without fire.These factors are similar to the image information of the smoke;hence,the image data set with or without smoke is not much different.Studies have shown that the Sigmoid activation functions tend to perform well in situations where there is little difference between the features or when more subtle classification judgments are required.Therefore,in the following discussion,the activation functions all choose the Sigmoid function.

Evaluation of the multi-channel convolution kernel optimization strategy

The multi-channel convolution optimization strategy can better extract and identify the characteristics of smoke.In order to determine the appropriate convolution kernel size,this study applied the 3 × 3 and 5 × 5,3 × 3 and 7 × 7,and 5 × 5 and 7 × 7 to set the convolution kernel size in the second convolution layer.The performance of the improved CNN for smoke recognition is shown in Table 2 .

The optimal combination of the convolution kernel sizes is 3 × 3 and 7 × 7 (Table 2).For the convenience of description,the improved CNN model that applies the combination of the 3 × 3 and 7 × 7 size in the second convolution layer is referred to as CNN_S.

Comparative test

In order to verify the ability of the improved CNN proposed in this study to extract the key features of smoke and accurately identify them,the performance of the proposed model on the experimental data set is compared to a classic model that is suitable for classification tasks.The models compared here are the SVM and the K-nearest neighbor (K-,KNN).SVM is a machine learning method proposed (Widodo and Yang 2007).SVM maps the input vector to a high-dimensional feature space through predetermined non-linear mapping,hence,it can construct an optimal classif ciation surface in this high-dimensional space (Widodo and Yang 2007).The basic idea of KNN is to calculate the distance between the sample points in the data set and the current point (Gursel and Aghaei 2018),and then sorts the distances so they can be classified.In recent years,SVM and KNN have beenwidely used in classification tasks.The classification results of the different models are shown in Table 3.

Table 1 Comparison of the activation functions

Table 2 Multi-channel convolution kernel evaluation

Table 3 Model comparison

Fig.3 Change of Model accuracy

Fig.4 Change of Model loss

During the training and testing of the model,as the number of iterations increases,the model’s loss and recognition accuracy values change as shown in Figs.3 and 4.The test set recognition accuracy was 94.1% and the test set loss was at least 0.201.At the same time,it has the highest detection efficiency.

In order to obtain a better ref lection of the performance of the CNN,a confusion matrix was used to represent the smoke recognition ability of the model.This is a specific matrix used to visualize the performance of the algorithm.Each column represents the predicted value and each row represents the actual category.The confusion matrix of the CNN_S model is shown in Fig.5.

Fig.5 Confusion matrix

Conclusion

By quickly detecting smoke in the early stages of a forest fire,allows a faster response and the capability to extinguish the fires and save more forest resources.This study identifies a convolutinal neural network (CNN) that can detect smoke in the early stages of a fire.The CNN avoids reliance on relevant experience and knowledge;instead,it directly uses color images as experimental inputs to avoid information loss.

To better improve the performance of the discrimination method,the model uses batch normalization and multi-convolution kernel strategies to speed up the model training process.In addition,it avoids overf itting problems through model parameter L2 regularization.The experimental results indicate that the proposed model does not require complicated manual feature extraction operations.Finally,the model can simultaneously learn the features with a high recognition rate simultaneously.


登录APP查看全文