Key Frame Extraction of Surveillance Video Based on Fractional Fourier Transform
2021-10-12YunzuoZhangJiayuZhangRanTao
Yunzuo Zhang,Jiayu Zhang,Ran Tao
Abstract:With the vigorous development of national infrastructure construction and public information construction,video surveillance systems have gradually penetrated various fields.The current key frame extraction technology has inadequate target details and inaccurate judgment of local actions.Addressing this problem,a key frame extraction method based on fractional Fourier transform is proposed.This method obtained the phase spectra information of different orders by performing fractional Fourier transform on the surveillance video frames.Next,the method designed an adaptive algorithm based on the golden section point to select the transformation order.Then,the phase spectrum information of two adjacent frames was used to characterize the changes in the global and local motion states of the target.The final step was to extract key frames based on this.Experimental results show that,compared with the previous methods,the key frames extracted by the method proposed in this paper can correctly capture the changes in the global and local motion states of the target.
Keywords:fractional Fourier transform (FRFT);phase spectrum;key frame extraction;adaptation;local motion status
1 Introduction
In recent years,the safety awareness of people has gradually increased.The safety protection system has progressively penetrated various fields and has played an active role in maintaining the public order of the society and ensuring the safety of people’s lives and property.The video surveillance system is a vital part of the security system [1].The system can monitor the situation in a fixed area in real-time and play a massive role in preventing potentially dangerous events[2].With the continuous improvement of the surveillance video system,thousands of cameras have been deployed in the streets and lanes to monitor and collect information continuously.The exponential growth of video information in different fields and a large amount of data redundancy and other issues follow [3].
In order to overcome the difficulties of a large amount of video data and information redundancy while improving the efficiency of video retrieval and enhancing the accuracy of video retrieval,people have proposed key frame extraction technology.Luo et al.[4]proposed a method of key frame extraction based on Sub-shot Cluster.This method uses the color histogram features between adjacent frames to extract key frames by clustering,effectively avoiding key frame redundancy.[5]proposed a key frame extraction based on optimal distance clustering and feature fusion expression.This method considers the changes of the environment and the target motion state and further improves the quality and efficiency of video key frame extraction.[6]extracts the key frame of the vehicle for road monitoring,takes the multi-feature fusion image as a reference,and selects the video frame corresponding to the fusion image with the most significant target as the key frame.Generally speaking,surveillance video does not have structural features,and the shooting scene is relatively single.Therefore,researchers have proposed a target-based key frame extraction technology to reduce the interference of background information on the target.Li et al.[7]proposed a key frame extraction based on the ViBe algorithm for motion feature extraction,which improved the problem of inaccurate extraction of the target’s motion feature.In addition,[8]proposed extracting key frames from video abstracts in the HEVC compression domain.This method uses the number of patterns obtained by statistics to construct a pattern feature vector,which effectively improves the quality of video abstracts.Zhou et al.[9]proposed a fast extraction of online video summaries in the compressed domain based on visual feature extraction,which avoided redundant or meaningless frames in video summaries.In addition,image processing techniques from the perspective of time and frequency domains are often used in target detection,and many effective results have been achieved [10−12].
However,the existing method is not very accurate in judging the local motion state of the target,and there is a problem of insufficient detail extraction.For this reason,this paper proposes a method for extracting key frames of surveillance video based on fractional Fourier transform.This method obtained the phase spectrum of the video frame under different orders by performing a fractional Fourier transform on the surveillance video frame.Next,it designed an adaptive algorithm based on the golden section point to select the transformation order.Then,the method used the phase spectrum of two adjacent frames to accurately reflect the changes in the target’s global and local motion state to extract the key frames of the surveillance video accurately.
2 Basic Theory
2.1 Fractional Fourier Transform
As a new image processing tool,the fractional Fourier transform (FRFT) is different from the traditional Fourier transform that analyzes and processes stationary signals[13].The difference is that the fractional Fourier transform can be understood as a counterclockwise rotation of the signal at any angle around the origin in the timefrequency domain plane.Fractional Fourier transform is a generalized form of Fourier transform.As the transform orderpcontinuously increases from 0 to 1,all the signal characteristics gradually change from the time domain to the frequency domain [14,15].Taking a one-dimensional signal as an example.Thep-order fractional Fourier transform ofx(t) is defined as follows from the perspective of integral transformation[16]

Among them,prepresents the transformation order,the rotation angleα=pπ/2,Kp(t,u)is the kernel function of fractional Fourier transform,and its expression is as follows

The next step is extending the one-dimensional fractional Fourier transform to the two-dimensional fractional Fourier transform (2DFRFT).Since the two-dimensional fractional Fourier transform can express richer time-frequency domain information,it will also be used in image processing related work.Define the 2DFRFT of the image as the following [17]

wherep1andp2represent the transformation orders in the two dimensions ofxandy.represents the kernel function of the two-dimensional fractional Fourier transform,which is defined as

Among them,α=p1π/2 andβ=p2π/2 represent the rotation angle after the two-dimensional fractional Fourier transform.Whenα=π/2,β=0,the 2D fractional Fourier transform only Fourier transformx.Whenα=0,β=π/2,the 2D fractional Fourier transform only performs Fourier transform ony.Whenα=0,β=0,the 2D fractional Fourier transform will only perform identity transformation,that is,f0,0(u,v)=f(x,y).Whenα=β=π/2,the 2D fractional Fourier transform can only perform the traditional Fourier transform.
The original image is shown in Fig.1.The amplitude spectrum of the original image under different transformation orders is shown in Fig.2.The phase spectrum of the original image under different transformation orders is shown in Fig.3.

Fig.1 Original image
It can be seen from Fig.2 that as the transformation orderpincreases,the energy of the image is more and more concentrated in the central area of the image.At the same time,as the transformation orderpcontinues to increase,the image information tends to be blurred.This information indicates that the energy of the image is constantly shifting from the time domain to the frequency domain.It can be seen from Fig.3 that whenp=0.05,the phase spectrum information in the corresponding fractional domain contains the contour feature information of the image.As the transformation orderpincreases,the contour information gradually becomes blurred,untilp=1.00,that is performing the traditional Fourier transform,the edge texture information of the image completely disappears.

Fig.2 Amplitude spectrum of the original image in different fractional domains

Fig.3 Phase spectrum of the original image in different fractional domains
The above conclusions show that in increasing the transformation order to 1,the time-domain information in the amplitude spectrum and phase spectrum will show a trend of negative correlation aspincreases.At the same time,the frequency domain information will show a trend of positive correlation as it increases.
Taking Fig.4 as an example,when the transformation orderpis a fixed value,after the target’s motion state changes from standing to squatting,the frequency spectrum and phase spectrum of the original image have changed.From the frequency spectrum,we can directly see that the image belongs to the change in the texture contour of the time domain part.In the phase spectrum,we can find that although the shape of the graph has not changed significantly,the phase spectrum data of the two video frames have changed.

Fig.4 Amplitude and phase spectrum of two video frames
The amplitude spectrum of the image contains the gray information of the original image,that is,the brightness value that each pixel in the image should represent.In contrast,the phase spectrum of the image contains information about the edges and overall structure of the original image [9].Compared with the amplitude spectrum,the phase spectrum contains more visual information and image information.For this reason,when performing key frame extraction operations,the method in this paper chooses the phase spectrum information of the video sequence for related calculations.
According to the above analysis,we know that when the motion state of the target in the video changes,the amplitude spectrum and phase spectrum of the corresponding image will change accordingly.Therefore,we can capture changes in the global or local motion state of the moving target in the video.And then achieve the purpose of improving the inadequate extraction of details and inaccurate judgment of local actions.Based on this,this paper proposes a surveillance video key frame extraction method based on fractional Fourier transform.
2.2 Golden Section Point
The golden section method,also known as the optimal method,is a well-known section method that is most aesthetically appealing.The golden section method contains strict proportionality and aesthetics,so it is widely used in painting,music,construction engineering,industrial and agricultural production and many other fields.The classic golden ratio can be applied to a single peak function to find its optimal solution[19].Because of this,this paper designs an adaptive algorithm based on the golden section point.This algorithm can select the optimal solution when the image undergoes two-dimensional fractional Fourier transform and use this to determine the position of the transform orderp.Taking Fig.5 as an example,the length of the line segmentABisa,one of the pointsCis taken as the golden section point close toB,and the length ofACisb.

Fig.5 Golden section method
According to the golden section method,we can get the proportional formula,shown as

The length of the line segmentABisa,the length ofACisb,and the length ofBCis (a–b).Putting the above length into the balanced equation (5) can getb2=a(a−b),after simplification,If the length ofABis unit 1,then pointCis at a position with an approximate value of 0.618.The above calculation shows that pointCis the optimal solution in theABinterval based on the golden section algorithm,and its numerical expression is 0.618.Therefore,in the adaptive algorithm based on the golden section point,we can determine the value of the transformation orderp=0.6,and further filter and select the key frames.
Based on the above analysis,this paper proposes an adaptive algorithm based on the golden section point.With this algorithm,the transform order of the fractional Fourier transform can accurately capture the phase spectrum information of two adjacent frames.The global and local motion state changes of the target reflected under this transformation order can accurately extract the key frames of the surveillance video.
3 Algorithm Description
The algorithm proposed in this paper starts from the point of view of fractional Fourier transform and reflects the change of the current target’s global or local motion state according to the changes in the time-frequency domain information of two adjacent video frames surveillance video sequence.Based on the above analysis of the adaptive algorithm for the fractional Fourier transform and the golden section point,this section proposes the basic framework of the key frame extraction method for surveillance video based on the fractional Fourier transform,shown in Fig.6.

Fig.6 Key frame extraction algorithm framework
The specific steps of the surveillance video key frame extraction algorithm based on two-dimensional fractional Fourier transform are:
Step 1Generating sequence of surveillance video frames.The selected surveillance video is segmented according to the duration to generate the surveillance video sequence.
Step 2Image preprocessing.Performing image preprocessing on the surveillance video frame sequence generated in the previous step.The specific operation is grayscale processing.Since the grayscale processed image only contains brightness information,the grayscale image is more convenient to store.At the same time,it can improve computing efficiency.
Step 3Selecting the transformation orderp.According to the adaptive algorithm based on the golden section point,this step selects the appropriate transformation orderp.The value of the transformation orderpis finally determined to be 0.6.
Step 4Performing a two-dimensional fractional Fourier transform.Step 4 used the transformation orderpdetermined in the previous step to perform a two-dimensional fractional Fourier transform on the surveillance video frame sequence to obtain

Step 5Obtaining the phase spectrum information of the video frame.According to (6),can be obtained,and then the fractional Fourier transform phase spectrumcan be obtained

Step 6Calculating the phase spectrum MSE of two adjacent frames.After obtaining the corresponding phase spectrum information of the surveillance video frame sequence,the next step is to calculate the mean square error between two adjacent frames.The mean square error can be used for image quality evaluation,that is,to determine the difference between the original reference image and the distorted image.Using the evaluation index of mean square error,we can more accurately judge the change of the target in the surveillance video frame sequence.Assuming that the length and width of the video frame areN1andN2,respectively,the phase spectrum of the current video frame is denoted asf(x,y),and the phase spectrum of the previous video frame is denoted asg(x,y).We can obtain the mean square error MSE of two adjacent frames according to the definition

Step 7Forming the MSE curve.According to the obtained mean square error of adjacent video frames in Step 6,this step form a mean square error curve,that is,an MSE curve.
Step 8MSE peak point detection.The method was extracting the peak point from the MSE curve formed in Step 7.This step extracts the peak point detected in this step and use it as the basis for judging the key frame.
Step 9Determining candidate key frames.This step will extract all the local peak points in the MSE curve.We set the MSE of the first frame of the video to 0,so the second frame must change suddenly.Because the first and second frames will be redundant,start from the second peak.The video frame extracted comprises the video frame at the peak mutation and the first frame.The total number of candidate key frames is set toS.
Step 10Determining the final key frame.Step 10 selects the final required key frame from the defined candidate key frames.Firstly,the method calculates the difference between all the local peak points and the value of the previous frame and the next frame,and then,it calculates the average value simultaneously.

whereaiandbirepresent the value difference between the peak point and the previous frame and the next frame,respectively.avgaand avgbrepresent the average difference between the peak point and the value of the previous frame and the next frame,respectively.Next,extract the video frame whose value difference between the video frame and the previous and next video frames is greater thanM1avgaandM2avgbas key frames.Among them,M1andM2are parameters,respectively.Finally,since the method set the MSE value of the first frame to 0,the first frame in the video sequence is separately extracted as the key frame.According to the above operation,all the key frames on the MSE curve and the first frame in the video sequence are sequentially extracted and determined as the final key frame.
4 Experimental Results and Analysis
In order to verify the correctness and effectiveness of the key frame extraction method of surveillance video based on the fractional Fourier transform,this paper selects multiple public surveillance videos on SISOR to conduct multiple experiments.This experiment uses AMD Ryzen 7 4800U with Radeon Graphics 1.80 GHz processor,and the graphics card is AMD Radeon Graphics.The operating system is 64-bit Windows10 Professional Edition.The specific information of the video data set used in this experiment is shown in Tab.1.

Tab.1 Video dataset
4.1 Algorithm Correctness
Through the analysis of the experiment,it can be found that using the fractional Fourier transform to process the image and transform it into the time-frequency domain can effectively obtain the changes in the local and global motion states of the target in the surveillance video.Take the test video Video2 as an example.In the video,firstly,the moving target appeared in front of the camera,and then,he walked to the center of the camera,bended knees,jumped,landed,and finally walked out of the camera.In this video,the changes in the target motion state include changes in the global and local motion states.The changes in the global motion state include the appearance of the target in the surveillance area,the target walking,and the jumping.Changes in local motion state include the squat,landing,and changes in the hands and legs during the exercise of the target.In this experiment,we set the experimental parametersN=2 andM=1.Fig.7 below shows the Video2 key frame extraction results except for the first frame.
From the extraction results in Fig.7,we can see that the 7th frame and 61st frames are “when the target enters” and “leave the surveillance area”,respectively.The 16th and 54th frames are the walking processes of the target,which reflects the changes in the state of the target’s legs.The 32nd frame is the action of the target jumping.The above few frames represent the transformation of the global motion state of the target in the surveillance video,which has been successfully extracted.In addition,the 19th and 25th frames reflect changes in the target’s local motion state,such as raising the arm and bending the knee.In frame 37th,the target leans forward and lifts the toes after landing.These frames represent that the transformation of the local motion state of the target in the surveillance video has been successfully extracted.

Fig.7 Video2 extraction result of key frame extraction method based on two-dimensional fractional Fourier transform
The analysis of the above experiments verifies that the time-frequency domain information can obtain the global and motion state changes of the target.At the same time,it supports that the proposed method can more accurately extract the local motion state changes of the target.Based on this,the experiment in this section verifies the correctness of the surveillance video key frame extraction method based on the fractional Fourier transform.
4.2 Algorithm Effectiveness
In order to verify the effectiveness of the method proposed in this paper,we compare with the key frame extraction method of surveillance video based on frequency domain analysis [18].Fig.8 below is the extraction result of the test video Video4.Besides,setting the experimental parameters of this method toN=M=2.We will conduct experiments from both subjective and objective perspectives.Among them,the subjective evaluation method is qualitative evaluation,the most intuitive and effective evaluation method;the objective evaluation method is also called quantitative evaluation,which is suitable for automatic video analysis.

Fig.8 Video key frame extraction result of Video3
It can be seen from Fig.8 that both the proposed method and the comparison method can extract the target entry and exit monitoring screen.Nevertheless,the target extracted by the method proposed in this paper is more complete than the control method.At the same time,both methods can also extract the action of the target taking off the jacket,but the image extracted by the method proposed in this paper can better reflect the changes in the local motion state of the target,such as the image of the target lifting and removing clothes.
In order to further verify the effectiveness of the algorithm proposed in this paper,we use precision,recall,and F1 criteria to evaluate the performance of the algorithm.The F1 criterion represents the harmonic average of precision and recall.The experiment verifies the universality and robustness of the algorithm by conducting experiments on the 6 videos in Tab.1.Among them,the experimental parameters of this method are set toN=2,M=1,and the parameters in method[8]are set toN=1.The statistics of the extraction results of the two algorithms are shown in Tab.2.

Tab.2 Comparative experimental results
By comparing the method proposed in this paper and the comparison method,we can see that the method proposed in this paper is better than the comparison method.The algorithm proposed in this paper uses an adaptive method to convert the video sequence to the time-frequency domain and then calculate it to determine the key frame.This experiment finally achieved good results.The above analysis shows that the method proposed in this paper can more accurately not capture the changes in the global and local motion states of the target in the surveillance video.Simultaneously,the accuracy of the extracted key frames is also better than the comparison method,thus demonstrating the effectiveness of the method proposed in this paper.
5 Conclusion
This paper proposes a surveillance video key frame extraction algorithm based on fractional Fourier transform.Firstly,we use an adaptive algorithm based on the golden section to select the transformation order.Next,we perform a fractional Fourier transform on the surveillance video and obtain the phase spectrum of the image at the same time.Then,we calculate the mean square error of the phase spectrum of two adjacent video frames to form a mean square error curve.Finally,we select the final key frame according to the change of the local peak value of the curve.The experimental results demonstrate the correctness and effectiveness of the method.This experiment compares the method proposed in this paper with the previous algorithm from the perspective of objectivity and subjectivity and concludes that the method proposed in this paper can more accurately capture the changes in the target ’s global and local motion states.Therefore,it can more accurately solve the problems of insufficient detail extraction and inaccurate judgment of local actions in the current surveillance video key frame extraction technology.
杂志排行
Journal of Beijing Institute of Technology的其它文章
- Adaptive Turbo Equalization for Probabilistic Constellation Shaped Underwater Acoustic Communications
- Discrete Convolution Associated with Fractional Cosine and Sine Series
- Speech Encryption in Linear Canonical Transform Domain Based on Chaotic Dynamic Modulation
- Detection of T-wave Alternans in ECG Signals Using FRFT and Tensor Decomposition
- A New Tensor Factorization Based on the Discrete Simplified Fractional Fourier Transform
- Adaptive Short-Time Fractional Fourier Transform Based on Minimum Information Entropy
