APP下载

Adaptive Optimal Control of Space Tether System for Payload Capture via Policy Iteration

2021-09-26,,,*

,,,*

1.School of Automation,Northwestern Polytechnical University,Xi’an 710129,P.R.China;

2.Beijing Institute of Aerospace Systems Engineering,Beijing 100076,P.R.China

Abstract:The libration control problem of space tether system(STS)for post-capture of payload is studied.The process of payload capture will cause tether swing and deviation from the nominal position,resulting in the failure of capture mission.Due to unknown inertial parameters after capturing the payload,an adaptive optimal control based on policy iteration is developed to stabilize the uncertain dynamic system in the post-capture phase.By introducing integral reinforcement learning(IRL)scheme,the algebraic Riccati equation(ARE)can be online solved without known dynamics.To avoid computational burden from iteration equations,the online implementation of policy iteration algorithm is provided by the least-squares solution method.Finally,the effectiveness of the algorithm is validated by numerical simulations.

Key words:space tether system(STS);payload capture;policy iteration;integral reinforcement learning(IRL);state feedback

0 Introduction

With the development of aerospace industry,more and more spacecraft have been launched,resulting in a large amount of space debris in low earth orbit(LEO).Due to the complex space disturbance,the orbital altitude of the spacecraft changes,which can cause the collision between different spacecraft,resulting in a large number of space debris.Therefore,the safe and efficient capture of space debris is of great significance for the safe completion of space missions.Space tether system(STS)has been widely studied in debris removal[1-3],orbit transfer[4-5]and artificial gravity generation[6-7]due to its flexible structure and higher reliability of approaching space target than the space manipulator on the floating platform.

An on-orbit capture mission operated by space tether can be mainly divided into three stages:The deployment of tether before capture,rendezvous for capture,post-capture stabilization and retrieval[8].Due to complex space environment,there are possibly some errors of tether length and pendulum angle in the payload capture process.The capture mechanism installed at the end of the tether is integrated with the target payload in the post-capture period.Uncertain inertial dynamic and undesirable rendezvous position can cause the tether swing and oscillation in the orbital plane,which can lead to the tether winding up with payload if unstable.Therefore,analysis of the the equilibrium position and corresponding stabilization control of STS are essential.Recently,numerous studies regarding to dynamic and control of STS for payload capture have been conducted by scholars from all over the world.Hovell et al.[9]compared the ability of different tether structures for space debris removal by air-floating test platform.Through two-dimensional plane ex-periment,it was verified that the sub-tethered structure has better performance of towing space debris.In Refs.[10-11],the swing characteristics and stability of STS in the process of in-plane orbit transfer were studied,and some effective orbit maneuver schemes and swing suppression strategy of the tethered system based on tension and continuous constant thrust were proposed.Although the proposed method can suppress the in-plane motion quickly and accurately,the mothod using external input such as electrodynamic force and thrust is not suitable for stability control in station-keeping stage.Tension control is more suitable to stabilize the tether system in the post-capture stage because the control time is not limited in the long orbit period.

In order to ensure STS fulfill the space missions quickly and stably,scholars have designed a variety of control methods for the release and retrieval process of the tether system.Ref.[12]proposed adaptive sliding mode control for deployment of space tether to overcome the disturbance in loweccentricity orbits.Energy-based control framework was employed into deployment of tethered spacecraft with input saturation[13].Apart from the above methods,several linear or nonlinear methods were also studied for STS,such as incremental nonlinear dynamic surface control,robust performance control,model predictive control,and so on[14-16].It should be noted that most existing studies discussed above depend on the accurate model of STS.However,the system parameters will change abruptly in the period of payload capture,so the accurate dynamic of STS cannot be obtained.Due to highly complex dynamic characteristic and unmeasurable dynamic parameters,designing model-free controller for STS is significant.In recent years,with the development of artificial intelligence,a variety of intelligent optimization algorithms have emerged in solving control issues of aerospace system[17-18].As a representative technology in the field of artificial intelligence,model-free control scheme based on reinforcement learning(RL)has gained wide attention in solving optimization problems with unknown internal dynamics and external disturbances.The optimal controller design of dynamic system can be converted into solving the Hamilton-Jacobi-Bellman(HJB)equations.However,the analytical solutions to the HJB equations are hard to obtain[19].The key superiority of RL is to approximately solve the HJB equations through an iterative method,including policy iteration(PI)and value iteration[20].PI is the most widely technique used in RL to approximate the HJB equations.Generally,PI method has twostep iterations:Policy evaluation and policy improvement.For optimal control problem with parametric uncertainties or even unknown dynamics,the online learning algorithms are of great significance,which can be integrated with adaptive control to develop adaptive optimal control algorithms[21].An online PI algorithm was first presented for optimal control of continuous time system in Ref.[22].Vrabie et al.[23]proposed an integral reinforcement learning(IRL)algorithm for linear continuous-time systems using only partial knowledge about the system dynamics.Furthermore,Ref[.24]presented online model-free RL algorithm for completely unknown continuous-time linear systems.

Based on the above discussion,an online IRL control scheme based on the policy iteration technique is designed to stabilize STS for payload capture with dynamical uncertainty.A state feedback controller for payload capture is developed by using the online information of the system states and inputs without requiring prior knowledge of the system internal dynamics.

1 Problem Formulation

1.1 Dynamic model

STS is considered as an elastic rod model with mass.Some reasonable assumptions are given to simplify the system dynamics modeling as follows.

(1)The space tug(main satellite)and capture mechanism are connected by an elastic tether,and the centroid of the system is on Kepler’s orbit.

(2)The main satellite and sub-satellite in tethered system can be considered as mass points,without regard to the attitude of the satellites.

(3)The tether is regarded as an elastic rod with uniform mass distribution.Only the longitudinal vibration along the tether is considered.

(4)Some space environment effects are ignored,such as solar light pressure,atmospheric resistance,and the oblateness of the earth.

Fig.1 shows the coordinate frames of the space tether.The inertial frameOXYZis attached to the center of earth.TheOXYplane is the same as orbit plane.The axis ofOXpoints to orbital perigee,and theOZaxis is along the equatorial plane normal toward the celestial north pole.TheOYaxis represents the third axis of righthanded orthogonal frame.The orbit coordinate frameCX oY o Zois located at the mass center of STS,withCX oaxis outward from the Earth center along the local vertical.CZ oaxis is directed toward the orbit normal direction,andCY oaxis along the local horizon represents the third axis of right-handed orthogonal frame.The body-fixed frameCX tY t Z thas the same origin as orbit coordinate frame,withCXtaxis along the opposite direction of the tether tension.The direction ofCY tandCZtare determined by the in-plane angleθand out-plane angleφrelative to frameCX oY o Z o.The payload target is in the same orbital plane as the space tether.ηis the true anomaly of target in the inertial frame.

Fig.1 Schematic of capture process by STS

The states of STS can be described by five generalized coordinates:The orbital radiusr,true anomalyϑ,in-plane angleθ,out-plane angleφ,and elastic deformation of tetherε.According to the Euler-Lagrange equation,the differential equations of the system model can be derived as follows

wheremAdenotes the mass of the tug andm Bthe mass of the combination of the capture mechanism and target;mt=ρl0represents the mass of tether,withρthe density of tether andl0the original length;the actual length of tether isl=l0(1+ε).In addition,Tis the tether tension andμthe geocentric gravitation constant.The mass coefficients are defined as

Remar k 1The tether swings around the equilibrium position disturbed by target and the nonnominal libration motion occur in the post-capture stage.The elastic elongation of tether is ignored and the length of tether remains unchanged without control input.The effects from the variation of the tether mass and deformation are ignored.It is assumed that the system moves on the circular orbit,with orbital angular velocityϑ˙=Ω=μ/r3.Based on theabove assumptions,the system model can be linearized near the equilibrium position.

Due to little impact on the in-plane stability,the out-plane motion of the system is ignored.For STS in the circular orbit,the equilibrium positions are along the radial direction of the orbit,with inplane angleθ1,2=0,π.Eq(.1)can be linearized around the equilibrium point by ignoring higher-order terms of nonlinear system.By introducing a dimensionless timeτ=Ωt,the dimensionless linear model of the system can be derived as follows

where the superscripts“'”“''”mean the first and the second derivative versusτ.The dimensionless length is denoted asε0=l/l c-1,andl cis the nominal tether length in the post-capture stage.

Then the state space equation of the system near the equilibrium point can be compactly expressed as

whereAdenotes the state matrix,Bthe input matrix,X=[ε0ε0'θθ′]Tthe state vector,andUthe dimensionless control input.

Remar k 2According to Eq(.3),while adjusting the tether length and velocity,the in-plane angle can be stabilized by the term ofε0′/(ε0+1)in the differential equation.Adjustment of the tether length and velocity can be realized by releasing/rewinding mechanism and tension control.In this paper,tension control scheme is adopted to stabilize the swing motion of tether in the post-capture stage.

1.2 Pr eliminar ies

The aim of this paper is to design an online adaptive control scheme based on policy iteration to drive the real-time asymptotic stability defined dynamic system.In this section,some concepts and propositions for control design are given.

Definition 1(Bellman’s optimality principle[25])Bellman optimality principle is a basic foundation of reinforcement learning.According to the Bellman optimal equation,there exists an optimal control strategy to obtain the optimal cost function of any Markov decision process(MDP).For the linear system,its cost function is the quadratic of the state vector and the control input,so the corresponding optimal feedback control can be derived by solving the basic algebraic Riccatiequation(ARE).

Consider the linear time-invariant(LTI)dynamical system described by

wherex(t)∈Rn,u(t)∈Rm,and the pair(A,B)is controllable,subject to the following optimal control problem

where the infinite horizon quadratic cost function to be minimized is expressed as

withQ≥0,R>0,and(Q1/2,A)detectable.

Based on Bellman optimality principle,the solution of this optimal control problem is obtained byu(t)=-Kx(t)with

where the matrixPis the symmetric positive definite solution of the ARE as Eq.(9).And the unique solution determines the stable close-loop controller.

Lemma 1[26]Consider the linear system expressed as Eq.(5),initializeK0∈Rm×nto be any stabilizing feedback gain matrix,and letP ibe the symmetric positive definite solution of the Lyapunov equation

whereK i=R-1BTP i-1withi=1,2,…,n.Then the following properties are satisfied:(i)(A-BK i)is Hurwitz;

Remar k 3It should be noted that model information of the system is needed to solve the above ARE,which means that the system matrixAand control input matrixBare needed to be known.Therefore,designing a model-free controller with-out using knowledge regarding the system dynamics is particularly important research topic in the optimal control field.

2 Control Design

In this section,an adaptive optimal controller for STS is investigated in the case of unknown internal dynamics.The aim of this paper is to design an online learning control scheme to drive STS asymptotically stable in real time.The proposed IRL policy iteration control diagram is shown as Fig.2,including policy evaluation and policy improvement.Policy evaluation is to calculate the infinite horizon cost associated with the given stability controller,and the purpose of policy improvement is to improve the feedback gain of the system to reduce the cost.

Fig.2 Closed-loop control structure of the system

2.1 IRL policy iteration scheme

LetKbe a stabilizing feedback control gain for Eq(.5).Under the assumption that(A,B)is controllable,the close-loop system is stable with the control inputu(t)=-Kx(t).It directly comes from the state feedback control explanation.

Substituting the expression of control input into Eq.(7),the corresponding infinite horizon quadratic cost becomes

wherePis the real symmetric positive definite solution of the Lyapunov matrix equation

As a Lyapunov function candidate for controlled plant,the cost functionV(x(t))can be written as

Using the conventionalized expressionx(t)=x t,the cost function can be rewritten asV(x t)=,and the initial stable control gain is defined asK1.So we can get the following online policy iteration scheme

Note that Eqs.(14)and(15)form a new policy iteration algorithm without involving the plant matrixA.The whole design procedure of the online control scheme can be summarized as Algorithm 1.The algorithm design of state feedback control based on online IRL can be summarized as the following theorem.

Algorithm 1 Continuous-time IRL policy iteration algorithm

Input:Initial condition of the systemx0,initial stabilizing control gainK1,initialP0=0,initial iteration numberi=1,a positive error constantε,positive definite symmetric matricesQandR.

end while

returnu t,K i+1,P i

Theorem 1For the system model described in Eq(.4),ifK0is initialized to guaranteeA0=ABK0stable,andP iandK iare updated as the policy iteration scheme with proper positive definite symmetric matricesQandR,the closed-loop system is always stable during the iteration period.

ProofSince the positive definite cost functionV i(x t)=is defined as the Lyapunov function candidate and

then for anyδT>0,the unique solution of the Lyapunov equation satisfies

Taking the derivative ofV i(x t)along the state trajectories generated by the control policyut,one obtains

According to Eq(.12),the first term in Eq.(18)can be written as-and the second term can be rewritten using Eq(.15).Then the transformed form of Eq(.18)can be obtained as

With the consideration ofQ>0 andR>0,(x t)<0 is guaranteed when the state vector is nonzero,which proves the updated control policy in Algorithm 1 is stable.The proof of Theorem 1 is completed.

Remar k 4In terms of online implement of IRL algorithm,that is,the on-policy method,the target policy and behavior policy are unified into one policy.Thus,the gain matrix calculated in each iteration is immediately applied to the system,which enhances the running speed of the algorithm.Compared with off-policy method,there is no need to introduce exploration noise into the controller in the proposed algorithm.

2.2 Online implementation of IRL algorithm

In this section,an online adaptive learning algorithm is presented to implement the IRL policy iteration scheme in real time.The algorithm performs the online IRL iterations by measuring the present statex tand the next statex t+Twith fixed sampling timeT.The information of the system matrixAis involved in the measured states,which leads to the policy update without knowing the internal dynamic of the system.

The symmetric matrixP iin the value functionV i(x(t))can be calculated at each iterationiby measuring the states along the system trajectory.For the convenience of computing,the value function is written as

wheredenotes the Kronecker product quadratic polynomial basis vector with the elements{x i(t)x j(t)}i=1,n;j=1,n.The parameter vectorcontains the elements of the matrixP iordered by columns with the redundant elements removed.Then Eq(.14)can be rewritten as

whereis the vector of unknown parameters and((t)-(t+T))acts as a regression vector.The right hand side is the integral reinforcement on the time interval[t,t+T],which can be denoted as

whered((t),K i)represents a desired value or target function and the estimate of the parameteris to be found such that the parameter satisfies the equation as closely as possible.To compute it efficiently,define a new controller stateV(t)as the augmented state of the system,and add the state equation˙(t)=xT(t)Qx(t)+uT(t)Ru(t)to the controller dynamics.The value ofd((t),K i)can be computed by usingd((t),K i)=V(t+T)-V(t).

The unknown parameter vectorof the value function is involved in the scalar equation Eq.(21),which can be solved by a batch solution method in the least-squares sense.Firstly,some relevant vectors of parameters are defined as

Assume the square loss function asTo minimize the square loss,let the derivative ofJ()with respect toequals to zero,one can obtain

Noting thatis a scalar,the left hand side of Eq.(27)has the following form after rearrangement

Using the above transformation equation,Eq(.23)can be rewritten as

Then the batch least-squares solution ofis obtained in the matrix form

Until now,the least-squares problem can be solved online with a sufficient number of data collected along the state trajectory.According to Lemma 2,the convergence of the online adaptive IRL algorithm can be guaranteed in finite iteration steps.

Remar k 5The developed adaptive optimal control is a type of data-driven method,where the system matrix is not needed.In fact,the algorithm can be also employed into time-varying system.If matrixAof the system changes suddenly,as long as the current controller of the new matrix is stable,the algorithm converges to the corresponding solution of the new ARE.

3 Numerical Simulation

In this section,numerical simulations are conducted to validate the performance of the proposed online adaptive IRL control scheme for stabilizing the swing motion of STS after capturing payload.Some parameters of the system are provided in Table 1.The desired dimensionless state of STS isDue to the impact of payload,the tether swings up to certain libration angle with varying length,the initial state of STS is assumed as

Table 1 Specific parameters of the tethered system

In view of the unknown mass parameter of payload in general cases,the the initial parameters of the controller is deduced based on the linearized model of STS before capturing payload,which can guarantee the initial stability of controller.Then the feedback gain will be updated to satisfy the optimal control of STS in the post-capture period until convergence.It should be noted that the designed controller is applied into the original nonlinear plant.Since the onboard computational source limited,the sampling interval is set asτ=0.05.

In addition to parameters of mission scenario,the controller parameters are well selected to make sure the convergence of algorithm and control performance.It is reasonable that the better control performance of the in-plane libration angle can be obtained by selecting larger corresponding weights in cost function.Therefore,the symmetric weight matrices are chosen asQ=diag(10 2 1 1)andR=1.The parameterN Nrepresents that an iterative up-date is performed afterN Nsampling time steps.HereN Nis set to 20,which means that the system iterates once everyN N·τ=20×0.05=1(rad)to update the control gainKand critic parameter matrixP.The parameterε=10-3denotes the error threshold,representing that the iteration process will stop when the cost functionV tsatisfies the requirement.

The simulation results are shown in Figs.3—8.Fig.3 shows the evolution of the parameters of matrixPin the Riccati equation.The matrixPconverges to a constant optimal value after four iterations,which means that the online learning process is completed and the final control gain is determined after around 4 rad.The error norms of critic matrix and gain matrix in the policy iteration are presented in Fig.4.It can be seen that the matrixKwill be close to the ideal optimal control gain after several policy updates.From Fig.5,the dimensionless states of STS converge from the initial position in the post-capture phase to zeros within 15 rad.The maximum libration angle is no more than the initial value,and the amplitude of tether oscillation is less than its nominal length,avoiding the risk of collision between the tether and the satellite.

Fig.3 Critic matrix P of controller

Fig.4 Error norms of matrices P and K

Fig.5 Dimensionless state variables under IRL control

The variation curve of tether tension in Fig.6 indicates that the tether always keeps tense in the control stage,and magnitude of the tension meets the physical characteristics of the tether.Fig.7 shows the evolution of the total cost function(augmented state)and the integral reinforcement signal(one time-step cost)in the optimal control.It can be seen that the cost functionV(t)increases gradually over time and ultimately converges to a positive constant.On the contrary,the integral reinforcement has the trend of decreasing gradually,even though transient rise occurs in the online learning stage due to large deviation between initial control gain matrix and desired one.

Fig.6 Variation curve of tether tension in the control process

Fig.7 Cost function and integral reinforcement of the controller

In order to illustrate the robustness performance of the proposed algorithm,we compared it with the classic LQR controller by simulations under the presence of stochastic disturbance.For our simulations,the Gaussian noise with mean value of 0 and variance of 1 is adopted as stochastic disturbance,and the control parametersQandRare chosen to be the same for both controllers.The gain matrix of LQR controller is deduced by the known dynamic model of STS before payload capture.Simulation results are shown in Fig.8.Both methods can ensure that the tether length and libration angle converge to desired values,while the method based on policy iteration has better convergence speed and control performance.

Fig.8 Comparison between policy iteration and LQR

4 Conclusions

An adaptive optimal controller based on IRL policy iteration is conducted to address stabilization control of the tether libration after capturing the payload by STS.Due to lack of accurate dynamic model of the system in the post-capture stage,the classic model-based control methods will result in poor control effect.The proposed algorithm can achieve continuous time optimal control without accurately understanding the internal dynamics of the system,thus effectively solving the libration control problem of STS.Firstly,the basic dynamic model of STS is derived considering tether elasticity.Then the policy iteration based IRL algorithm is designed and the batch solution method is proposed for online implementing the algorithm.Finally,the effectiveness of the proposed control scheme is validated by the numerical simulation.The drawback is that the proposed method relies on linear dynamic system,causing poor control performance or even instability for the case of large libration amplitude.Our future work will focus on developing model-free adaptive optimal control scheme that can be directly applied into nonlinear dynamic system.

AcknowledgementsThiswork wassupported by the National Natural Science Foundation of China(No.62111530051),the Fundamental Research Funds for the Central Universities(No.3102017JC06002)and the Shaanxi Science and Technology Program,China(No.2017KW-ZD-04).

AuthorsMr.FENG Yiting received the M.S.degree in control science and engineering from Northwestern Polytechnical University,China,in 2020.He is currently pursuing the Ph.D.degree in control science and engineering at Northwestern Polytechnical University,China.His current research interests include nonlinear control,adaptive control,intelligent control and space tether system dynamic analysis and control.Prof.WANG Changqing is currently a professor with School of Automation,Northwestern Polytechnical University.He received the B.S.degree in mechanical design and manufacturing and the M.S.degree in navigation guidance and control from Northwestern Polytechnical University,Xi’an,China,in 1996 and 2001.He received the Ph.D.degree in system analysis,control and information processing from National Research University Moscow Power Engineering Institute,Russia,in 2006.His current research interests include adaptive control,and space tether system dynamic analysis and control.

Author contributionsMr.FENG Yiting designed the control algorithm,contributed to the simulation and the analysis of the study and wrote the manuscript.Mr.ZHANG Ming contributed to the simulation and the analyzation of the study.Dr.GUO Wenhao contributed to the data for the simulation and validation of model.Prof.WANG Changqing contributed to the research approach,background of the study and interpreted the results.All authors commented on the manuscript draft and approved the submission.

Competing interestsThe authors declare no competing interests.


登录APP查看全文