This paper investigates the guidance method based on reinforcement learning (RL) for the coplanar orbital interception in a continuous low-thrust scenario. The problem is formulated into a Markov decision process (MDP) model, then a welldesigned RL algorithm, experience based deep deterministic policy gradient (EBDDPG), is proposed to solve it. By taking the advantage of prior information generated through the optimal control model, the proposed algorithm not only resolves the convergence problem of the common RL algorithm, but also successfully trains an efficient deep neural network (DNN) controller for the chaser spacecraft to generate the control sequence. Numerical simulation results show that the proposed algorithm is feasible and the trained DNN controller significantly improves the efficiency over traditional optimization methods by roughly two orders of magnitude.
scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.