Background The China initiative Accelerator Driven System (CiADS) is a large-scale scientific facility for nuclear waste transmutation, whose stable operation requires accurate control of beam transport trajectories. Conventional orbit correction methods are generally effective under approximately linear and stable operating conditions, but their adaptability may deteriorate in the presence of nonlinear responses, coupled actuators, disturbances, and varying operating conditions. Therefore, data-driven intelligent control provides a promising approach for improving the autonomous beam tuning capability of accelerator systems.
Purpose This study aims to develop a multi-agent reinforcement learning control method for adaptive beam trajectory correction in the Medium Energy Beam Transport (MEBT) section of CiADS and to evaluate its feasibility, correction performance, robustness, and real-time implementation capability.
Methods A digital simulation environment was first constructed using a cascaded backpropagation neural network (CBPNN), in which four BP neural networks were sequentially connected according to the beam transport process to reproduce the nonlinear mapping between steering-magnet currents and beam position monitor (BPM) readings. On this basis, a customized multi-agent soft actor–critic (MASAC) controller was developed according to the physical topology and steering-magnet–BPM response relationships of the MEBT section. The global trajectory correction task was decomposed into several local control tasks, and individual agents generated continuous steering-magnet current adjustments from local beam states. The trained controller was evaluated using 50 previously unseen initial conditions and compared with response-matrix/SVD, single-agent SAC, and TD3 methods.
Results The MASAC controller reduced the mean beam offset at BPM5 from 0.8560 mm to 0.7487 mm, corresponding to an average correction gain of 12.54%. It achieved higher correction performance than SVD (5.11%), single-agent SAC (8.54%), and TD3 (10.07%), and converged within approximately 190 training episodes. Under disturbance conditions, MASAC maintained an average correction gain of 9.86%. FPGA-based implementation further reduced the single-step inference latency from 419 μs on the CPU to 55 μs, while achieving a power consumption of 1.572 W.
Conclusions The proposed CBPNN–MASAC framework effectively integrates data-driven beam transport modeling with topology-aware multi-agent reinforcement learning. It improves beam trajectory correction, convergence, and robustness in the MEBT simulation environment and demonstrates the potential for low-latency deployment in future online intelligent accelerator control systems.