基于强化学习的多智能加速器控制仿真系统

Multi-Intelligent accelerator control simulation system based on reinforcement learning

  • 摘要: 加速器驱动嬗变研究装置(CiADS)是我国核废料处理领域的重大科研基础设施,其稳定运行依赖于对束流传输状态的精确调控。针对传统调束方法在动态适应性与智能化方面的不足,探索了一种数据驱动的多智能体加速器控制方案。建了基于级联BP神经网络的中能束流传输段数字仿真环境,以精确模拟束流传输动力学。在此基础上,基于MASAC算法设计了一套多智能体强化学习控制系统,用于实现束流轨迹的在线自适应调校与优化。实验结果表明,该强化学习模型能通过自主学习有效优化控制策略,在50组测试数据上对束流位置的平均调校增益达到12.54%,使束流轨迹显著向传输中心集中。本研究为大型加速器装置的智能化运行控制提供了一种兼具高精度、强实时性与良好鲁棒性的新方法。

     

    Abstract:
    Background The China initiative Accelerator Driven System (CiADS) is a large-scale scientific facility for nuclear waste transmutation, whose stable operation requires accurate control of beam transport trajectories. Conventional orbit correction methods are generally effective under approximately linear and stable operating conditions, but their adaptability may deteriorate in the presence of nonlinear responses, coupled actuators, disturbances, and varying operating conditions. Therefore, data-driven intelligent control provides a promising approach for improving the autonomous beam tuning capability of accelerator systems.
    Purpose This study aims to develop a multi-agent reinforcement learning control method for adaptive beam trajectory correction in the Medium Energy Beam Transport (MEBT) section of CiADS and to evaluate its feasibility, correction performance, robustness, and real-time implementation capability.
    Methods A digital simulation environment was first constructed using a cascaded backpropagation neural network (CBPNN), in which four BP neural networks were sequentially connected according to the beam transport process to reproduce the nonlinear mapping between steering-magnet currents and beam position monitor (BPM) readings. On this basis, a customized multi-agent soft actor–critic (MASAC) controller was developed according to the physical topology and steering-magnet–BPM response relationships of the MEBT section. The global trajectory correction task was decomposed into several local control tasks, and individual agents generated continuous steering-magnet current adjustments from local beam states. The trained controller was evaluated using 50 previously unseen initial conditions and compared with response-matrix/SVD, single-agent SAC, and TD3 methods.
    Results The MASAC controller reduced the mean beam offset at BPM5 from 0.8560 mm to 0.7487 mm, corresponding to an average correction gain of 12.54%. It achieved higher correction performance than SVD (5.11%), single-agent SAC (8.54%), and TD3 (10.07%), and converged within approximately 190 training episodes. Under disturbance conditions, MASAC maintained an average correction gain of 9.86%. FPGA-based implementation further reduced the single-step inference latency from 419 μs on the CPU to 55 μs, while achieving a power consumption of 1.572 W.
    Conclusions The proposed CBPNN–MASAC framework effectively integrates data-driven beam transport modeling with topology-aware multi-agent reinforcement learning. It improves beam trajectory correction, convergence, and robustness in the MEBT simulation environment and demonstrates the potential for low-latency deployment in future online intelligent accelerator control systems.

     

/

返回文章
返回