Policy Update Magnitude is a measure of how much a reinforcement learning agent’s policy changes between successive gradient update steps, typically quantified as the KL divergence between the old and new policy distributions or as the Euclidean norm of the parameter change vector. Controlling this magnitude is essential to training stability: excessively large updates can cause catastrophic performance collapse, whilst excessively small updates slow convergence. Algorithms such as Proximal Policy Optimisation (PPO) and Trust Region Policy Optimisation (TRPO) impose explicit constraints on policy update magnitude to balance exploration, exploitation, and stability.

Policy Update Magnitude measures how much a reinforcement learning agent’s policy changes between successive gradient update steps, quantified as KL divergence between old and new policy distributions or as the norm of the parameter change vector. Constraining this magnitude balances training stability against convergence speed.

Content

In policy gradient methods, each parameter update step shifts the agent’s behaviour by an amount determined by the gradient magnitude and the learning rate. If this shift is too large, the new policy may land in a region of the parameter space where the gradient estimate (computed under the old policy) is no longer valid, causing destructive interference between the collected experience and the updated behaviour—a phenomenon sometimes called “catastrophic forgetting of recent experience.” Trust Region Policy Optimisation (TRPO) formalises this concern by solving a constrained optimisation problem: maximise the surrogate objective subject to the KL divergence between old and new policies being no greater than a hyperparameter δ. This constraint defines a “trust region” within which the first-order approximation of the objective is reliable.

Proximal Policy Optimisation (PPO) achieves a similar constraint more efficiently by clipping the probability ratio r(θ) = π_θ(a|s) / π_θ_old(a|s) to the interval [1−ε, 1+ε], where ε is typically 0.1–0.2. The clipped surrogate objective removes the gradient contribution from updates that would move the policy outside the allowed range, effectively penalising large updates without the computational overhead of the TRPO constrained optimiser. PPO has become the dominant algorithm for fine-tuning large language models via Reinforcement Learning from Human Feedback precisely because its update-magnitude control produces stable training at scale.

Monitoring policy update magnitude during training—by logging KL divergence or parameter norms per update step—is a standard diagnostic practice. A diverging KL divergence signals that the learning rate or clip parameter is too large and should be reduced. Conversely, a KL divergence that never exceeds a tiny fraction of the constraint suggests the policy is barely changing, indicating excessive conservatism and a need to increase the learning rate or reward signal scale.

Provenance