Policy optimisation is the process of searching over a parameterised space of decision-making policies to maximise a scalar objective, typically a cumulative reward signal in a reinforcement learning context. It encompasses both gradient-based methods such as proximal policy optimisation and gradient-free approaches such as evolutionary strategies, applied across discrete and continuous action spaces.

Content

  • Policy optimisation emerged as a formal discipline from the confluence of optimal control theory and statistical machine learning in the 1990s. Early work by Williams (1992) on REINFORCE established the policy-gradient theorem, showing that the gradient of expected return with respect to policy parameters could be estimated from sampled trajectories without knowledge of the environment model. Subsequent research by Sutton et al. (2000) generalised this to compatible function approximation, laying the groundwork for modern actor-critic architectures.
  • Technically, policy optimisation methods can be categorised along two axes: on-policy versus off-policy data usage, and first-order versus zeroth-order gradient estimation. On-policy methods such as Trust Region Policy Optimisation (TRPO) and its successor PPO collect fresh trajectories at each update step, computing policy gradients constrained by KL-divergence bounds to prevent destructive parameter updates. Off-policy methods such as Soft Actor-Critic (SAC) recycle past experience from a replay buffer, improving sample efficiency at the cost of higher variance. Zeroth-order evolutionary strategies treat the policy as a black box and estimate gradients through finite differences across a population of parameter perturbations.
  • The policy optimisation ecosystem spans simulation engines such as Gazebo Simulator, robotics middleware, GPU-accelerated environments, and distributed training frameworks. Libraries such as Stable-Baselines3, RLlib, and CleanRL provide reference implementations of canonical algorithms, while benchmarks such as MuJoCo locomotion tasks, Atari ALE, and robotics manipulation suites serve as standardised evaluation arenas. Reward shaping, curriculum design, and hierarchical decomposition are employed to tame sparse-reward environments where naive policy gradients fail to propagate useful learning signals.
  • In 2024–2025, policy optimisation has moved beyond game-playing benchmarks into industrial deployment: large language model alignment via Reinforcement Learning from Human Feedback relies on reward-model-guided policy optimisation, while robot foundation models trained with diffusion policies and flow-matching objectives challenge the dominance of classical gradient-based approaches. Multi-task and meta-learning formulations allow a single policy to generalise across families of tasks, and offline-to-online fine-tuning pipelines enable safe deployment by first imitating logged behaviour before online policy improvement.