Skip to content

Advanced policy optimization

Make policy gradient stable enough to trust: constrained updates and modern actor-critic.

  • A2C and A3C
  • Why unconstrained policy updates collapse
  • Trust regions and TRPO
  • PPO and the clipped objective
  • Generalised advantage estimation

Everything above this chapter in the sidebar.