Skip to content

Policy gradient methods

Optimise the policy directly instead of deriving it from value estimates.

  • Policy parameterisation and the softmax policy
  • The policy gradient theorem
  • REINFORCE and its variance problem
  • Baselines and advantage
  • Actor-critic methods

Everything above this chapter in the sidebar.