Policy gradient methods
Learning objective
Section titled “Learning objective”Optimise the policy directly instead of deriving it from value estimates.
Planned scope
Section titled “Planned scope”- Policy parameterisation and the softmax policy
- The policy gradient theorem
- REINFORCE and its variance problem
- Baselines and advantage
- Actor-critic methods
Prerequisites
Section titled “Prerequisites”Everything above this chapter in the sidebar.