Advanced policy optimization
Learning objective
Section titled “Learning objective”Make policy gradient stable enough to trust: constrained updates and modern actor-critic.
Planned scope
Section titled “Planned scope”- A2C and A3C
- Why unconstrained policy updates collapse
- Trust regions and TRPO
- PPO and the clipped objective
- Generalised advantage estimation
Prerequisites
Section titled “Prerequisites”Everything above this chapter in the sidebar.