Unified Notation for Episodic and Continuing Tasks
Episodic and continuing tasks can use the same return formula. The trick is to imagine that, after an episode ends, the process remains in a terminal state and receives zero rewards forever.
An absorbing terminal state
Section titled “An absorbing terminal state”A terminal state is absorbing when its next state is itself with probability one. After termination, we can imagine a dummy action that keeps the process there and produces reward 0. The real agent does not need to keep acting; this is a mathematical convention.
flowchart LR
accTitle: An absorbing terminal state
accDescr: Entering Goal gives plus five once. All imagined transitions afterward stay at Goal and give zero reward.
hall["Hall"] -->|"Right; reward +5"| goal["Goal: terminal"]
goal -->|"Remain here; reward 0"| goal
If the gridworld episode ends at , its reward sequence can be written as:
The goal reward is paid on entry. It is not paid again by the terminal self-loop.
One formula for both task types
Section titled “One formula for both task types”For an episode ending at , the finite return is:
Once rewards after are defined to be zero, this is the same as:
A continuing task uses the same infinite sum without an artificial terminal state. Its return still needs to be well-defined; with bounded rewards, choosing is sufficient.
| Quantity | Episodic task | Continuing task |
|---|---|---|
| End time | A terminal time exists for an episode that finishes. | No natural terminal time. |
| Rewards after termination | Set to zero by convention. | Rewards continue according to the environment. |
| Discounting | May use when returns remain well-defined. | Usually use for discounted-return objectives. |
| Return recursion | The same recursion. |
Let denote nonterminal states and include terminal states as well. When summing over possible next states, include the terminal states. Their future value is zero.
Why terminal value is zero
Section titled “Why terminal value is zero”At the end of an episode, because no rewards remain. Consequently, the terminal state’s value is also zero.
For the policy that chooses right from Hall:
The value functions are defined formally in the next section. Here the important point is that reward on arrival and value after arrival are different quantities.
Pole balancing: where the episode boundary matters
Section titled “Pole balancing: where the episode boundary matters”Suppose reward is 0 while the pole remains upright and −1 when it falls. If one episode ends with failure at time , then for :
With , delaying failure makes its negative contribution closer to zero and therefore improves the return. With , every trajectory ending in one failure gives return −1, regardless of its duration. That reward rule alone would not distinguish short and long balancing episodes; giving +1 per surviving step would define a different objective.
In a continuing formulation, assume each failure is followed by a reset, resets yield no additional reward, and future failure times are , all greater than the current time . Then:
The episodic return counts this episode’s failure. The continuing return counts all future failures, including those after resets. Choosing the episode boundary therefore affects what the agent is optimizing.
Reference
Section titled “Reference”Sutton, R. S. & Barto, A. G. Reinforcement Learning: An Introduction, 2nd ed., chapter 3, section 3.4. The explanations and small worked examples here are written for these notes.
Previous: 3.3 Returns and Episodes · Next: 3.5 Policies and Value Functions