Skip to content

Unified Notation for Episodic and Continuing Tasks

Episodic and continuing tasks can use the same return formula. The trick is to imagine that, after an episode ends, the process remains in a terminal state and receives zero rewards forever.

A terminal state is absorbing when its next state is itself with probability one. After termination, we can imagine a dummy action that keeps the process there and produces reward 0. The real agent does not need to keep acting; this is a mathematical convention.

flowchart LR
    accTitle: An absorbing terminal state
    accDescr: Entering Goal gives plus five once. All imagined transitions afterward stay at Goal and give zero reward.
    hall["Hall"] -->|"Right; reward +5"| goal["Goal: terminal"]
    goal -->|"Remain here; reward 0"| goal

If the gridworld episode ends at T=2T=2, its reward sequence can be written as:

R1=−1,R2=5,R3=R4=⋯=0.R_1=-1,\quad R_2=5,\quad R_3=R_4=\cdots=0.

The goal reward is paid on entry. It is not paid again by the terminal self-loop.

For an episode ending at TT, the finite return is:

Gt=∑k=0T−t−1γkRt+k+1.G_t = \sum_{k=0}^{T-t-1} \gamma^k R_{t+k+1}.

Once rewards after TT are defined to be zero, this is the same as:

Gt=∑k=0∞γkRt+k+1.G_t = \sum_{k=0}^{\infty} \gamma^k R_{t+k+1}.

A continuing task uses the same infinite sum without an artificial terminal state. Its return still needs to be well-defined; with bounded rewards, choosing γ<1\gamma<1 is sufficient.

QuantityEpisodic taskContinuing task
End time TTA terminal time exists for an episode that finishes.No natural terminal time.
Rewards after terminationSet to zero by convention.Rewards continue according to the environment.
DiscountingMay use γ=1\gamma=1 when returns remain well-defined.Usually use γ<1\gamma<1 for discounted-return objectives.
Return recursionGt=Rt+1+γGt+1G_t=R_{t+1}+\gamma G_{t+1}The same recursion.

Let S\mathcal{S} denote nonterminal states and S+\mathcal{S}^{+} include terminal states as well. When summing over possible next states, include the terminal states. Their future value is zero.

At the end of an episode, GT=0G_T=0 because no rewards remain. Consequently, the terminal state’s value is also zero.

For the policy that chooses right from Hall:

vπ(Hall)=5+γvπ(Goal)=5+γ(0)=5.v_\pi(\text{Hall}) = 5 + \gamma v_\pi(\text{Goal}) = 5 + \gamma(0) = 5.

The value functions are defined formally in the next section. Here the important point is that reward on arrival and value after arrival are different quantities.

Pole balancing: where the episode boundary matters

Section titled “Pole balancing: where the episode boundary matters”

Suppose reward is 0 while the pole remains upright and −1 when it falls. If one episode ends with failure at time TT, then for t<Tt<T:

Gt=−γT−t−1.G_t = -\gamma^{T-t-1}.

With 0<γ<10<\gamma<1, delaying failure makes its negative contribution closer to zero and therefore improves the return. With γ=1\gamma=1, every trajectory ending in one failure gives return −1, regardless of its duration. That reward rule alone would not distinguish short and long balancing episodes; giving +1 per surviving step would define a different objective.

In a continuing formulation, assume each failure is followed by a reset, resets yield no additional reward, and future failure times are T1,T2,…T_1,T_2,\ldots, all greater than the current time tt. Then:

Gt=−∑i=1∞γTi−t−1.G_t = -\sum_{i=1}^{\infty}\gamma^{T_i-t-1}.

The episodic return counts this episode’s failure. The continuing return counts all future failures, including those after resets. Choosing the episode boundary therefore affects what the agent is optimizing.

Sutton, R. S. & Barto, A. G. Reinforcement Learning: An Introduction, 2nd ed., chapter 3, section 3.4. The explanations and small worked examples here are written for these notes.


Previous: 3.3 Returns and Episodes · Next: 3.5 Policies and Value Functions