Skip to content

The Agent–Environment Interface

An agent makes decisions; its environment responds. A Markov decision process (MDP) describes this interaction when the current state provides enough information to predict what happens next, given the action.

At each time step tt:

  1. The agent observes the current state StS_t.
  2. It chooses an action AtA_t.
  3. The environment produces a reward Rt+1R_{t+1} and a next state St+1S_{t+1}.
  4. The agent repeats from the new state, unless the episode has ended.
flowchart LR
    accTitle: Agent–environment interaction
    accDescr: The agent sends an action to the environment, which returns a reward and the next state.
    agent["Agent"] -->|"Action A_t"| environment["Environment"]
    environment -->|"Reward R_(t+1) and state S_(t+1)"| agent

The reward is labeled Rt+1R_{t+1} because it arrives after action AtA_t. The resulting sequence is a trajectory:

S0,A0,R1,S1,A1,R2,S2,…S_0, A_0, R_1, S_1, A_1, R_2, S_2, \ldots

The agent–environment boundary describes what the decision-maker controls. For a robot, choosing a motor command is an action; the resulting motion is part of the environment’s response. The environment can include parts of the robot’s own body.

Imagine a robot moving through this grid:

Column 1Column 2Column 3
Top rowStartHallGoal
Bottom rowBCD

The state is the robot’s current cell. Its actions are up, down, left, and right. A move deterministically reaches the neighboring cell in that direction; attempting to cross the outer wall leaves the robot in place.

Entering Goal gives +5 and ends the episode. Every other move gives −1, including attempts to cross a wall. The final +5 replaces the usual −1; it is not +4.

For example, moving right twice produces:

flowchart LR
    accTitle: A two-step gridworld route
    accDescr: Moving right from Start reaches Hall with reward minus one. Moving right again reaches Goal with reward plus five and ends the episode.
    start["Start"] -->|"Right; reward −1"| hall["Hall"]
    hall -->|"Right; reward +5"| goal["Goal: episode ends"]
ConceptMeaningGridworld example
State ssCurrent situationThe robot is in Hall.
Action aaChoice available nowMove right.
TransitionNext state after an actionHall → Goal.
Reward rrNumerical feedback after an action+5 on entering Goal.
Policy π\piRule for choosing actionsChoose right in Start and Hall.
Return GtG_tTotal future reward, possibly discountedThe score for the rest of the route.
ValueExpected return under a policyHow good it is to start in Hall.

An MDP specifies the states S\mathcal{S}, available actions A(s)\mathcal{A}(s), and environment dynamics. A discounted-return objective also specifies a discount factor γ\gamma. The policy describes the agent’s behavior within this problem; changing the policy does not change the environment’s rules.

For a finite MDP, the joint model is:

p(s′,r∣s,a)=Pr⁡(St+1=s′,Rt+1=r∣St=s,At=a).p(s',r \mid s,a) = \Pr(S_{t+1}=s', R_{t+1}=r \mid S_t=s, A_t=a).

Read it as: “Given this state and action, how likely are this next state and reward?” For each valid state–action pair, the probabilities over all possible outcomes sum to one:

∑s′,rp(s′,r∣s,a)=1.\sum_{s',r} p(s',r \mid s,a) = 1.

In our gridworld, p(Goal,5∣Hall,right)=1p(\text{Goal},5 \mid \text{Hall},\text{right})=1. A slippery version could move right with probability 0.8 and stay put with probability 0.2. The agent chooses an action; it does not choose which random outcome occurs.

Sutton & Barto’s Example 3.3 gives a p(s′,r∣s,a)p(s',r \mid s,a) table with real randomness in it. A robot collects cans and has a battery that is either HIGH or LOW. It can SEARCH for cans, WAIT in place, or RECHARGE. Searching pays the most but risks draining the battery; recharging is free but forfeits a turn.

StateActionNext stateProbabilityReward
HIGHSEARCHHIGHα\alpharsearchr_{\text{search}}
HIGHSEARCHLOW1−α1-\alpharsearchr_{\text{search}}
HIGHWAITHIGH1rwaitr_{\text{wait}}
LOWSEARCHLOWβ\betarsearchr_{\text{search}}
LOWSEARCHHIGH1−β1-\beta−3-3 (rescued)
LOWWAITLOW1rwaitr_{\text{wait}}
LOWRECHARGEHIGH10

This is a genuinely stochastic, continuing task (no terminal state), which makes it a better test case than the gridworld for checking that you can read a dynamics table rather than just a deterministic diagram.

toy-recycle-bot.py in this chapter’s folder implements it as a Gymnasium environment. It checks its own deterministic transitions with assertions before running a random-policy demo, so python toy-recycle-bot.py both verifies and demonstrates the dynamics above.

The current state contains the information needed to predict the next state and reward, given the action. Earlier history adds nothing to that prediction:

Pr⁡(St+1=s′,Rt+1=r∣St,At,earlier history)=Pr⁡(St+1=s′,Rt+1=r∣St,At).\Pr(S_{t+1}=s', R_{t+1}=r \mid S_t,A_t,\text{earlier history}) = \Pr(S_{t+1}=s', R_{t+1}=r \mid S_t,A_t).

Location is enough for our gridworld. But if entering Goal requires a key, “Hall with a key” and “Hall without a key” have different possible outcomes. The state must then include both location and whether the robot has the key. Battery level or time remaining may also need to be included when they affect the rules.

Markov does not mean deterministic: the slippery robot can still be Markov if its outcome probabilities depend only on its current state and action.

What if the agent cannot observe the full state?

A partially observable MDP (POMDP) distinguishes the true state StS_t from the observation OtO_t available to the agent. For example, a robot’s camera may not reveal what is behind a wall.

One approach maintains a belief state, a probability distribution over possible states based on the action and observation history HtH_t:

bt(s)=Pr⁡(St=s∣Ht).b_t(s) = \Pr(S_t=s \mid H_t).

This chapter uses fully observed states. The distinction matters when deciding whether the information available to an agent really is a Markov state.

Sutton, R. S. & Barto, A. G. Reinforcement Learning: An Introduction, 2nd ed., chapter 3, section 3.1. The explanations and small worked examples here are written for these notes.


Chapter overview · Next: 3.2 Goals and Rewards