We can compute the average reward per time step. Even for an infinite policy, this will usually be finite.