Run a simplified value iteration. The policy is fixed, so we know what action to do in each state. Repeat the following a fixed number of times: