Choices, Games & Branching Paths

Learning What to Do in Each Situation

Also called: Markov decision process

  • Personal interest
  • Formal theory
  • Personal metaphor

In one kind of decision model, the current situation, the actions available and the rewards that follow together shape a policy: a rule for what to do in each situation. My personal comparison is that each lesson I learn updates how I choose in the next situation.

Which of my current rules for choosing came from an old lesson I should revisit?

Why it attracts me

I like models that are honest about what they assume. A Markov decision process is one of the cleanest models of choosing over time, and it gives words to something I feel: each lesson changes how I choose next time.

The idea

The model has four parts. There is the situation I am in, which the model calls the state. There are the actions I can take. Each action brings a reward or a cost and moves me to a new situation, partly by chance (Probability, Risk & Uncertainty). And there is a policy, a rule that says what to do in each situation. The model assumes the current situation carries everything that matters for what happens next. Methods called reinforcement learning improve a policy from experience: try, see what happens, adjust. A central puzzle there is when to try something new and when to use what already works (Try Something New or Use What Works).

What I think (and don't know)

The comparison to a life is striking. Habits act like a policy learned over years (Small Habits Shape Who I Become). A lesson from one experience shifts what I do the next time something similar happens. Thinking this way helps me ask a useful question: which of my rules did I choose on purpose, and which did I simply absorb?

I hold the comparison loosely. A person is not literally this kind of model. My present situation never fully captures my past, the rewards I care about change, and purpose cannot be reduced to a single score (AI Workflows That Catch Their Own Mistakes). The model is a mirror for reflection, not a description of a mind.

I also use the image in a more personal way. In my own philosophy, I sometimes picture a soul as a traveler that learns from each experience (A Soul That Learns Along the Way). The policy-updating picture gives that belief a shape. It does not give it evidence.

Where it connects

This idea sits between the branching map of choices (Mapping the Futures a Choice Opens) and the larger picture of life as a game that changes its player (Life as an Evolving Game). It also links to the AI agents I am interested in, which revise their plans as results come in (AI Workflows That Catch Their Own Mistakes).

An example

My Steady side quest is a small practice tool built on this idea. It has 33 skills. You pick a setting (Home, Work or Self), rehearse a response, and save a "When this happens, I will do this" plan. Then you try it, reflect on how it went, and the app brings it back for recall after 1, 3, 7, 14 and 30 days. In policy terms, each plan is a rule written for one situation, and practice is what makes that rule come to mind when the situation really happens. Whether it helps is the experiment.

Questions I am still carrying

  • Which of my rules for choosing were learned in situations that no longer exist?
  • What reward am I really following, and is it the one I would choose?
  • How do I learn from a bad outcome without letting it rewrite the whole rule?

What this does not establish

Comparing a life to a decision model does not show that people are reward-maximizing machines, or that a soul learns across experiences. The model is a useful mirror for habits and choices, not a description of what a person is.

Questions I'm still exploring

  • Which of my rules for choosing were learned in situations that no longer exist?
  • What reward am I really following, and is it the one I would choose?
  • How do I learn from a bad outcome without letting it rewrite the whole rule?

Sources and further reading

  • Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd edition (MIT Press, 2018)
  • Richard Bellman, Dynamic Programming (Princeton University Press, 1957)