Posts

Showing posts from September, 2020

Homo Deus

The value-based and policy-based distinction isn't one I'm familiar with, but it makes total sense to me. Two very meaningful ways of assessing progress. I think my usual mindset aligns more closely with policy-based, but with some distinctions. My goal is usually to push the limits of my potential. I don't really care about comparisons to others (and when I do find myself caring, I try to stop). I want to be better than I was yesterday or last year, which I think aligns to your policy-based, but my mental model is that I have a certain amount of potential inherent, and I try to use it and develop it as much as possible.  In other words, I think I blur the two together a bit. My 'value-based' is about achieving my maximum potential: if I'm at 100% of what I could possibly achieve, that's a win. It doesn't really matter whether my 100% is more or less than someone else's. Of course, the problem is I have no real way to know what my potential is, and s...

Ex Machina

When you are training a reinforcement learning system, there are two broad approaches: value based and policy based. For value based, you try to get the system to understand the value of each action in each state so that it has a running estimate of "how well it's doing". For policy based, you simply ask the machine at each point in time how it could do better, but don't expect it to be able to evaluate how well it's doing in any absolute sense. For the past few days, I've been debating which of these two methods I should adopt at present. I've been feeling a bit lost from a value perspective, I don't really know how well I'm doing in any absolute sense. However, I can see before me lots of directions for improvement. It's tempting therefore to take a policy based approach: "I don't know where I am, but I do know how to do a bit better". (Parenthetically, let's expand on the specific strategies we each use for this type of thi...