
Very few will care about the actual reinforcement learning details here, but watching me grope around for understanding in the chat log might be of broader interest:
I have often noticed that our estimated Q values are higher than the observed returns, sometimes substantially so, and that can’t be good for performance.
For policy decisions, only the relative values matter, but an offset negatively impacts bootstrapping, so I was excited to see this paper on Relative Value Learning:
arxiv.org/pdf/1901.09732
After struggling a bit with “the Banach space of bounded antisymmetric pairwise functions“, I realized all that is essentially going on is subtracting two value functions. Given this framing, Chat was able to simplify the machinery into a particularly elegant form:
Subtracting the mean TD error from the individual sample TD errors gives all the benefits of relative value learning. Summarized as "relative Bellman regression is TD-error centering".
Unfortunately, while you can create examples where this should be very valuable (someone should write a proper paper on it!), it didn’t actually improve performance on my tasks.
However, I think this has usefully narrowed down what is actually happening. Bellman iteration naturally corrects towards a correct absolute value, but only when the bootstrap values are also training targets. With Q-learning, you are often / mostly bootstrapping from a max-action that was not actually taken, so it never sees any downward Bellman pressure.
This is consistent with another result I have noted: learning state value functions offline from a frozen policy without actions doesn’t seem to suffer from any value overestimation.