Good teachers don’t cheat
TL;DR: Policy gradient RL, self-distillation techniques like SDFT , and Pedagogical RL can all be viewed as optimizing the same objective , just with slightly different optimization procedures. The privileged information that some of these methods feed in context is simply a tool to make the optimization of easier. The punchline is that, at optimality, the teacher’s use of has to vanish: good…