NP-OPD
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- NP-OPD adds negative-policy rollouts to on-policy distillation
The paper introduces Negative-Policy OPD (NP-OPD), which adds a lower-performing, lower-capability “negative policy” at the rollout stage of on-policy distillation (OPD), continuously supplying tokens that the negative policy prefers over the teacher so they stay under teacher supervision during training — without changing the distillation reward formulation. The authors report improvements over OPD across model scales, generation modes, reasoning domains and different OPD variants, and analyses indicating NP-OPD suppresses negative-policy-preferred tokens and moves the student away from the negative policy. The 25-page preprint (7 figures, 24 tables) says code will be released.
Hugging Face · Papers · 🔥 0
Experience and discussion from the community
Share my NP-OPD experienceAsk about NP-OPD
Nobody has shared their experience with NP-OPD yet.