Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

NP-OPD

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 1

  1. NP-OPD adds negative-policy rollouts to on-policy distillation

    The paper introduces Negative-Policy OPD (NP-OPD), which adds a lower-performing, lower-capability “negative policy” at the rollout stage of on-policy distillation (OPD), continuously supplying tokens that the negative policy prefers over the teacher so they stay under teacher supervision during training — without changing the distillation reward formulation. The authors report improvements over OPD across model scales, generation modes, reasoning domains and different OPD variants, and analyses indicating NP-OPD suppresses negative-policy-preferred tokens and moves the student away from the negative policy. The 25-page preprint (7 figures, 24 tables) says code will be released.

    Hugging Face · Papers · 🔥 0

Experience and discussion from the community

Share my NP-OPD experienceAsk about NP-OPD

Nobody has shared their experience with NP-OPD yet.