Sherpa
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- Sherpa: a multi-turn RL framework that trains LLMs to teach adaptively
An arXiv preprint introduces Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing those students' learning outcomes. The authors report that Sherpa-trained teachers improve instructed students' performance by an average of 20.5 percentage points across all archetypes, and raise the overall pedagogy score on MathTutorBench from 52.5% to 79.2%. In human studies, the trained teacher was preferred over the base model in 79.6% of pairwise comparisons; the 32-page paper says code and model are available.
Hugging Face · Papers · 🔥 0
Experience and discussion from the community
Share my Sherpa experienceAsk about Sherpa
Nobody has shared their experience with Sherpa yet.