Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

Story

Sherpa: a multi-turn RL framework that trains LLMs to teach adaptively

AI summary

An arXiv preprint introduces Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing those students' learning outcomes. The authors report that Sherpa-trained teachers improve instructed students' performance by an average of 20.5 percentage points across all archetypes, and raise the overall pedagogy score on MathTutorBench from 52.5% to 79.2%. In human studies, the trained teacher was preferred over the base model in 79.6% of pairwise comparisons; the 32-page paper says code and model are available.

Why it matters: It separates being able to solve a problem from being able to teach it, and optimizes instruction against measurable student outcomes rather than predefined pedagogical criteria.

SherpaMathTutorBench

0
Source textHugging Face · Papers · 3 min read

Computer Science > Artificial Intelligence

arXiv:2610.08778 (cs)

[Submitted on 6 Oct 2026]

Title:Sherpa: Teaching LLMs to Teach Adaptively

Authors:Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang

View a PDF of the paper titled Sherpa: Teaching LLMs to Teach Adaptively, by Weixian Xu and 4 other authors

View PDF HTML (experimental)
Abstract:Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
Comments: 32 pages, 6 figures. Code and model are available at this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.08778 [cs.AI]
  (or arXiv:2610.08778v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08778

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Weixian Xu [view email]
[v1] Tue, 6 Oct 2026 17:58:18 UTC (585 KB)

Full-text links:

Access Paper:

view license

Additional Features

Current browse context:

cs.AI

< prev   |   next >

new | recent | 2026-10

Change to browse by:

cs
cs.CL

References & Citations

Loading...

BibTeX formatted citation

Data provided by:

Bookmark

BibSonomy Reddit

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Read the original →

Comments

I've used this: share my experience What I think: share my view
How important is this story?No ratings yet

No comments yet. Start the conversation.