GRPO
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 2
- KLPO: a critic-free, KL-regularized policy optimization method for LLM agents
A new arXiv paper, "On KL-Regularized Policy Optimization" by Yifan Zhang (submitted 6 Oct 2026), proposes KLPO, which anchors the KL regularizer at the sampler so asynchronous RL for LLM agents can train on trajectories from stale checkpoints and mismatched inference probabilities without importance weights. The paper says the regularized improvement step has a closed-form Gibbs solution, and that for token-level policy mirror descent targets the gradient can be computed from terminal returns without a critic, using one rollout per prompt. It further proves independent Monte Carlo estimates of the KL term keep gradients unbiased and shows SPPO, GPO, REBEL and BPO arise as special cases of KLPO.
Hugging Face · Papers · 🔥 0 - HiPLEX: Hierarchical Policy Factorization for Full-Duplex Speech Language Models
The paper introduces HiPLEX, a reinforcement learning framework that factorizes a pretrained full-duplex text policy into a control policy deciding when to emit content (choosing among pad, epad and con) and a conditional content policy that picks a token only when con is selected. The authors report that across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities and shortens post-interruption response latency versus GRPO, while keeping comparable judged interruption-response quality, and better matches pooled human turn-timing and backchannel-rate marginals on Moshi and PersonaPlex.
Hugging Face · Papers · 🔥 0
Experience and discussion from the community
Share my GRPO experienceAsk about GRPO
Nobody has shared their experience with GRPO yet.