Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

Story

HiPLEX: Hierarchical Policy Factorization for Full-Duplex Speech Language Models

AI summary

The paper introduces HiPLEX, a reinforcement learning framework that factorizes a pretrained full-duplex text policy into a control policy deciding when to emit content (choosing among pad, epad and con) and a conditional content policy that picks a token only when con is selected. The authors report that across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities and shortens post-interruption response latency versus GRPO, while keeping comparable judged interruption-response quality, and better matches pooled human turn-timing and backchannel-rate marginals on Moshi and PersonaPlex.

Why it matters: Turn-taking and interruption handling are central pain points for real-time voice agents, and separating timing from content optimization offers a reusable training idea.

HiPLEXMoshiPersonaPlex

0
Source textHugging Face · Papers · 3 min read

Computer Science > Sound

arXiv:2610.07727 (cs)

[Submitted on 6 Oct 2026]

Title:HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models

Authors:Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park, Jinhyeok Yang, KiHyun Nam, Jaegul Choo, Jinkyu Lee

View a PDF of the paper titled HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models, by Kyudan Jung and 7 other authors

View PDF HTML (experimental)
Abstract:As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.
Comments: 34 pages, 9 figures, 17 tables,
Subjects: Sound (cs.SD)
Cite as: arXiv:2610.07727 [cs.SD]
  (or arXiv:2610.07727v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.07727

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kyudan Jung [view email]
[v1] Tue, 6 Oct 2026 04:23:23 UTC (340 KB)

Full-text links:

Access Paper:

license icon view license

Additional Features

Current browse context:

cs.SD

< prev   |   next >

new | recent | 2026-10

Change to browse by:

cs

References & Citations

Loading...

BibTeX formatted citation

Data provided by:

Bookmark

BibSonomy Reddit

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Read the original →

How we got here

  1. KLPO: a critic-free, KL-regularized policy optimization method for LLM agentsHugging Face · Papers · GRPO

Comments

I've used this: share my experience What I think: share my view
How important is this story?No ratings yet

No comments yet. Start the conversation.