Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

Recurrent Looped Transformer

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 1

  1. Recurrent Looped Transformer adds per-token feedback to boost length generalization

    The paper introduces the Recurrent Looped Transformer (RLT), which splits its eight layers between a parallel causal encoder and a recurrent decoder that merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, two RLT splits trained on at most 40 bits generalize parity to 256 bits with 100% accuracy in every seed while an eight-layer Transformer stays at chance; swap-based S5 permutation tracking at eight times the training length reaches 97% versus under 1%, and modular arithmetic reaches up to 93% versus 33%. Ablations show the gains depend on the feedback: removing it drops parity and S5 to chance, and updating feedback once per four-token chunk keeps 64-bit parity at 99% but lowers length-64 S5 from 100% to 20%.

    Hugging Face · Papers · 🔥 0

Experience and discussion from the community

Share my Recurrent Looped Transformer experienceAsk about Recurrent Looped Transformer

Nobody has shared their experience with Recurrent Looped Transformer yet.