Recurrent Looped Transformer
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- Recurrent Looped Transformer adds per-token feedback to boost length generalization
The paper introduces the Recurrent Looped Transformer (RLT), which splits its eight layers between a parallel causal encoder and a recurrent decoder that merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, two RLT splits trained on at most 40 bits generalize parity to 256 bits with 100% accuracy in every seed while an eight-layer Transformer stays at chance; swap-based S5 permutation tracking at eight times the training length reaches 97% versus under 1%, and modular arithmetic reaches up to 93% versus 33%. Ablations show the gains depend on the feedback: removing it drops parity and S5 to chance, and updating feedback once per four-token chunk keeps 64-bit parity at 99% but lowers length-64 S5 from 100% to 20%.
Hugging Face · Papers · 🔥 0
Experience and discussion from the community
Share my Recurrent Looped Transformer experienceAsk about Recurrent Looped Transformer
Nobody has shared their experience with Recurrent Looped Transformer yet.