A research paper published on arXiv proposes a new mechanism for large language models, termed 'sleep-like consolidation,' aimed at improving performance on long-horizon tasks. The proposed method involves the model periodically converting recent context into persistent fast weights and clearing its key-value cache. This 'sleep' phase consists of offline recurrent passes over accumulated context, updating state-space model blocks through a learned local rule.
The researchers tested this approach on synthetic tasks like cellular automata and multi-hop graph retrieval, as well as a math reasoning task. They report that their method, when compared to standard transformers and hybrid SSM-attention models, showed improved results, particularly on the math reasoning task where others failed.
According to the paper, increasing the 'sleep' duration, denoted by N, led to enhanced performance, with the most substantial gains observed in examples requiring deeper reasoning. The mechanism is designed to shift additional computation to the 'sleep' phase while aiming to preserve the latency of real-time predictions during the 'wake' phase.
The paper, authored by Sangyun Lee, Sean McLeish, Tom Goldstein, and Giulia Fanti, was submitted on May 25, 2026. It falls under the categories of Computation and Language and Artificial Intelligence.




