Skip to slide
Chapter 5 · Recursive Language Models, Thinking Beyond the Memory Limit
38 / 53

CHAPTER 05 · Recursive Language Models, Thinking Beyond the Memory Limit · 4 / 6

Why this works so well

The payoff reported in the paper is striking. RLMs can handle inputs up to a hundred times larger than the model's normal context window. And even on inputs that would fit, they often beat a plain model, because they avoid the accuracy-drop-with-length problem by only ever looking closely at small, relevant pieces. Remarkably, they do this at a cost per query that is comparable to, or even cheaper than, the alternatives, because the model spends effort only on the parts that matter instead of processing everything.

This is another face of the theme running through this whole folder: inference-time compute, spending more thinking effort at the moment of answering. Chapter 1 spent it on writing steps; Chapter 3 spent it on generating and checking many solutions; here it is spent on programmatically exploring and recursively breaking down a massive input. The model decides, on its own, how to divide and conquer.

flowchart LR
    Theme[Spend more effort on harder problems] --> S1[CoT: more steps]
    Theme --> S2[Verify: more solutions checked]
    Theme --> S3[RLM: explore and recurse over huge inputs]
← → arrow keys work too