WEDNESDAY, SEPTEMBER 16, 2026|No. 15299
Artificial Intelligence

Dream-RSI Framework Promises Accelerated AI Self-Improvement

A new framework called Dream-RSI introduces a novel approach to recursive self-improvement in AI, aiming to enhance exploration strategies and reduce discovery costs.

An abstract representation of artificial intelligence and data processing.
An abstract representation of artificial intelligence and data processing. · Photo by Steve A Johnson on Unsplash
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Authors: Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo

Abstract: Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce \textsc{Dream-RSI}, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By performing dreaming in the replay simulator constructed from historical discovery trees, \textsc{Dream-RSI} secures immediate, low-cost off-policy feedback to evaluate and refine exploration policies without invoking repetitive, expensive online evaluations. The improved policy is subsequently redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, \textsc{Dream-RSI} achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.

Comments: 12 pages Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.14858 [cs.CL] (or arXiv:2609.14858v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.14858Focus to learn morearXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tong Zheng view email

[v1] Mon, 14 Sep 2026 00:10:47 UTC (777 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Dream-RSI: Recursive Self-Improvement through Evolving Worlds, by Tong Zheng and 16 other authors

view license

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →