Authors: Tong Zheng , Xidong Wu , Zheng Zhang , Zhankui He , Chaoyi Zhang , Benjamin Coleman , Ruoqiao Wei , Di Bai , Haolin Liu , Rui Liu , Xue Wang , Yue Zhuan , Wang-Cheng Kang , Renkai Xiang , Heng Huang , Xinwu Cheng , Yunsong Guo View PDF HTML (experimental) Abstract: Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce \textsc{Dream-RSI}, a framework for scalable and recursively self-improving exploration.…