Google researchers and several universities developed RRSI, a method that stops self-improving AI agents from overspecializing and memorizing test tasks. The approach lifts scores on unseen benchmarks by up to 4.7 points while using about 30 percent fewer tokens than unregularized versions.