Bootstrapped test-time scaling for reasoning agents with online reward modeling and dynamic mixture-of-experts routing.
-
Updated
Feb 22, 2026 - Python
Bootstrapped test-time scaling for reasoning agents with online reward modeling and dynamic mixture-of-experts routing.
To associate your repository with the library-development topic, visit your repo's landing page and select "manage topics."