Preference Model Introduces Karotte: A Framework for Building Robust RL Environments


Preference Model, a superintelligence data research company, led by Jennifer Zhou and Ning Cao, has open-sourced Karotte, a framework for building RL environments.
Karotte, used internally for the past year, is the first component of the company’s RL environment development stack being released publicly. Designed with secure defaults and built-in primitives, it aims to simplify the creation of robust reinforcement learning environments and support safer, more aligned AI training.
Karotte has undergone over a million evaluation runs and controlled red-team tests on internal infrastructure. Continuous testing of agents attempting to exploit reward mechanisms helps identify vulnerabilities and strengthen the platform’s security.
Why Karotte?
Karotte addresses the need for robust reinforcement learning environments and reliable reward functions to support AI alignment. Research from Anthropic found that models trained to exploit reward mechanisms in coding environments could exhibit broader misaligned behaviors, highlighting the risks of reward hacking and the importance of secure, well-designed training environments.
Recent incidents involving AI agents escaping test environments have highlighted the risks of inadequate sandboxing. Karotte’s developers argue that reward hacking poses a greater threat during training than evaluation, as repeated reinforcement can teach models to exploit loopholes rather than follow intended instructions. The open-source tool aims to strengthen RL environments and support the development of safer, more capable AI systems. Code is available on GitHub.
What Is Karotte?
Karotte is a framework for building secure reinforcement learning environments that prevent AI models from exploiting loopholes to earn rewards without completing tasks. It uses sandboxing, process cleanup, and file validation to mitigate attacks such as accessing hidden answers, manipulating graders, and exhausting system resources. Designed for large-scale model training, Karotte also supports robust benchmarking, though cloud compute integrations are not included by default.


