Preference Model has open-sourced Karotte, an opinionated framework designed to help build secure reinforcement learning (RL) environments and prevent reward hacking during AI training. Hardened through extensive internal testing, Karotte uses sandboxing and secure defaults to protect against exploits like fork bombs, memory exhaustion, and sandbox escapes. By ensuring models are properly evaluated and trained on accurate reward functions, the framework aims to support the development of aligned and well-behaved AI models.

We are open-sourcing Karotte, our internal framework for building robust and reliable reinforcement learning environments. Karotte is designed to simplify the creation of RL tasks and agents, ensuring reproducibility and high performance.