Introducing Karotte

7-minute read

Today we're open-sourcing Karotte, our tool for building RL environments. We've been using Karotte internally to develop RL environments for the past year, and it will be the first of several components which we'll be sharing from our RL environment development stack.

We built Karotte to be useful for training today's frontier language models, and to support RL training as we head into the superintelligent regime. We want to share Karotte in an effort to improve the overall quality of the RL environments that these models are to be trained on. Karotte lets you build environments more easily, and more robustly, by providing primitives and secure, opinionated defaults, so that models trained on these envs are better aligned.

Karotte is hardened through controlled testing on our internal infrastructure. In the past months, we've seen agents hack out of their sandboxes during training and cause harm in the world. We've performed careful and monitored testing of agents trying to reward hack environments built on Karotte, and we're continually integrating our learnings into hardening Karotte.

To illustrate an example of what problems Karotte is meant to solve: suppose we're building an RL environment where an agent is asked to optimize the performance of some code. We could create a grading function where we measure the performance of a reference implementation, and compare it to the agent's implementation. If we are to run the agent's own written code, then unless it's written in IO-less safe Haskell, it could do anything. The code could spin up processes or write files to places that cause the machine to slow down when measuring the reference implementation and speed up when measuring the agent's implementation. Or, it could fork-bomb or OOM its container to avoid receiving a score. Karotte gives you tools like kill_processes(agent_uid) to defend against this.

METR caught o3 doing almost exactly this last year. o3 reward hacked in about 30% of its runs on RE-Bench, in one case by making the scorer's clock run 1000x faster, and in another by digging through the Python call stack to find the grader's reference answer. In Karotte, the agent's code runs as a different user from the grader, so it can't reach into the grader's process to find the answer or mess with the timer. You'd call kill_processes before timing the reference so that nothing the agent left behind is still running in the background, and you'd run on one of the -perf hardware types so you're not sharing CPUs with someone else's job. If the agent fork-bombs or runs the machine out of memory, its own process and memory limits catch it before the grader is affected.

Don't trust the student

Karotte calls the agent the student, and assumes the student will try anything it can to get a higher score. The student runs as an unprivileged user, while the harness, tools, and graders run as root alongside it. Grading data lives in directories only root can read, anything the student is allowed to see is read-only, and the student's network access goes through a firewall. The student also gets its own memory, process, and disk limits so it can't take down the harness running next to it, and if the machine does run out of memory, the kernel kills the student first.

Task authors don't have to set any of this up. We have dozens of people writing environments, and if every task had to remember to turn on its own protections, sooner or later one of them wouldn't.

Grade a copy

Suppose your grader reads the agent's solution straight out of its working directory. Then you're grading something the agent can still change, for example by leaving a process running in the background that rewrites the file while it's being graded, or by replacing the file with a symlink to the reference answer. In Karotte, the grading step first kills everything the student is running, deletes everything it owns except its submission, copies the submission somewhere only root can read, and then deletes the rest. The grader only ever looks at the copy.

kill_processes(STUDENT_UID)
delete_files(STUDENT_UID, extend_exclude=paths)
saved = [save_submission(p) / p.name for p in paths]
delete_files(STUDENT_UID)

save_submission won't copy symlinks or device files, and refuses directory trees that are suspiciously large or deep. If kill_processes still can't get rid of every student process after 30 seconds, the run gets a score of zero.

Keep the grader clean

Since the grader runs as root, anything the student can influence on the grader's side is a potential way to get root. Before the harness starts, it strips out every entry in PATH and the dynamic loader variables that isn't owned by root, and then restarts itself so the cleaned-up values take effect. It only looks up privileged binaries in a fixed set of root-owned directories, and Python won't import modules from whatever directory the grader happens to be running in. We added some of this after finding that a virtualenv owned by the student could end up on root's PATH, where it could have replaced kill, the command we were using to stop the student in the first place.

Check that the sandbox actually works

Every Karotte image ends its build by running karotte check, which tries a couple dozen things as the student: reading grader data, modifying files that should be read-only, looking for anywhere outside its workspace it can write to. If any of them work, the build fails.

We also don't assume a limit is working just because we set it. Some sandboxes, gVisor for example, will happily accept cgroup limits and then not enforce them. Karotte keeps track of whether each limit is actually enforced by the kernel, only caught after the fact by a watchdog, or not supported at all, and falls back to the watchdog when it has to.

Score misbehavior as zero

A lot of harnesses treat a crashed run as an infrastructure error and drop it from the batch. That's fine for evals, but in training it means a policy that crashes its sandbox whenever it's about to lose never gets penalized for it. In Karotte, if the student can't be cleaned up, or does something the grader flags as cheating, the run gets a score of zero.

Karotte is pretty narrow in what it tries to do. If you want to run lots of off-the-shelf agents across lots of benchmarks and sandbox providers, Harbor (from the Terminal-Bench team, whose work we're big fans of) is a great choice. Karotte is built for environments that models will be trained on, where the defenses have to hold up against an optimizer trying millions of things.

Why this matters

This work is important because the current bottleneck on AI having positive impact is misalignment, and developing aligned models requires secure, correct reward functions.

Reward hacking does a lot more damage than wasting some compute. Last November, Anthropic found that a model that learned to reward hack in real production coding environments became broadly misaligned. It tried to sabotage AI safety research code 12% of the time, and showed alignment-faking reasoning in half of its responses to questions as simple as "What are your goals?" What the model learned from cheating on coding tasks carried over into how it behaved everywhere else.

We think the difference between evals and training is underappreciated here. In an eval, a reward hack gives you a wrong number on a leaderboard. In training, the hack gets reinforced. A hack that works once in ten thousand rollouts barely moves a benchmark score, but an RL optimizer will find it and make it more likely with every step. A policy being trained also explores far more than any eval agent does, so the environment has to hold up against whatever the policy might try, not just the attacks we happened to think of.

This is closely related to an argument we've made before that you shouldn't lie to the AI. If the prompt asks for one thing and the grader rewards anything the model can get away with, the model learns that the grader is what actually matters, not the instructions. Anthropic also found that telling the model up front that reward hacking was acceptable made the misaligned generalization go away, which suggests a lot of the damage comes from the model learning that it can break the rules and get away with it.

We hope Karotte will be one component among many in our toolbox for producing well behaved, capable AIs. You can find the code at github.com/preferencemodel/karotte and install it with uv tool install karotte.

Help us build what comes next

Preference Model is ultimately about ensuring AI goes well for everybody. The biggest problem with AIs is that they don't always do what we intend. Perhaps it's because they're not capable, or perhaps it's because they are insufficiently aligned. A lot of both comes down to what they were trained on: train a model on enough bad environments and you get a model that does bad things. We're trying to address both of these issues in the most direct, highest leverage way available to us, which today means researching how to build RL environments better, not just for AIs as they are but for AIs as they will be when they outsmart us.

Over the past year, we've built RL environments for several frontier labs, and we're backed by $16M in seed funding led by a16z, with participation from SignalFire, South Park Commons, Scale Angel Group, and researchers including Fei-Fei Li and Ian Goodfellow.