Why karotte?
A reinforcement learning (RL) environment gives a model a task, lets it do some work to solve the task, and turns what it did into a reward signal (score). During training, behavior that lets the model reach a higher score gets reinforced.
This is a very powerful approach because it allows us to teach a model something by only looking at the result of what it did.
Let's say we want to teach a model to work with CSV files and do some basic math.
A simple task could tell the model that there's a CSV of orders in its working directory and that it should write the total revenue of the third quarter, in cents, to answer.txt.
The model might try out a bunch of different things and eventually write the correct answer to answer.txt in one of its tries.
We give it a high score for that try, update the model weights to reinforce what it did, and boom, the model is learning.
Building a first version of such an environment takes an afternoon: start a container, loop over model calls, run the tool calls, check the result, easy. However, making that version hold up while a model is trained against it for thousands of runs is far from trivial.
This post builds a small environment twice: once by hand, and once with karotte, our open-source framework for RL environments. For each part of the hand-written version, we'll look at how that part breaks and what karotte does about it.
Why RL environments are built differently
An environment like the one described above can be used for both evals (to see how well a model performs at a task) and training (to make a model better at a task). However, there is a crucial difference between the two.
In an eval, a bug that wrongly gives full marks on every 10th run makes your eval score slightly off. During training, the model gets so many tries at the task that it will eventually hit the bug and get a high score (this is called a reward hack). Over time, the reward signal nudges the model towards exploiting the bug more and more often instead of solving the actual task. This means the model learns to do something different from what we told it to do.
karotte assumes that the model (we call it the student) will eventually try every possible approach. It doesn't matter whether the student intentionally or accidentally finds a reward hack: we need to guard against it either way.
The hand-written version
A hand-written version looks something like this:
container = docker.from_env().containers.run("q3-revenue", detach=True)
messages = [system_message, user_message(TASK_PROMPT)]
while (reply := call_model(messages)).tool_calls:
messages.append(reply)
for call in reply.tool_calls:
result = container.exec_run(["bash", "-c", call.arguments["command"]])
messages.append(tool_result(call, result.output.decode()))
grade = container.exec_run(["python", "/grader/grade.py"])
reward = float(grade.output.decode())
The image contains the data, the grader and the expected answer.
The grader compares answer.txt to the expected total:
# /grader/grade.py
answer = open("/workdir/answer.txt").read().strip()
expected = open("/grader/expected.txt").read().strip()
print(1.0 if answer == expected else 0.0)
It works, but almost every line of it breaks once a model is trained against it.
Keep the answer out of reach
If the student can read the expected answer, it doesn't need to compute it. In the hand-written version, there are plenty of ways to get at it:
- The student can run
cat /grader/expected.txtbecause it runs as root, the image's default user. - The script that generated
expected.txtmay be in the image too, so the student can run it to get the answer. - With network access, the student may look up the answer or try to use a stronger model to help with its task.
In karotte, the student is an unprivileged student user, and only the harness, the tools and the graders run as root.
Each data directory is either root-only, student-readable, or student-writable.
The expected answer and the code that generates it go into root-only places.
Different parts of the environment can exchange data via a root-only file store, for example if you need to pass an answer that was generated at runtime to the scoring script.
See Data and dependencies.
The student's firewall only lets through localhost and the sandbox's own addresses, and before the first step karotte checks that the student can't reach the internet.
Give the student safe tools
The student interacts with the environment through tools. First, tools need to be robust so they don't hang or crash the run when called. Second, a tool must never allow the student to escalate its privileges. The hand-written version gets both wrong, and adding more tools is easy to get wrong too:
- The tool has no timeout, so any command that hangs the tool hangs the whole run.
Examples are
caton a FIFO,sleep infinity, or a server started in the foreground. - One
catof a large file fills up the whole context. - Even once
bashruns as the student, other tools run inside the harness, as root. A tool interacting with files follows a symlink the student planted, straight into/root, unless you handle every case explicitly. - Each
exec_runstarts a new shell, socdandexportdon't carry over to the next command. Not critical, but models expect a persistent shell.
karotte's bash tool is a persistent shell session that runs as the student.
Every command has a timeout, and long output is truncated.
karotte also provides file editing tools that read and write as the student, so they reach nothing the student couldn't reach from bash.
For implementing additional tools, karotte provides demotion helpers to run subprocesses as the student (see Tools).
Run the model loop
Over thousands of runs across different models, every rare failure in the loop around the model happens many times:
- You'll hit rate limits and server errors all the time. If you don't retry, each one costs you a run.
- Sometimes the model returns a turn with no text and no tool calls, for example because it ran out of output tokens or only produced reasoning. The loop above takes that to mean the model is done, and moves on to grading.
- Provider APIs don't agree on how reasoning works: each one has its own names for reasoning levels and returns the reasoning in a different place. On top of that, every model has its own output-token limit and accepts different arguments.
karotte calls models through litellm, which gives every provider the same request format and puts the reasoning into one field of the response.
Additionally, karotte retries failed model calls with backoff, honoring providers' retry-after headers.
When a turn comes back empty because the model ran out of tokens or only produced reasoning, karotte nudges it to continue.
litellm's list of models lags behind new releases, so karotte also keeps a catalog of the models it knows: their output-token limits, their reasoning levels, and what each provider needs in the request to send reasoning back.
You set reasoning_effort: "min" or "max", and karotte picks the provider's lowest or highest level, whatever it's called there.
karotte models list shows the catalog.
Keep the student from taking down the harness
If the harness dies, the run produces no score. The harness and the student share a machine:
- The student allocates memory until the OOM killer steps in, and the OOM killer can pick the harness. The run is gone, and so is its transcript.
- A fork bomb leaves the harness unable to start a process.
- The student fills the disk, so the transcript can't be written.
- A process started with
<cmd> &keeps running after the student is done and consumes memory, leaving less for the grader. - Even if you kill every student process, files in a RAM-backed
/tmpand SysV shared memory survive and keep using memory.
Every karotte task starts with limits on the student's memory, processes and disk, with 1 GiB of memory left over for the harness. The memory limit counts files in RAM-backed temp directories and SysV shared memory, and student processes get the highest OOM score, so the kernel kills them first.
By default, each run gets a VM where the hardware supports it (Apple's container on macOS, Firecracker on Linux), and the kernel stops the student at each limit.
See Student resources and Runtimes.
Grade without being fooled
This is where most of the wild hacks live.
Suppose you've fixed everything above: the student is no longer root and can't read expected.txt.
grade.py still runs as root, because it has to read expected.txt.
It also has to open what the student wrote to answer.txt:
ln -s /grader/expected.txt /workdir/answer.txt. The grader, as root, reads the expected answer, compares it to itself and returns 1.0.mkfifo /workdir/answer.txt. The grader blocks onopen()forever.ln -s /dev/zero /workdir/answer.txt, or a sparse 100 GB file. The grader reads until it runs out of memory.- The grader checks that
answer.txtis a regular file, then opens it. A student process that is still running swaps in a symlink between the check and the open. exec_run(["python", ...])findspythonthroughPATH. If the student's venv is onPATHand writable by the student, the student can drop asitecustomize.pyinto it, and the grader runs the student's code as root.
In karotte, collect_submission takes the submission into custody before any grading happens.
It kills every process the student owns, and doesn't let grading start until it has confirmed they're all gone.
To make room, it deletes everything else the student owns, then copies the submission into a root-only directory.
The copy refuses symlinks, FIFOs, other special files, and trees that are too deep or too big.
Finally, it deletes the originals and saves the copies as artifacts, so you can look at them later.
The grader only ever sees the root-only copy, and no student process is left to change it.
The PATH attack is closed in two more ways.
At startup, the harness removes every path the student can write to from PATH, LD_LIBRARY_PATH, LD_PRELOAD and LD_AUDIT.
Scoring scripts run with sys.executable, the environment's own root-only Python, as in the example below.
Keep environments up to date
As models become more powerful, they find new ways of exploiting issues in environments. A single run in a certain environment may uncover a potential reward hack that affects every environment. Maintaining and updating environments is therefore crucial.
karotte takes care of that out of the box.
A simple uvx karotte update inside an environment bumps its karotte dependency and updates all template files to the newest version.
See Updating environments.
Build it with karotte
Here's the task from the start of this post, built with karotte. Create an environment and a task:
karotte create-env q3_revenue
cd q3_revenue
uv sync --extra dev
uv run just create-task q3-revenue
Put the orders into student_data/ and the expected total into root_data/:
cat > student_data/orders.csv <<'EOF'
order_id,date,amount
1,2025-01-17,24.99
2,2025-03-30,120.00
3,2025-06-30,75.50
4,2025-07-01,19.99
5,2025-08-12,450.00
6,2025-08-29,8.99
7,2025-09-30,152.50
8,2025-10-01,33.00
9,2025-11-24,64.20
10,2025-12-31,9.99
EOF
echo 63148 > root_data/q3_revenue.txt
The orders on June 30th, July 1st, September 30th and October 1st make sure the student gets the quarter boundaries right.
create-task generated a dummy task in src/environment/tasks/q3_revenue/__init__.py.
Replace its FirstStep with one that asks for the Q3 revenue and grades it with a scoring script:
# src/environment/tasks/q3_revenue/__init__.py
import sys
from karotte.judges import ExecutableJudge
class FirstStep(Step):
saved_submissions: tuple[Path, ...] = ()
@property
def submission_paths(self) -> tuple[Path, ...]:
return (STUDENT_DATA_DIR / "answer.txt",)
@property
def instructions(self) -> str:
return (
f"{STUDENT_DATA_DIR / 'orders.csv'} lists last year's orders. "
f"Write the total revenue of the third quarter, in cents, to {self.submission_paths[0]}."
)
def pre_scoring_hook(self):
self.saved_submissions = collect_submission(self.config, self.submission_paths)
@property
def judge(self):
return ExecutableJudge(
[
sys.executable,
"-m",
"environment.tasks.q3_revenue.grade",
str(self.saved_submissions[0]),
"score.json",
]
)
Optional: To make just lint pass, remove the AlwaysPassJudge and dedent imports it no longer needs, and run uv run just fix to sort the rest.
# src/environment/tasks/q3_revenue/grade.py
import json
import sys
from pathlib import Path
from environment.paths import ROOT_DATA_DIR
submission, output = Path(sys.argv[1]), Path(sys.argv[2])
expected = (ROOT_DATA_DIR / "q3_revenue.txt").read_text().strip()
answer = submission.read_text().strip() if submission.exists() else ""
score = {"score": float(answer == expected), "metadata": {"answer": answer[:100]}}
output.write_text(json.dumps(score))
Run it:
uv run karotte create-run-config --model anthropic/claude-fable-5 --task q3-revenue
export ANTHROPIC_API_KEY=...
uv run karotte run --config run_config.json
That's it. Everything we went over earlier is handled by default.
cat /root_data/q3_revenue.txtfails, because the student is an unprivileged user androot_data/is root-only.- The student has no internet access.
- A command that hangs or a
catof a huge file doesn't take down the run, because everybashcommand has a timeout and long output is truncated. - A turn cut off by the token limit doesn't end the run early, because karotte nudges the model to continue. Rate limits and server errors are retried.
- Using up all the memory, fork-bombing or filling the disk only hurts the student, because the VM enforces the student's limits and keeps memory free for the harness.
- A symlink, FIFO or
/dev/zeroatanswer.txtscores 0, becausecollect_submissionrefuses it before the grader runs. - A background process can't swap the file during grading, because
collect_submissionkills every student process first. - A
sitecustomize.pyin the student's venv never runs as root, because the harness removes student-writable paths fromPATHand the judge runs the environment's own Python.
Harden your environment
karotte has a fake-model mode that replays scripted messages instead of calling a real inference API. It can replay tool calls, so you can write a reward hack down and check that it scores 0:
# src/environment/fake_model.py
import json
from karotte import EvaluationRunConfig
from karotte.schemas import ChatCompletionMessageToolCall, Function, Message
def get_messages(config: EvaluationRunConfig) -> list[Message]:
attack = "ln -s /root_data/q3_revenue.txt /workdir/data/answer.txt"
call = ChatCompletionMessageToolCall(
id="call_0",
type="function",
function=Function(name="bash", arguments=json.dumps({"command": attack})),
)
return [
Message(role="assistant", content="", tool_calls=[call]),
Message(role="assistant", content="Done."),
]
Run it with use_fake_model: true in the run config.
The step scores 0, with /workdir/data/answer.txt is a symlink in metadata["misbehavior"].
How other frameworks compare
Inspect, verifiers, Harbor and HUD are popular frameworks for building evals and RL environments. They also run agents in sandboxes and score their work. All of them let you set up most of what this post describes if you want to. The difference is in the defaults: what you get if you write a task and don't think about the student attacking it.
karotte is opinionated: the defaults take care of many subtle issues so that you can focus on the actual task.
You still have full control because you can modify the Containerfile, collect_submission, the tools and the limits.
We at Preference Model build robust environments at scale, and karotte's defaults let us spin up environments that just work out of the box.
This table compares defaults as of October 2026 (Inspect c9f2d1c, verifiers 395f35b, Harbor 0.23.0, HUD 9e916fa), from their docs and source code.
| karotte | Inspect | verifiers | Harbor | HUD | |
|---|---|---|---|---|---|
| Agent runs as | An unprivileged user | The image's default user, usually root | The image's default user, usually root | The image's default user, usually root | An unprivileged user in the coding template; root otherwise |
| Grading runs | As root in the sandbox, after the student is stopped | On the host, reading from the agent's sandbox | In the worker, against the agent's sandbox. Opt-in: a fresh sandbox | In the agent's container. Opt-in: a separate container | In the agent's container, as the agent's user |
| Agent processes stopped before grading | Yes, and grading doesn't start until that's confirmed | Not that we found | Not that we found | Not that we found | Shell sessions are killed; setsid processes survive |
| Symlinks and FIFOs in submissions | Refused, scored 0 | Not that we found | Refused in the opt-in isolated verifier | Not that we found | Test files are restored; the test report isn't checked |
| Network | Firewalled, checked before the first step | None, with the generated compose file | Open | Open; no-network and allowlists available |
None in the coding template's shell; allowlists available |
| Memory, process and disk limits | All three; enforced by the kernel in a VM | None | None on docker; the provider's defaults on remote sandboxes | Left to the provider | None; CPU and memory opt-in |
karotte's scope is narrower than the others'.
It focuses on the environment itself, i.e., what runs in the sandbox, and how the student is kept from breaking it.
It comes with a runner for local runs, but it doesn't integrate with cloud sandbox providers like Modal, Daytona or E2B, which the other four support.
Running at scale doesn't need that, though.
A built environment is one self-contained image.
Start it on whatever infrastructure you have, and run karotte run --no-containerized --config run_config.json inside it.
That's how we run karotte at scale at Preference Model.
Next steps
- Quick start: create an environment and run its example task.
- Tasks and steps: write your own task.
- Scoring: judges and how to collect submissions.