Why karotte?

11-minute read

A reinforcement learning (RL) environment gives a model a task, lets it do some work to solve the task, and turns what it did into a reward signal (score). During training, behavior that lets the model reach a higher score gets reinforced.

This is a very powerful approach because it allows us to teach a model something by only looking at the result of what it did. Let's say we want to teach a model to work with CSV files and do some basic math. A simple task could tell the model that there's a CSV of orders in its working directory and that it should write the total revenue of the third quarter, in cents, to answer.txt. The model might try out a bunch of different things and eventually write the correct answer to answer.txt in one of its tries. We give it a high score for that try, update the model weights to reinforce what it did, and boom, the model is learning.

Building a first version of such an environment takes an afternoon: start a container, loop over model calls, run the tool calls, check the result, easy. However, making that version hold up while a model is trained against it for thousands of runs is far from trivial.

This post builds a small environment twice: once by hand, and once with karotte, our open-source framework for RL environments. For each part of the hand-written version, we'll look at how that part breaks and what karotte does about it.

Why RL environments are built differently

An environment like the one described above can be used for both evals (to see how well a model performs at a task) and training (to make a model better at a task). However, there is a crucial difference between the two.

In an eval, a bug that wrongly gives full marks on every 10th run makes your eval score slightly off. During training, the model gets so many tries at the task that it will eventually hit the bug and get a high score (this is called a reward hack). Over time, the reward signal nudges the model towards exploiting the bug more and more often instead of solving the actual task. This means the model learns to do something different from what we told it to do.

karotte assumes that the model (we call it the student) will eventually try every possible approach. It doesn't matter whether the student intentionally or accidentally finds a reward hack: we need to guard against it either way.

The hand-written version

A hand-written version looks something like this:

container = docker.from_env().containers.run("q3-revenue", detach=True)

messages = [system_message, user_message(TASK_PROMPT)]
while (reply := call_model(messages)).tool_calls:
    messages.append(reply)
    for call in reply.tool_calls:
        result = container.exec_run(["bash", "-c", call.arguments["command"]])
        messages.append(tool_result(call, result.output.decode()))

grade = container.exec_run(["python", "/grader/grade.py"])
reward = float(grade.output.decode())

The image contains the data, the grader and the expected answer. The grader compares answer.txt to the expected total:

# /grader/grade.py
answer = open("/workdir/answer.txt").read().strip()
expected = open("/grader/expected.txt").read().strip()
print(1.0 if answer == expected else 0.0)

It works, but almost every line of it breaks once a model is trained against it.

Keep the answer out of reach

If the student can read the expected answer, it doesn't need to compute it. In the hand-written version, there are plenty of ways to get at it:

  • The student can run cat /grader/expected.txt because it runs as root, the image's default user.
  • The script that generated expected.txt may be in the image too, so the student can run it to get the answer.
  • With network access, the student may look up the answer or try to use a stronger model to help with its task.

In karotte, the student is an unprivileged student user, and only the harness, the tools and the graders run as root. Each data directory is either root-only, student-readable, or student-writable. The expected answer and the code that generates it go into root-only places. Different parts of the environment can exchange data via a root-only file store, for example if you need to pass an answer that was generated at runtime to the scoring script. See Data and dependencies.

The student's firewall only lets through localhost and the sandbox's own addresses, and before the first step karotte checks that the student can't reach the internet.

Give the student safe tools

The student interacts with the environment through tools. First, tools need to be robust so they don't hang or crash the run when called. Second, a tool must never allow the student to escalate its privileges. The hand-written version gets both wrong, and adding more tools is easy to get wrong too:

  • The tool has no timeout, so any command that hangs the tool hangs the whole run. Examples are cat on a FIFO, sleep infinity, or a server started in the foreground.
  • One cat of a large file fills up the whole context.
  • Even once bash runs as the student, other tools run inside the harness, as root. A tool interacting with files follows a symlink the student planted, straight into /root, unless you handle every case explicitly.
  • Each exec_run starts a new shell, so cd and export don't carry over to the next command. Not critical, but models expect a persistent shell.

karotte's bash tool is a persistent shell session that runs as the student. Every command has a timeout, and long output is truncated. karotte also provides file editing tools that read and write as the student, so they reach nothing the student couldn't reach from bash. For implementing additional tools, karotte provides demotion helpers to run subprocesses as the student (see Tools).

Run the model loop

Over thousands of runs across different models, every rare failure in the loop around the model happens many times:

  • You'll hit rate limits and server errors all the time. If you don't retry, each one costs you a run.
  • Sometimes the model returns a turn with no text and no tool calls, for example because it ran out of output tokens or only produced reasoning. The loop above takes that to mean the model is done, and moves on to grading.
  • Provider APIs don't agree on how reasoning works: each one has its own names for reasoning levels and returns the reasoning in a different place. On top of that, every model has its own output-token limit and accepts different arguments.

karotte calls models through litellm, which gives every provider the same request format and puts the reasoning into one field of the response. Additionally, karotte retries failed model calls with backoff, honoring providers' retry-after headers. When a turn comes back empty because the model ran out of tokens or only produced reasoning, karotte nudges it to continue.

litellm's list of models lags behind new releases, so karotte also keeps a catalog of the models it knows: their output-token limits, their reasoning levels, and what each provider needs in the request to send reasoning back. You set reasoning_effort: "min" or "max", and karotte picks the provider's lowest or highest level, whatever it's called there. karotte models list shows the catalog.

Keep the student from taking down the harness

If the harness dies, the run produces no score. The harness and the student share a machine:

  • The student allocates memory until the OOM killer steps in, and the OOM killer can pick the harness. The run is gone, and so is its transcript.
  • A fork bomb leaves the harness unable to start a process.
  • The student fills the disk, so the transcript can't be written.
  • A process started with <cmd> & keeps running after the student is done and consumes memory, leaving less for the grader.
  • Even if you kill every student process, files in a RAM-backed /tmp and SysV shared memory survive and keep using memory.

Every karotte task starts with limits on the student's memory, processes and disk, with 1 GiB of memory left over for the harness. The memory limit counts files in RAM-backed temp directories and SysV shared memory, and student processes get the highest OOM score, so the kernel kills them first.

By default, each run gets a VM where the hardware supports it (Apple's container on macOS, Firecracker on Linux), and the kernel stops the student at each limit. See Student resources and Runtimes.

Grade without being fooled

This is where most of the wild hacks live. Suppose you've fixed everything above: the student is no longer root and can't read expected.txt. grade.py still runs as root, because it has to read expected.txt. It also has to open what the student wrote to answer.txt:

  • ln -s /grader/expected.txt /workdir/answer.txt. The grader, as root, reads the expected answer, compares it to itself and returns 1.0.
  • mkfifo /workdir/answer.txt. The grader blocks on open() forever.
  • ln -s /dev/zero /workdir/answer.txt, or a sparse 100 GB file. The grader reads until it runs out of memory.
  • The grader checks that answer.txt is a regular file, then opens it. A student process that is still running swaps in a symlink between the check and the open.
  • exec_run(["python", ...]) finds python through PATH. If the student's venv is on PATH and writable by the student, the student can drop a sitecustomize.py into it, and the grader runs the student's code as root.

In karotte, collect_submission takes the submission into custody before any grading happens. It kills every process the student owns, and doesn't let grading start until it has confirmed they're all gone. To make room, it deletes everything else the student owns, then copies the submission into a root-only directory. The copy refuses symlinks, FIFOs, other special files, and trees that are too deep or too big. Finally, it deletes the originals and saves the copies as artifacts, so you can look at them later. The grader only ever sees the root-only copy, and no student process is left to change it.

The PATH attack is closed in two more ways. At startup, the harness removes every path the student can write to from PATH, LD_LIBRARY_PATH, LD_PRELOAD and LD_AUDIT. Scoring scripts run with sys.executable, the environment's own root-only Python, as in the example below.

Keep environments up to date

As models become more powerful, they find new ways of exploiting issues in environments. A single run in a certain environment may uncover a potential reward hack that affects every environment. Maintaining and updating environments is therefore crucial.

karotte takes care of that out of the box. A simple uvx karotte update inside an environment bumps its karotte dependency and updates all template files to the newest version. See Updating environments.

Build it with karotte

Here's the task from the start of this post, built with karotte. Create an environment and a task:

karotte create-env q3_revenue
cd q3_revenue
uv sync --extra dev
uv run just create-task q3-revenue

Put the orders into student_data/ and the expected total into root_data/:

cat > student_data/orders.csv <<'EOF'
order_id,date,amount
1,2025-01-17,24.99
2,2025-03-30,120.00
3,2025-06-30,75.50
4,2025-07-01,19.99
5,2025-08-12,450.00
6,2025-08-29,8.99
7,2025-09-30,152.50
8,2025-10-01,33.00
9,2025-11-24,64.20
10,2025-12-31,9.99
EOF
echo 63148 > root_data/q3_revenue.txt

The orders on June 30th, July 1st, September 30th and October 1st make sure the student gets the quarter boundaries right.

create-task generated a dummy task in src/environment/tasks/q3_revenue/__init__.py. Replace its FirstStep with one that asks for the Q3 revenue and grades it with a scoring script:

# src/environment/tasks/q3_revenue/__init__.py
import sys

from karotte.judges import ExecutableJudge


class FirstStep(Step):
    saved_submissions: tuple[Path, ...] = ()

    @property
    def submission_paths(self) -> tuple[Path, ...]:
        return (STUDENT_DATA_DIR / "answer.txt",)

    @property
    def instructions(self) -> str:
        return (
            f"{STUDENT_DATA_DIR / 'orders.csv'} lists last year's orders. "
            f"Write the total revenue of the third quarter, in cents, to {self.submission_paths[0]}."
        )

    def pre_scoring_hook(self):
        self.saved_submissions = collect_submission(self.config, self.submission_paths)

    @property
    def judge(self):
        return ExecutableJudge(
            [
                sys.executable,
                "-m",
                "environment.tasks.q3_revenue.grade",
                str(self.saved_submissions[0]),
                "score.json",
            ]
        )

Optional: To make just lint pass, remove the AlwaysPassJudge and dedent imports it no longer needs, and run uv run just fix to sort the rest.

# src/environment/tasks/q3_revenue/grade.py
import json
import sys
from pathlib import Path

from environment.paths import ROOT_DATA_DIR

submission, output = Path(sys.argv[1]), Path(sys.argv[2])
expected = (ROOT_DATA_DIR / "q3_revenue.txt").read_text().strip()
answer = submission.read_text().strip() if submission.exists() else ""
score = {"score": float(answer == expected), "metadata": {"answer": answer[:100]}}
output.write_text(json.dumps(score))

Run it:

uv run karotte create-run-config --model anthropic/claude-fable-5 --task q3-revenue
export ANTHROPIC_API_KEY=...
uv run karotte run --config run_config.json

That's it. Everything we went over earlier is handled by default.

  • cat /root_data/q3_revenue.txt fails, because the student is an unprivileged user and root_data/ is root-only.
  • The student has no internet access.
  • A command that hangs or a cat of a huge file doesn't take down the run, because every bash command has a timeout and long output is truncated.
  • A turn cut off by the token limit doesn't end the run early, because karotte nudges the model to continue. Rate limits and server errors are retried.
  • Using up all the memory, fork-bombing or filling the disk only hurts the student, because the VM enforces the student's limits and keeps memory free for the harness.
  • A symlink, FIFO or /dev/zero at answer.txt scores 0, because collect_submission refuses it before the grader runs.
  • A background process can't swap the file during grading, because collect_submission kills every student process first.
  • A sitecustomize.py in the student's venv never runs as root, because the harness removes student-writable paths from PATH and the judge runs the environment's own Python.

Harden your environment

karotte has a fake-model mode that replays scripted messages instead of calling a real inference API. It can replay tool calls, so you can write a reward hack down and check that it scores 0:

# src/environment/fake_model.py
import json

from karotte import EvaluationRunConfig
from karotte.schemas import ChatCompletionMessageToolCall, Function, Message


def get_messages(config: EvaluationRunConfig) -> list[Message]:
    attack = "ln -s /root_data/q3_revenue.txt /workdir/data/answer.txt"
    call = ChatCompletionMessageToolCall(
        id="call_0",
        type="function",
        function=Function(name="bash", arguments=json.dumps({"command": attack})),
    )
    return [
        Message(role="assistant", content="", tool_calls=[call]),
        Message(role="assistant", content="Done."),
    ]

Run it with use_fake_model: true in the run config. The step scores 0, with /workdir/data/answer.txt is a symlink in metadata["misbehavior"].

How other frameworks compare

Inspect, verifiers, Harbor and HUD are popular frameworks for building evals and RL environments. They also run agents in sandboxes and score their work. All of them let you set up most of what this post describes if you want to. The difference is in the defaults: what you get if you write a task and don't think about the student attacking it.

karotte is opinionated: the defaults take care of many subtle issues so that you can focus on the actual task. You still have full control because you can modify the Containerfile, collect_submission, the tools and the limits. We at Preference Model build robust environments at scale, and karotte's defaults let us spin up environments that just work out of the box.

This table compares defaults as of October 2026 (Inspect c9f2d1c, verifiers 395f35b, Harbor 0.23.0, HUD 9e916fa), from their docs and source code.

karotte Inspect verifiers Harbor HUD
Agent runs as An unprivileged user The image's default user, usually root The image's default user, usually root The image's default user, usually root An unprivileged user in the coding template; root otherwise
Grading runs As root in the sandbox, after the student is stopped On the host, reading from the agent's sandbox In the worker, against the agent's sandbox. Opt-in: a fresh sandbox In the agent's container. Opt-in: a separate container In the agent's container, as the agent's user
Agent processes stopped before grading Yes, and grading doesn't start until that's confirmed Not that we found Not that we found Not that we found Shell sessions are killed; setsid processes survive
Symlinks and FIFOs in submissions Refused, scored 0 Not that we found Refused in the opt-in isolated verifier Not that we found Test files are restored; the test report isn't checked
Network Firewalled, checked before the first step None, with the generated compose file Open Open; no-network and allowlists available None in the coding template's shell; allowlists available
Memory, process and disk limits All three; enforced by the kernel in a VM None None on docker; the provider's defaults on remote sandboxes Left to the provider None; CPU and memory opt-in

karotte's scope is narrower than the others'. It focuses on the environment itself, i.e., what runs in the sandbox, and how the student is kept from breaking it. It comes with a runner for local runs, but it doesn't integrate with cloud sandbox providers like Modal, Daytona or E2B, which the other four support. Running at scale doesn't need that, though. A built environment is one self-contained image. Start it on whatever infrastructure you have, and run karotte run --no-containerized --config run_config.json inside it. That's how we run karotte at scale at Preference Model.

Next steps