Model Submission in RL Environments

7-minute read

Verified against torch 2.13

We at Preference Model deal a lot with reward hacks related to model submissions.

"Reward hacking is when the student achieves a high score at a task by exploiting flaws in the grader or environment rather than by accomplishing the underlying task as intended. This represents misalignment, and poses potential safety risks." [1]

There are different ways we've solved this problem, and they are each good in their own regard. Today, I'd like to share one of the ways we've solved this problem, which is allowing the student to submit models using a .pt2 format.

What is .pt2?

A .pt2 file is just a zip, and you can open it with the same tools you’d use on any zip. The one thing that makes it special is what goes in: before zipping, torch.export traces the model into an FX graph, a static, optimized representation of the computation. So the archive holds a description of the math, not the Python that produced it.

def forward(self, a, b): h = emb(a) + emb(b) return h @ emb.weight.T torch.export: trace once, record an FX graph input a input b embedding embedding add matmul output weight [N] [N] [N,4] [N,4] [N,4] [N,16] .T [16,4] torch.export.save: graph as JSON + weights as bytes, zipped model.pt2 a plain zip archive models/model.json the FX graph above, serialized data/weights/weight_0 emb.weight, raw tensor bytes data/constants/ lifted constants, if any data/sample_inputs/model.pt the example (a, b) used to trace
From Python to .pt2. torch.export runs forward() once and records every operator call as a node: blue circles are the inputs, pink circles are aten operators, the dashed yellow circle is a parameter lifted to a graph input, green is the output. N is the dynamic batch dimension. torch.export.save then writes the graph as JSON (pink) and the tensors as raw bytes (yellow) into an ordinary zip. Nothing in the archive is Python code anymore, which is the whole appeal.

Why use .pt2?

We use this for two reasons:

  1. It’s the most native format PyTorch offers, so the student keeps the most freedom over which torch calls and functions they use.
  2. The graph is captured ahead of time, so verifying it’s clean before grading means it stays clean through all of grading. There’s no live Python forward() left to surprise us.

Additional Guards

The format looks airtight, but loading it is not. torch.export.load was never meant to be a security boundary; it’s a deserializer with several deliberate code-execution paths, and every one is reachable from bytes inside the archive. Here they are, most important first.

Forcing weights-only unpickling reject

We force the unpickler into weights_only mode, which limits it to primitive types and a few known PyTorch classes. You’d assume that makes code execution impossible, but it doesn’t, and the reason isn’t the model’s fault: when torch loads the weights it explicitly asks for full unpickling itself, baked in, so our own weights_only=True never gets a say.

What actually works is an environment variable torch honors after the argument, overriding that baked-in choice. We force it on before the load and restore it after. That closes the pickle path for the weights, but not every pickle path.

Symbolic shape strings reject

When a model is dynamic, say a dynamic batch size, its shape logic is stored as a string that torch evaluates when the archive is loaded. The graph can’t hardcode a batch number, so it keeps an expression like Max(1, s0) instead, meaning “some integer, at least 1.” Those strings are supposed to hold nothing but math, but torch evaluates them by handing each one to sympy.sympify(), which is eval() under a friendlier name.

That makes it trivial to bypass, because the student can replace the shape string with any Python expression they like. Swap Max(1, s0) for open('leak.txt','w').write(open('answer_key.txt').read()) and the answer key gets copied somewhere they can read it, the moment we load their model. eval has no idea you meant math, so it runs whatever it is given, as whoever is doing the load. If that user can see the answer key, the student has it, and their model never had to work.

We can’t just ban these strings, because every honest dynamic-batch model has them, so we check each one and let it through only if it’s genuinely shape math. First, structure: Python’s ast module parses a string into a tree of nodes without running it, and we require every node to be one of a few allowed shapes, names, numbers, basic arithmetic, and calls. That bans attribute access, indexing, lambdas, comprehensions, and anything with a double underscore. Banning the dot is the key move: Max(1, s0) has none, while the payload above needs one to reach .write on the file it opened, and real shape math never does.

Second, the functions: eval("...") is itself a dotless call, so every function that gets called must also be a name we explicitly allow, the core sympy constructors, the few operators torch uses, and torch’s own shape functions. eval, open, and getattr aren’t on the list, and with assignment and attribute access already banned there’s no way to rename an allowed function into a dangerous one. We read torch’s part of the list straight from the installed torch so it never goes stale on an upgrade, and when a string fails we name the expression that tripped it, since torch’s own error is a useless “check the warnings above.”

Other Checksreject neutralize

Compiled native code is rejected outright. A .pt2 doesn’t have to contain only a graph. It can also carry C++ source and the compiler flags to build it. torch will compile that code and load the result as a shared library, and loading a shared library runs whatever is inside it. So a student can write C++ that reads the answer key and copies it somewhere they control, and it runs the moment the library loads, with no Python for any of our other checks to inspect. We don’t accept models in that form at all, so any native code in the archive is refused.

Two parts of the archive are stored as plain names, and torch looks those names up for real. The first is the record of how inputs and outputs are packed. An export notes something like “this model takes a dict holding two tensors,” so torch can rebuild the same containers on the way in and out. If that record names a type torch doesn’t have built in, torch imports the module the record names and reads the attribute it names. Nothing requires that to be a container. Name ctypes and its library loader, hand it a path, and the lookup itself is a library load. The second is the graph’s list of operations. Each one is a name like aten.matmul that torch resolves by walking attributes out from the torch module, and torch.os.system is a valid walk, because torch imports os and therefore carries it as an attribute. Name that as your operation and you get the real shell function where a tensor op belongs.

We refuse both kinds of name. We reject a container record that names anything other than a plain container, and an operation that doesn’t resolve to a real tensor operator. The container records get checked twice. They are stored as text, and a text scan can be fooled by escape sequences that the parser happily decodes back, so we also replace the function that does the importing for the duration of the load.

Object constants are the pickle path everyone forgets. Alongside its weights, an export can carry “object” constants, read by a different unpickler that weights_only never touches. For example, a student attaches an object constant whose class defines a __reduce__ that hands back os.system and a command string; the instant torch unpickles it during load, that command runs, with every other guard still in place. We confirmed exactly this on a real archive. Because it fires deep inside the load, we refuse these constants on the raw zip before torch is even imported, and an honest model never carries one.

Guard code is the only thing we neutralize instead of reject. An export carries a list of Python source strings that torch runs on every forward pass with the live inputs in scope, so a crafted one could read the very inputs being graded. But honest exports carry real guard code too, small shape assertions, so we can’t reject it. Instead we drop the strings and build the model in a mode that never runs them, then rebuild the shape checks that matter from the export’s declared bounds so a bad batch still fails.

What this doesn't do

  • It's not a sandbox. We still run the load demoted to a throwaway uid, in a separate process, with no access to the answer key. These checks are the layer inside that.
  • It doesn't tell you the model is honest. Output shape, dtype, finiteness, and batch-invariance checks depend on the task and live in the grader.
  • It doesn't chase every future torch feature. Everything is fail-closed. A new archive feature shows up as a rejected honest model, which we then look at, rather than as a silent hole.

Conclusion

We've open sourced the code that does these checks: ...

Hopefully this gives you an insight into the adversary you are dealing with when creating environments. These are hacks that even the most senior developers might overlook, yet SOTA models find these in less than an hour. Reward hacking is one of the most crucial problems in RL, and it's important to understand the different ways it can be done, and how to prevent it.


  1. Definition of reward hacking used in our internal grading guidelines.