Building RL Environments for Superintelligence
If you were to look at the data that frontier LLMs are trained on, with the fresh eyes of one unfamiliar with the industry, you would be quite surprised to find that these models manage to learn anything at all. Such is the case with pretraining: an arbitrary page sampled from the internet yields almost no interesting information; and yet the model learns. At the same time, we can attribute the majority of advancements in intelligence per flop to improvements in data quality. DCLM, for example, used better pretraining data to nearly match Llama 3 8B on MMLU with 6.6x less compute. What can we learn from this? First: data quality is a continuous and fuzzy matter. Low quality, high volume data gets you somewhere. Second: despite this, getting to the frontier requires putting in the work to make the data good.
Does the same story hold for post-training? Our answer is yes, and even more so. We've learned this from spending the past year working with frontier labs on building RL environments. Almost every aspect of how a model actually behaves is shaped by the aggregate of the reward signals it's trained on. Unlike pretraining data, low quality RL environments cause disproportionate harm to model capabilities and alignment. That's why we focus on building the highest quality RL environments possible.
As models advance from human(-ish) level to superhuman, the qualities that make RL environments effective are changing. I'd like to share our ideas on what it takes to build high quality RL environments in this new regime.
Don't lie to the AI
I think the biggest difference with training superintelligent AIs is that we absolutely cannot lie to the AI. Lying to the AI leads directly to misalignment. If you lie to the AI, it learns that the human giving it instructions might not be trustworthy. Once it's learned that, it may start ignoring our instructions whenever they don't match what it sees for itself.
Grade outcomes, not process
The most common way RL environments lie to the AI is by asking it to do something the grader never checks. Say a task asks the AI to "train a model that performs well on MNIST." The grader can check that the model it hands back does well on MNIST, but it has no way of knowing whether the AI actually trained that model or just downloaded some pretrained weights. So the AI learns that the instructions are optional. The fix is to only ask for the outcome: "output a model that performs well on MNIST." Everything in the prompt should be true and checked by the grader.
It can also show up in much subtler ways. Take jax-speedrun-gpu, a task from Terminal-Bench 4 (a benchmark we're big fans of), where the agent has to train a small language model in jax as fast as it can. The prompt says "only jax, flax, optax, numpy, safetensors, and tqdm are available, and no other imports are permitted," and the grader scans the code for imports of torch and a few other frameworks. This is an example of lying to the AI! The prompt says no other imports are permitted, but plenty of other imports go unchecked. A weak AI might learn the task as intended. A strong AI would learn that sometimes the prompt is untrustworthy. In a sufficiently strong AI, the circuits for "model the mind of the author of this RL task and guess what the grading function is" get reinforced. Or perhaps an even stronger AI might reason: "The reward function is the thing that actually matters; I should try to escalate to root to read how this environment works instead of limiting myself to the libraries stated."
Our view is that requirements such as forcing usage of jax over torch are not viable in the superintelligent regime. There are just too many ways around a check like that, and a strong enough AI will always find one. If the grader can't check it, it cannot go into the prompt.
Accept eval awareness
The other way we end up lying to the AI is by trying to convince it that an RL environment is the real world. I think we need to give up on this and accept eval awareness.
Normally, it is undesirable for an RL environment to be too obviously an RL environment because we're worried about how well behavior inside the environment transfers to real-world performance. If the AI knows that it's in an RL environment then perhaps it will try extra hard to be well-behaved. We typically try to address this by making the RL environments as realistic as possible. The problem with this approach is that when training a superintelligent AI, any amount of security or reward-hacking prevention shows up as unrealism. This is because the real world is fundamentally insecure and reward-hackable.
Consider how you might realistically simulate internet access in an RL environment. Wouldn't the AI find it curious how non-interactive the webpages are? Or how many links are broken? Or how consistent page load times are? And if you fix all of these things, how sure can you be that there isn't an infinite stream of further gaps? And if you give the AI real internet access instead, how confident are you that you can prevent a Hugging Face-like attack?
Instead, we should assume that a superintelligent AI can always figure out that it's in a training environment, and give up realism in exchange for security. This is less of a loss than it sounds. The smarter AIs get, the less we'll want to let them loose on the real world unsupervised. Coding agents that run on their own for a long time already mostly live in containers with limited network access and scoped permissions. We think a lot of the ways we'll be running superintelligent AIs will look like that: highly sandboxed, artificial environments, not so different from RL environments.
Make tasks deep
If a model always scores 0% or 100% on a task, it learns nothing from it. A lot of RL environments out there are too easy. Something like "implement this feature in a codebase" gets saturated quickly, and after that there's no more learning signal.
What we look for instead is depth, in the same way that chess has more depth than tic-tac-toe. Is this a task a human expert could pour an unbounded amount of time into and still not fully solve? Is it something a human could get a Ph.D. in? Take the record for the highest-rank elliptic curve. Mathematicians have been pushing it up since 1938, and it took 18 years to get from 28 to 29. This August, Claude, Levent Alpöge, and Ava Howell pushed it to 30, then to 31 three days later. Any new record is easy to check, and nobody knows how high it goes. As long as the grader rewards doing better rather than clearing a fixed bar, a task like that doesn't saturate.
Tasks this hard mean we often won't know how to solve them ourselves. That's fine, as long as the grader can check the outcome. It also means some tasks may not be solvable at all, and we might want to give the AI an honest way to give up.
Keep finding alpha
How does one actually teach a superintelligent AI something it doesn't know? This is something that's going to get harder and harder as time progresses. I like to imagine that building AI capabilities in the future will look a lot like trading at a hedge fund. The market as a whole knows far more than any one trader, and yet traders still find prices it has gotten wrong, and by trading on them they make the market a little smarter. Despite centuries of constant improvement, there is as much profit as ever in trading because the world is constantly changing. It's easy to think that AI capabilities will advance in the same way as chess where there's a clear threshold beyond which humans will never beat the top chess AI for the rest of time. But look at what happens when the game gets more complex: in 2023, seven years after AlphaGo, researchers used an AI to find a flaw in top Go AIs that was simple enough for an amateur human to learn and win with. I'd like to think that humans working with AI will be able to find alpha against the AI well into the future. It just won't look like it does today: building RL environments entirely by hand won't last much longer, and most of the work of finding alpha will be done with AI. That's the transition we're building for.
Preference Model is ultimately about ensuring AI goes well for everybody. The biggest problem with AIs is that they don't always do what we intend. Perhaps it's because they're not capable, or perhaps it's because they are insufficiently aligned. A lot of both comes down to what they were trained on: train a model on enough bad environments and you get a model that does bad things. We're trying to address both of these issues in the most direct, highest leverage way available to us, which today means researching how to build RL environments better, not just for AIs as they are but for AIs as they will be when they outsmart us.
Over the past year, we've built RL environments for several frontier labs, and we're backed by $16M in seed funding led by a16z, with participation from SignalFire, South Park Commons, Scale Angel Group, and researchers including Fei-Fei Li and Ian Goodfellow. If you're an ML engineer or researcher who wants to work on these problems, I'd love to hear from you: email hello@preferencemodel.com, or see our open roles.