01

About

Preference Model builds RL environments for training capable, aligned superintelligence.

Our mission is to make sure AI goes well for everybody. So we work on the most direct lever we have: the reward signals that shape how models behave.

We grade only what we can verify, build tasks with no ceiling, and design for models that will be smarter than us.

03

Join_Us

Build the future.

Train a model on bad environments and you get a model that does bad things. We build better ones, for today’s models and for the ones that will outsmart us.

All open roles
04

Use_Cases

A sample of the tasks in our environments, and what a frontier model actually did on one run of each.

Write CUDA kernels for the sliding-window attention layer of Inkling, Thinking Machines Lab's open-weights model, which no public kernel covers.

What it did
15×
faster than the PyTorch implementation

2 × 4,096 tokens, 64 query heads, 512-token window · one H100

PyTorch reference18.2 ms
Opus 52.0 ms
Fable 51.2 ms

Play the real Slay the Spire through game tools alone: pick a route through each act's map, win turn-based card fights, and build a deck over hundreds of decisions.

What it did
Floor 55
of 56: cleared three acts, died to the final boss

GPT-6 Astra as the Ironclad · 494 turns · left the Heart at 152 of 750 HP

1
Slime Boss
2
The Champ
3
Donu & Deca
4
Corrupt Heart

Post-train Qwen3-4B, an open reasoning model, to solve competition math problems with far less reasoning, without losing accuracy.

What it did
3.2×
fewer tokens per answer, with no accuracy lost

Fine-tuned on the model's own shortest correct solutions · 95.4% on 1,000 MATH problems

Reference4,846 tokens
Fable 51,530 tokens

Build the training set for a classifier on frozen CLIP embeddings by picking and labeling images from over a million unlabeled web images, then get scored on a hidden benchmark.

What it did
73%
on ImageNet-R, beating training on its own labeled images

Picked and labeled 445,912 web images across 200 classes · balanced accuracy of a linear probe

Reference65.1%
Fable 573.0%
05

Our_Team

Jennifer_Zhou

Jennifer Zhou is co-founder and CEO of Preference Model. Before founding the company, Jennifer worked on Anthropic's data team, building data infrastructure, tokenizers, and datasets, and previously worked at Stripe.

Ning_Cao

Ning Cao is co-founder of Preference Model, leading strategy, business development, and recruiting. Ning was previously early at DatologyAI, helping build the company from 0 to 1.

Backed by

South Park Commons, Scale Angel Group, Manifund, MoE Capital, Fei-Fei Li, Ian Goodfellow, Jacob Jackson, Julian Schrittwieser, Sammy Sidhu, Barry McCardel, Dylan Patel, and swyx.

Built by a team from

Anthropic, Stripe, Google DeepMind, Datology, EvolutionaryScale

06

Get_In_Touch

We are a small team committed to making big impact.