Write CUDA kernels for the sliding-window attention layer of Inkling, Thinking Machines Lab's open-weights model, which no public kernel covers.
2 × 4,096 tokens, 64 query heads, 512-token window · one H100
Preference Model builds RL environments for training capable, aligned superintelligence.
Our mission is to make sure AI goes well for everybody. So we work on the most direct lever we have: the reward signals that shape how models behave.
We grade only what we can verify, build tasks with no ceiling, and design for models that will be smarter than us.
Notes from building RL environments: how we design and grade them, and what frontier models do inside them.
Train a model on bad environments and you get a model that does bad things. We build better ones, for today’s models and for the ones that will outsmart us.
A sample of the tasks in our environments, and what a frontier model actually did on one run of each.
Write CUDA kernels for the sliding-window attention layer of Inkling, Thinking Machines Lab's open-weights model, which no public kernel covers.
2 × 4,096 tokens, 64 query heads, 512-token window · one H100
Play the real Slay the Spire through game tools alone: pick a route through each act's map, win turn-based card fights, and build a deck over hundreds of decisions.
GPT-6 Astra as the Ironclad · 494 turns · left the Heart at 152 of 750 HP
Post-train Qwen3-4B, an open reasoning model, to solve competition math problems with far less reasoning, without losing accuracy.
Fine-tuned on the model's own shortest correct solutions · 95.4% on 1,000 MATH problems
Build the training set for a classifier on frozen CLIP embeddings by picking and labeling images from over a million unlabeled web images, then get scored on a hidden benchmark.
Picked and labeled 445,912 web images across 200 classes · balanced accuracy of a linear probe
Jennifer Zhou is co-founder and CEO of Preference Model. Before founding the company, Jennifer worked on Anthropic's data team, building data infrastructure, tokenizers, and datasets, and previously worked at Stripe.
Ning Cao is co-founder of Preference Model, leading strategy, business development, and recruiting. Ning was previously early at DatologyAI, helping build the company from 0 to 1.
South Park Commons, Scale Angel Group, Manifund, MoE Capital, Fei-Fei Li, Ian Goodfellow, Jacob Jackson, Julian Schrittwieser, Sammy Sidhu, Barry McCardel, Dylan Patel, and swyx.
Anthropic, Stripe, Google DeepMind, Datology, EvolutionaryScale