Active
Preference Learning at Scale
We train proprietary reward models on the behavioral signal of 1,000,000+ creators, systems that learn to judge images and video the way our audience does. They work ahead of the user: scoring candidates before results appear, catching weak generations for automatic retries, and learning which of the frontier models in production will serve a given prompt best, per credit spent.
The same program runs our internal evaluation harness. Every new model release is benchmarked within days of shipping, paired comparisons, confidence intervals, category-level breakdowns, so what we route to is a measurement, never a guess.