About AI Studio & The Lab
AI Studio puts every great AI model in one app, on iPhone, iPad and the web: cinematic video, 4K images, songs and voiceovers, actively used by over 10,000 creators every month and by more than 1,000,000 creators to date. We are self-funded with no outside investors, we ship every week, and decisions get made in days, not quarters. The company is registered in Türkiye and the team works remotely from Germany, Ukraine, Canada, Italy and Türkiye. The Lab is our in-house research group: a small team of ML engineers and data scientists building the proprietary models behind the app. The founders come from data science and statistics; this role reports straight into that conversation.
The Role
You will own our preference learning program; the most defensible thing we build. The raw material is behavioral: which generations creators keep, discard, regenerate, and share, across 10+ frontier models in production and more than a million creators to date. From it you will build three systems: reward models that judge generations the way our creators do, learned routing that decides which model serves each prompt best per credit, and the statistically rigorous evaluation harness every model decision at AI Studio runs through.
- Type: Full-time
- Location: Fully remote; the team is spread across Germany, Ukraine, Canada, Italy and Türkiye
- Contract: Freelance / contract, long-term collaboration intended
- Data: Production-scale behavioral data from day one; no cold start
What You Will Do
- Model preferences: Turn implicit behavioral signal into clean preference data, handling selection bias, position effects, and noise, and train reward models for image and video quality on top of it.
- Put judgments to work: Deploy reward models for best-of-N selection, automatic retry of weak generations, and ranking; measurably lifting what users see without them ever knowing.
- Learn the routing: Replace hand-written provider rules with a model that predicts quality-per-credit for each prompt across every backend; improving output and unit economics at once.
- Build the harness: Design our internal benchmark; paired comparisons, Bradley-Terry / Elo aggregation, confidence intervals, category breakdowns, so every new frontier release gets a verdict within days.
- Design the experiments: Own experimental methodology for the Lab's launches: power analysis, guardrail metrics, and honest reads of A/B results.
What We Are Looking For
Must-Haves
- MSc or PhD in Statistics, Machine Learning, or a closely quantitative field, plus 3+ years of applied experience.
- Hands-on experience with preference and ranking models; reward modeling, RLHF/DPO pipelines, learning-to-rank, or recommender systems trained on implicit feedback.
- Serious statistical footing: paired-comparison models, experimental design, bias correction in observational data. You should enjoy that this job is half statistics.
- Strong Python and PyTorch, and the SQL / data-pipeline skills to build your own datasets from raw production events.
- Experience shipping models into production and owning their metrics afterward.
- Good written English; clear analysis memos are half the influence of this role.
Nice-to-Haves
- Reward modeling or evaluation work specifically for generative models (image, video, or LLM).
- Familiarity with perceptual and generative quality metrics and human-eval methodology.
- Contextual bandits or off-policy evaluation for the routing problem.
- Publications, or evidence of research taste in an applied setting.
How We Work
- Your data moat, your program: This dataset exists nowhere else and compounds daily. You define what gets built on it.
- Decisions run through you: Which models we onboard, how we price them in credits, what we route where; your harness is the referee.
- Ship weekly: Offline gains become online experiments in days, with 10,000+ monthly active users powering your tests.
- Competitive compensation: We value senior research talent and structure offers accordingly.
If You Are Interested
Use the form below. Alongside your CV, we would love to see:
- Your CV and LinkedIn profile, plus a short introduction.
- The preference, ranking, or evaluation system you are proudest of; the data, the method, and what it moved.
- In a paragraph: how you would turn keep / discard / regenerate events into unbiased preference pairs.
- Your expected compensation for this position.
We read every application and reply to all of them.
Apply for this role
About five minutes. We read every application and reply to all of them by email.