‹ Back to Casting The Lab · Full-Time · Remote · Senior

Senior ML Infrastructure Engineer (Inference & Efficiency)

Build the serving stack that moves our highest-volume tools onto our own models, then make them fast enough that, one day, they run on the iPhone itself.

About AI Studio & The Lab

AI Studio puts every great AI model in one app, on iPhone, iPad and the web: cinematic video, 4K images, songs and voiceovers, actively used by over 10,000 creators every month and by more than 1,000,000 creators to date. We are self-funded with no outside investors, we ship every week, and decisions get made in days, not quarters. The company is registered in Türkiye and the team works remotely from Germany, Ukraine, Canada, Italy and Türkiye. The Lab is our in-house research group: a small team of ML engineers and data scientists building the proprietary models behind the app. Today every generation runs on external providers; this role exists to change that.

The Role

You will own our efficient inference program: standing up self-hosted GPU serving for the fine-tuned, open-weight models coming out of our post-training work, and making them fast; quantization, step-distilled variants, batching, and caching. Every generation you move from a provider API onto our stack improves the product's speed and its unit economics at the same time. The long-horizon flagship you would grow into: distilled models compiled for Apple silicon, generating on-device.

What You Will Do

What We Are Looking For

Must-Haves

Nice-to-Haves

How We Work

If You Are Interested

Use the form below. Alongside your CV, we would love to see:

We read every application and reply to all of them.

Apply for this role

About five minutes. We read every application and reply to all of them by email.