About AI Studio & The Lab
AI Studio puts every great AI model in one app, on iPhone, iPad and the web: cinematic video, 4K images, songs and voiceovers, actively used by over 10,000 creators every month and by more than 1,000,000 creators to date. We are self-funded with no outside investors, we ship every week, and decisions get made in days, not quarters. The company is registered in Türkiye and the team works remotely from Germany, Ukraine, Canada, Italy and Türkiye. The Lab is our in-house research group: a small team of ML engineers and data scientists building the proprietary models behind the app. Today every generation runs on external providers; this role exists to change that.
The Role
You will own our efficient inference program: standing up self-hosted GPU serving for the fine-tuned, open-weight models coming out of our post-training work, and making them fast; quantization, step-distilled variants, batching, and caching. Every generation you move from a provider API onto our stack improves the product's speed and its unit economics at the same time. The long-horizon flagship you would grow into: distilled models compiled for Apple silicon, generating on-device.
- Type: Full-time
- Location: Fully remote; the team is spread across Germany, Ukraine, Canada, Italy and Türkiye
- Contract: Freelance / contract, long-term collaboration intended
- Scope: Greenfield; you design the serving stack from the first GPU up
What You Will Do
- Stand up self-hosted inference: Design and run our GPU serving layer for open-weight image and video models; provisioning, autoscaling, queuing, and observability, built to sit behind our existing Firebase generation pipeline.
- Make models fast: Quantization, torch.compile / TensorRT graph optimization, attention and caching tricks, batched throughput tuning; chasing latency and cost per generation as first-class metrics.
- Deploy distilled variants: Productionize the Lab's few-step distilled models for near-instant draft previews.
- Engineer the economics: Own cost per generation across self-hosted and provider backends; capacity planning, spot strategies, and the dashboards that prove the margin story.
- Pave the road to on-device: Prototype Core ML / MLX conversions of our smallest distilled models and map what on-iPhone generation will take.
What We Are Looking For
Must-Haves
- 3+ years in ML systems or backend infrastructure, with real GPU inference services you built and operated in production.
- Deep PyTorch deployment experience: quantization, compilation (torch.compile, TensorRT, or ONNX), and memory / throughput profiling.
- Hands-on with modern serving and orchestration; containerized GPU workloads, autoscaling, queue architectures, and the failure modes of long-running generation jobs.
- Strong grasp of diffusion-model inference specifically; samplers, step counts, attention costs, VRAM pressure; enough to co-design with researchers, not just host their checkpoints.
- Cost-engineering instincts: you can say what a generation costs and defend the number.
- Good written English for async collaboration and runbooks.
Nice-to-Haves
- Core ML, MLX, or Metal experience, or a serious itch to compile models for Apple silicon.
- Custom CUDA / Triton kernel work.
- Experience serving video models in particular (long jobs, chunked decoding, webhook pipelines).
- Familiarity with Firebase / GCP, where our backend lives.
How We Work
- Direct P&L impact: Your latency and cost graphs are the company's margin story; few infra jobs are this legible.
- Greenfield, not legacy: No inherited stack to babysit. You make the architecture calls and live with them.
- Ship weekly: Small team, direct access to founders, production from week one.
- Competitive compensation: We value senior infrastructure talent and structure offers accordingly.
If You Are Interested
Use the form below. Alongside your CV, we would love to see:
- Your CV and LinkedIn profile, plus a short introduction.
- The inference system you are proudest of; scale, latency, cost, and what you specifically owned.
- In a paragraph: how you would serve a 14B-parameter open video model to consumer traffic without lighting money on fire.
- Your expected compensation for this position.
We read every application and reply to all of them.
Apply for this role
About five minutes. We read every application and reply to all of them by email.