ironquill.tech/board

$ cat jobs/mlops-engineer-whitecircle-e042d6bb03c4.json

MLOps Engineer

Whitecircle·Worldwide·Paris·mid
kubernetesllmdatadog
Apply on ashby → Get AI match score →
TLDR: We're looking for an MLOps Engineer to sit at the boundary between Research and Production. You'll own the infrastructure that takes a trained model and makes it production-safe: rollout pipelines, quality and latency gates, canary deployments, and the dashboards that decide whether a release ships or rolls back. About us White Circle https://whitecircle.ai/ is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do. We automatically test, enforce, and continuously improve these policies at scale. - We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others - We process over 100M+ API calls every month - We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model We’re a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built – you’re the one we need. You will: - Integrate new text and multimodal models into our serving paths and verify they behave correctly under production-like traffic. - Build and maintain rollout pipelines for frequent model releases. - Create smoke, quality, and performance gates for model promotion. - Operate local and cluster GPU deployments on Kubernetes. - Build dashboards for latency, throughput, queue depth, GPU usage, fallback rate, and quality drift. - Run A/B and canary rollouts for model, prompt, routing, and serving config changes. - Debug production issues across model config, tokenizer, serving API, router, queue, Kubernetes, GPU runtime, and CI jobs. - Optimize serving cost and reliability across mixed GPU capacity. Who you are - Experience with an inference serving engine such as SGLang, vLLM, Dynamo, or TensorRT-LLM, and a working un

Similar remote roles

DevOps Engineer
Jobsforhumanity · Worldwide · mid
Senior AI Engineer
Temus · APAC · senior
Senior Full Stack Engineer, Acquisition
CookUnity · LATAM · senior
Machine Learning Engineer
Oak Tree Software · Worldwide · mid
FinOps Engineer
Sutherland · Worldwide · mid
Associate Distinguished Engineer (Agentic AI Architect)
Nagarro1 · APAC · senior
Engineering Manager, Applied AI & Machine Learning Engineering
SPD Technology · Worldwide · mid
AI Engineer
Chabre · Worldwide · mid