ironquill.tech/board

$ cat jobs/staff-ml-engineer-buildkite-59d500b3b454.json

Staff ML Engineer

Buildkite·Worldwide·ANZ Region·senior
ml
Apply on greenhouse → Get AI match score →
Push a one-line fix. Then watch CI grind through forty minutes of tests, ninety-five percent of which never had a chance of touching what you changed. You already know the handful that mattered. The test suite doesn't — so it runs everything, every time, just in case. That "just in case" is the most expensive habit in software delivery. Every engineering team pays it, because the alternative — knowing which tests actually matter for a given change — has been too hard to get right. We're building the team that gets it right. This role sits at the centre of it. 🔧 The problem worth solving Test Engine already ingests billions of test runs. We can see the tests, the code underneath them, and how the two move together — at a scale very few people ever get to work with. The raw material for the answer is already here. Nobody's turned it into predictions yet. That's the step to take: for a given change, work out the slice of tests most likely to fail, and run only those. Get it right and teams stop re-running what hasn't changed, and spend that time where it counts — like fixing the two percent of tests most likely to break. It's a genuinely difficult ML problem — sparse signal, cold-start on new repos, generalising across languages and frameworks, and latency tight enough to sit in the critical path. It's also close to a blank page. There's no ML org above you setting the direction — you'd set it. And not alone: we've just hired another ML engineer, so there's someone to think out loud with from day one. 🚀 What you'll own Machine learning in Test Engine, end-to-end — the strategy, the architecture, and the models running in production. That means shaping the whole path: pulling features out of code changes and test history, training and evaluating models, building the serving layer that keeps predictions fast, and closing the loop so the system keeps improving. You'd make the trade-offs that matter — accuracy versus latency, what happens when confidence is low — and bui

Similar remote roles

Director, Global Advisory EMEA
biocatch · EU · mid
Data Engineer
Wabtec · Worldwide · mid
Senior Python Backend Developer / ML Engineer (IR-535)
See posting · EU · senior
Senior ML Engineer with Python (IR-536)
See posting · EU · senior
STA0031 .Net Developer
See posting · EU · mid
Senior Threat Researcher- Threat Detection Engineer
sophos · APAC · senior
Senior Machine Learning Engineer - Forecasting
lifelancer · US · senior
Payroll & Admin Officer
Akur8 · Worldwide · mid