ironquill.tech/board

$ cat jobs/product-reliability-engineer-us-opsmill-4a8f30d57d24.json

Product Reliability Engineer | US

Opsmill·Canada·US/Canada·mid
kubernetes
Apply on ashby → Get AI match score →
Shipping infrastructure software is only half the job. The other half is making it work in environments you don’t control—across messy reality, strict security constraints, and endless platform variations. The difference between a good product and a trusted one is how quickly you can diagnose issues and how effectively you prevent them from happening again. At OpsMill, we're building Infrahub, a schema-driven infrastructure source of truth that helps teams unify data and scale automation reliably. Our customers deploy Infrahub on-prem, which means reliability is a product feature, not just an operational concern. When something breaks in the field, it's not just a support ticket—it's a signal about what we need to fix, test, or instrument better. WHY THIS ROLE EXISTS We need someone who can operate in both worlds: diving deep on gnarly customer escalations while systematically eliminating entire classes of problems. You'll be the crucial bridge between "customer is blocked right now" and "this type of issue can't happen again." You'll build the diagnostics, tests, and automation that turn on-prem deployment chaos into predictable, debuggable, fixable reliability. WHAT YOU'LL BE DOING - Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. - Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements - Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster - Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integrati

Similar remote roles

Senior Full Stack Engineer, Acquisition
CookUnity · LATAM · senior
Site Reliability Engineer (SRE/ DevOps) - Engineering Productivity
Aristanetworks · Worldwide · mid
DevOps Engineer
mhymatch · APAC · mid
Data Engineer
SumUp · EU · mid
Staff Engineer, Big Data
Nagarro1 · Worldwide · senior
Senior Software Engineer, Infrastructure - Compute Platform
Coinbase · Worldwide · senior
Senior Software Engineer, Infrastructure - Compute Platform
Coinbase · Worldwide · senior
Associate Staff Engineer, ERP Dynamics 365 (POS Developer)
Nagarro1 · Worldwide · senior