$ cat jobs/site-reliability-engineer-xpertdirect-8aff22372f48.json
Site Reliability Engineer
Listed as a remote role based in European Union.
Site Reliability Engineer Remote (Europe) – Company based in Dublin, Ireland Enterprise Software | Cloud Infrastructure | Site Reliability Engineering Our client, a fast-growing Enterprise Software company headquartered in Dublin, is looking for a Site Reliability Engineer to help scale and operate the cloud platform that powers mission-critical applications for customers across Europe. What You'll Be Working On • Improving the reliability, availability, and scalability of a multi-tenant cloud platform • Operating and optimising Kubernetes environments supporting production workloads • Managing cloud infrastructure using Terraform and Infrastructure as Code best practices • Building comprehensive observability across applications and infrastructure using Prometheus, OpenTelemetry, and Grafana • Automating operational workflows to reduce manual intervention and improve platform resilience • Developing monitoring, alerting, and incident response capabilities that minimise customer impact • Collaborating with software engineers to improve application performance and production readiness • Participating in incident management, root cause analysis, and continuous service improvement initiatives • Driving reliability engineering best practices throughout the engineering organisation • Supporting platform growth while maintaining high levels of security, availability, and operational efficiency Experience Required • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, or DevOps • Strong commercial experience managing Kubernetes in production environments • Experience with AWS cloud infrastructure and cloud-native architectures • Advanced Terraform and Infrastructure as Code experience • Hands-on experience implementing monitoring and observability using Prometheus and OpenTelemetry • Strong understanding of distributed systems, networking, Linux, and cloud operations • Experience supporting high-availability production environments with demanding uptime requirements • Passion for automation, reliability engineering, and operational excellence Nice to Have • Experience with service meshes such as Istio or Linkerd • Knowledge of GitOps practices using Argo CD or Flux • Experience with distributed tracing and performance optimisation • Familiarity with OpenSearch, Elasticsearch, or cloud logging platforms • Experience implementing SLOs, SLIs, and error budgets • AWS or Kubernetes certifications are an advantage Show more Show less Seniority level Mid-Senior level Employment type Full-time Job function Engineering Industries Software Development
Similar remote roles
DevOps Engineer (all genders)
Holidu · Worldwide · mid
Site Reliability Engineer, Enterprise Technology Services
Apple · Worldwide · mid
Sr Dev-Ops Engineer
Renesaselectronics · Worldwide · senior
DevOps / Platform Engineer
Veocareers · Worldwide · mid
DevOps Engineer
Playtech · Worldwide · mid
SR DevOps Azure
codea it · US · senior
DevOps Engineer
Jobsforhumanity · Worldwide · mid
Data Engineer
Betsson Group · Worldwide · mid