ironquill.tech/board

$ cat jobs/senior-cloud-site-reliability-engineer-nice-34af76bfe5ee.json

Senior Cloud Site Reliability Engineer

NICE·APAC·India - Pune·senior
kubernetesterraformhelmawssreci/cdgithub actionsjenkinsprometheusgrafana
Apply on greenhouse → Get AI match score →
At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer you the ultimate career opportunity that will light a fire within you. So, what's the role all about? NICE is looking for a Senior Site Reliability Engineer to join our core Reliability Engineering team, responsible for ensuring the scalability, reliability, and performance of mission-critical systems and observability platforms across multiple environments and regions. This role is ideal for someone who thrives in fast-paced environments, enjoys automation, and has a strong background in cloud-native operations, observability stacks, and incident management. You'll collaborate closely with product, platform, and development teams to drive reliability-first design, proactive observability, and operational excellence. How will you make an impact? Reliability & Performance · Design and implement scalable, reliable, and resilient systems across hybrid or multi-cloud environments (primarily AWS/EKS/ECS/Lambda) · Drive improvements in system uptime, latency, and overall service health metrics (SLOs, SLIs, SLAs) Automation & Infrastructure as Code · Build and manage infrastructure automation using Terraform, Helm, and Kubernetes · Improve CI/CD pipelines using Jenkins, GitHub Actions, ensuring safe and automated rollouts, monitoring, and rollbacks Observability & Monitoring · Own and enhance the observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Mimir, etc.) · Define and implement SLOs and error budgets; enable development teams to monitor and act on reliability metrics Incident Management · Lead major incident response, root cause analysis (RCA), and blameless postmortems · Partner with product teams to define and enforce operational readiness standards before production releases Security & Compliance · Ensure platform-level security a

Similar remote roles

Specialist Cloud Site Reliability Engineer
NICE · APAC · mid
DevOps Engineer
Playtech · Worldwide · mid
DevOps Engineer
Reactive Technologies Limited · Worldwide · mid
DevOps Engineer - Remote, Latin America
bluelight consulting · Worldwide · mid
Staff/Principal DevOps Engineer, AI Inference
Lila Sciences · US · senior
DevOps / SRE Engineer
Avangarde Consulting · Worldwide · mid
DevOps Engineer
COLIBRIX ONE · Worldwide · mid
DevOps-Engineer (m/w/d)
IT42morrow IFT GmbH · Worldwide · mid