ironquill.tech/board

$ cat jobs/mla-senior-site-reliability-engineer-sre-kubernetes-software-mind-0e191437dfe5.json

[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes

software mind·EU·Poland·senior
node.jsjavakubernetesdevopssreci/cdprometheusgrafana
Apply on himalayas → Get AI match score →
Project – the aim you'll have We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work. Position – how you’ll contribute Support the deployment, operation, and reliability of production services running on Kubernetes. Monitor service health and investigate production incidents across distributed applications. Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements. Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams. Support CI/CD, GitOps-based deployments, observability, and production monitoring. Work within a client-directed backlog and established priorities. Expectations – the experience you need 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering , or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services. 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene Splunk experience for log aggregation, search, and production troubleshooting Prometheus and Grafana experience, specifically building alert rules a

Similar remote roles

ANALISTA DEVOPS JR | HOME OFFICE
solinftec · LATAM · junior
Middle+/Senior DevOps/System Engineer (SE Platform Ops Sun)
Softswiss · Worldwide · senior
Senior Java Fullstack Developer (ST)
alter solutions · EU · senior
Backend Software Engineer
Northflank · EU · mid
Software Engineer (TelCo)
LanceSoft Europe · Worldwide · mid
Senior ML Engineer - Offline Team
Voodoo · Worldwide · senior
Software Engineer Telco
ISG International Service Group · Worldwide · mid
Senior Full-Stack Software Engineer
Transperfect · Worldwide · senior