$ cat jobs/mla-senior-site-reliability-engineer-ai-experience-framework-software-mind-7abb7cdd36c9.json
[MLA] Senior Site Reliability Engineer - AI Experience Framework
Project – the aim you'll have We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work. Position – how you’ll contribute Own Kubernetes deployment and operational health for Framework services, including scaling, rollout/rollback strategy, and resource tuning Build and maintain production observability - Grafana dashboards and Prometheus alerting rules - across the SSR runtime and the Glide platform layer Diagnose and resolve Node.js production incidents: event-loop stalls, heap growth, V8 isolate exhaustion, and isolate-pool scheduling issues under concurrent versioned traffic (vN/vN-1) Diagnose and resolve JVM production incidents on the Glide/Java side: GC pressure, thread dumps, and platform-service latency Own incident response for the team: runbooks, on-call rotation, postmortems, and paging hygiene Drive CI/CD and infrastructure-as-code for Kubernetes manifests/Helm and deployment pipelines Partner with the framework engineering team to identify reliability gaps before they become incidents - capacity planning, load testing, chaos/failure-injection where useful Represent production reliability concerns in architecture reviews for new framework capabilities Expectations – the experience you need Production operations/SRE experience, including hands-on Kubernetes deployment, scaling, and incident response Direct operational experience troubleshooting Node.js in production: reading heap snapshots and CPU pr
Similar remote roles
Sr Dev-Ops Engineer (Altium)
Renesaselectronics · Worldwide · senior
Software Engineer III - Python
JPMorganChase · Worldwide · mid
Back-End Developer
talent sam · Worldwide · mid
Cloud Engineer - Fully Remote | Upto $85/hr
mercor · Worldwide · mid
DevOps Engineer - AI Model Evaluator
mercor · EU · mid
Backend Developer Semi Senior | Java + AWS + OpenShift (Remoto)
babel · LATAM · senior
Staff Engineer I (AI, Java)
cotiviti · US · senior
DevOps Engineer - AI Model Evaluator
mercor · Worldwide · mid