ironquill.tech/board

$ cat jobs/principal-iii-sre-herbalife-87b817ed6370.json

Principal III, SRE

herbalife·Canada·senior
awsgcpazuresre
Apply on himalayas → Get AI match score →
Overview POSITION SUMMARY: The SRE Principal Engineer III role is responsible for leading, designing, and implementing robust Site Reliability Engineering (SRE) practices to ensure high availability, scalability, and resilience of critical business systems and applications. The SRE Principal Engineer III will focus on improving system reliability through automation, monitoring, and performance tuning, working closely with development and operations teams to foster a culture of continuous improvement and operational excellence.The SRE organization spans key disciplines including:• SRE Engineering• Deployment Automation• Incident Response and Postmortem Analysis• Observability and MonitoringOperating in a fully remote capacity, this role will drive the adoption of best practices in multi-cloud and hybrid-cloud platforms, managing services from major cloud providers like Microsoft Azure, Amazon AWS, Oracle OCI, Google GCP, and Alibaba Cloud. The SRE Principal Engineer III will focus on automation, incident management, performance monitoring, and optimizing infrastructure to support scalable, reliable systems. The position will also be responsible for fostering collaboration between development, operations, and security teams to streamline system operations across the organization. DETAILED RESPONSIBILITIES/DUTIES: ● Lead the implementation and optimization of SRE practices, ensuring system reliability, performance, and scalability.● Architect and maintain automation for infrastructure provisioning, deployment, and incident response.● Establish and enforce SLOs (Service Level Objectives) and SLIs (Service Level Indicators) for key services.● Collaborate with development teams to design and deliver reliable software systems, ensuring that production environments are optimized for uptime and performance.● Create and maintain monitoring, alerting, and observability solutions to provide real-time insights into system health and performance.● Respond to production incidents,

Similar remote roles

Lead SRE
66degrees · Worldwide · senior
DevOps Engineer - AI Model Evaluator
mercor · Worldwide · mid
Analista de Engenharia de Software
dxc technology · LATAM · mid
Senior Application Developer
fsr llc · US · senior
Tech Lead (TypeScript)
orcrist technologies · EU · senior
Director, Cloud Operations
Elastic · US · mid
Data Scientist H/F
Aneo · EU · mid
Senior Software Engineer -Golang (Security Services Integrations)
Elastic · US · senior