ironquill.tech/board

$ cat jobs/systems-reliability-engineer-bright-vision-technologies-f0261d190595.json

Systems Reliability Engineer

bright vision technologies·US·United States·mid
sreprometheusgrafana
Apply on himalayas → Get AI match score →
Systems Reliability Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: Systems Reliability Engineer Location: 100% Remote (United States) Position Type: Full-time, Direct W2 Salary Range: $100,000–$150,000 Annually Experience: 6+ years Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. Job Summary We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern. Key Responsibilities Define, instrument, and continually refine service-level objectives (SLOs), service-level indicators (SLIs), and error budgets for critical services, and use those measures to drive concrete engineering and prioritization decisions. Lead incident response and resolution for production issues, acting as a calm and effective incident commander when needed, and ensuring high-quality post-incident reviews that drive lasting improvements. Design and implement comprehensive monitoring, logging, and tracing strategies using Prometheus, Grafana, OpenTeleme

Similar remote roles

DevOps Engineer
Jobsforhumanity · Worldwide · mid
Site Reliability Engineer, Enterprise Technology Services
Apple · Worldwide · mid
Senior Platform Engineer
State Street · Worldwide · senior
DevOps Engineer
Higgsfieldai · Worldwide · mid
Senior Platform Engineer / SRE
Codurance · Worldwide · senior
DevOps Engineer Senior
Undelucram.ro · Worldwide · senior
Middle Platform Automation Engineer
Miratech1 · Worldwide · mid
Platform Engineer
Chronograph (chronograph.pe) · Worldwide · mid