$ cat jobs/site-reliability-engineer-boson-ai-72c917bd8236.json
Site Reliability Engineer
About The Role Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work. Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from “it works” to dependable, observable, and scalable. You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area—networking, cluster scheduling, storage, GPU systems, or AI infrastructure— and the curiosity and judgment to collaborate across the rest. Responsibilities Design, operate, and improve reliable infrastructure for AI training and inference workloads Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements Improve provisioning, configuration management, testing, and deployment automation Help plan cluster growth, capacity allocation, upgrades, and lifecycle management Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards Minimum Qualifications 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role Strong hands-on expertise in at least one of the follo
Similar remote roles
Senior Ruby on Rails Engineer
fifth third bank · UK · senior
Site Reliability Engineer - Dedicated Hosted Runners
GitLab · APAC · mid
Site Reliability Engineer (SRE/ DevOps) - Engineering Productivity
Aristanetworks · Worldwide · mid
Executive Director, AI Infrastructure & Platform Engineering
lifelancer · US · mid
Consultant Ingénieur Avant-Vente & Architecte Technique - F/H/N
Octotechnology · Worldwide · mid
Data Engineer
SumUp · EU · mid
Senior Site Reliability Engineer, Workforce Identity
Coinbase · Worldwide · senior
Senior SRE/DevOps Engineer (Remote, Colombia/ Brazil) - Min.
exceptionly · US · senior