$ cat jobs/staff-principal-devops-engineer-ai-inference-lila-sciences-39987d7be42b.json
Staff/Principal DevOps Engineer, AI Inference
Your Impact at LILA The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency. What You'll Be Building GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking Observability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints CI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI AWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances
Similar remote roles
DevOps Engineer (all genders)
Holidu · Worldwide · mid
Sr Dev-Ops Engineer
Renesaselectronics · Worldwide · senior
DevOps Engineer
nagarro · Worldwide · mid
DevOps / Platform Engineer
Veocareers · Worldwide · mid
DevOps / SRE Engineer
Avangarde Consulting · Worldwide · mid
DevOps Engineer
Playtech · Worldwide · mid
Senior DevOps Engineer
MWDN · Worldwide · senior
DevOps Engineer
Jobsforhumanity · Worldwide · mid