ironquill.tech/board

$ cat jobs/staff-engineer-inference-optimizations-digitalocean-5a74d481d983.json

Staff Engineer, Inference Optimizations

DigitalOcean·Worldwide·Austin·senior
node.js
Apply on greenhouse → Get AI match score →
Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world. DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be the Inference Cloud of choice for digitally native companies and you will help ensure we can offer the industry-leading performance for our inference services. You will be responsible for the architectural decisions that maximize throughput and minimize latency for the world’s most advanced large models. As an IC leader, you will act as a force multiplier for the engineering organization, solving the most complex bottlenecks in memory bandwidth and compute utilization while guiding the technical roadmap for our high-performance inference fleet. What You’ll Do: Performance Architecture: Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, ensuring our infrastructure extracts maximum value from every TFLOP. Deep-Dive Optimization: Engineer solutions for complex performance issues, including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters. Technological Innovation: Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape. Some examples of projects you may work on: Improving batch size performance using AMD's AITER library for AMD MI355X - identify and tune AITER's CK (composable kernel) or ASK (assembly) to optimize FP8 / BF16 Identify kernel fusion

Similar remote roles

Staff Engineer, Inference Optimizations
DigitalOcean · Worldwide · senior
Staff Engineer, Inference Optimizations
DigitalOcean · Worldwide · senior
Staff Engineer, Inference Optimizations
DigitalOcean · US · senior
Freelance Full-Stack Web App Developer
mindrift · APAC · mid
AI & Data Engineer (m/w/d)
PeakSoft GmbH · Worldwide · mid
Flutter Full Stack Mobile Developer (FinTech).
ebizon · APAC · mid
Freelance Full-Stack Web App Developer
Mindrift · Worldwide · mid
Fullstack Developer
Schrder · Worldwide · mid