$ cat jobs/senior-director-inference-products-and-optimizations-digitalocean-79bbc24785ff.json
Senior Director, Inference Products and Optimizations
Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world. Our Inference Engine organization is seeking an experienced Senior Director of Engineering to lead a high-performing engineering team building and scaling our Large Language Model (LLM) inference products across the control plane, model optimization, and model architecture layers. This organization sits at the heart of DigitalOcean's mission to bring our signature simplicity to optimized LLM inference. In this role, you will own DigitalOcean's inference product suite — Serverless Inference, Dedicated Inference, Inference Router, Batch Inference, and Multimodal Inference — along with the model optimization and architecture stack that underpins them. You'll be responsible for delivering robust, cost-efficient systems that serve millions of users globally at scale and high performance. What You'll Do: Team Leadership & Development: Recruit, mentor, and coach engineers on the team, fostering a culture of ownership, technical excellence, and continuous improvement. Build Performant and Scalable Inference Products: Work with Product teams to define and execute on the Product roadmap for all of DigitalOcean’s Inference Products - including Serverless Inference, Dedicated Inference, Inference Router, Batch Inference and Multimodal Inference Inference Optimizations and Model Architecture: Lead the design and evolution of our inference serving stack, driving deep technical strategy across vLLM, SGLang, and LLM-D to optimize throughput, latency, and GPU utilization at scale. Architect the model-s
Similar remote roles
Machine Learning Engineer
Persistent Systems · Worldwide · mid
Founding Product Engineer
clera · UK · mid
Chatbot Developer (WhatsApp, Telegram, Discord) - Freelance
Mindrift · Worldwide · mid
Chatbot Developer (WhatsApp, Telegram, Discord) - Freelance
Mindrift · Worldwide · mid
Chatbot Developer (WhatsApp, Telegram, Discord) - Freelance
Mindrift · Worldwide · mid
Chatbot Developer (WhatsApp, Telegram, Discord) - Freelance
Mindrift · Worldwide · mid
Chatbot Developer (WhatsApp, Telegram, Discord) - Freelance
Mindrift · Worldwide · mid
Chatbot Developer (WhatsApp, Telegram, Discord) - Freelance
Mindrift · Worldwide · mid