ironquill.tech/board

$ cat jobs/machine-learning-infrastructure-tech-lead-reducto-9357c3817336.json

Machine Learning Infrastructure Tech Lead

Reducto·US·San Francisco Office·senior
node.jsml
Apply on ashby → Get AI match score →
ABOUT REDUCTO Reducto is the agentic document platform for leading AI teams who demand enterprise performance at scale. We provide a comprehensive toolkit for working with documents the way a human would, combining custom in-house and leading frontier models to power efficient and accurate document workflows. We’ve grown rapidly, increasing revenue 8x year over year and partnering with hundreds of companies, from leading AI teams like Harvey, Vanta, and Scale, to enterprise customers across FAANG and top trading firms. Reducto has raised over $100M from world-class investors including a16z, Benchmark, and First Round Capital. THE OPPORTUNITY As our ML Infrastructure Tech Lead, you'll own the systems that make high-performance model training and inference possible at Reducto. This is a deeply hands-on role: roughly 80% of your time will be spent building, debugging, and optimizing our infrastructure. The remaining 20% will focus on setting technical direction - identifying bottlenecks, planning our infrastructure roadmap, and helping the ML and Platform teams make strong architectural decisions. You'll work across the stack, from model-serving kernels and GPU utilization to distributed systems and Kubernetes. We're looking for someone with the experience and judgment to lead ambiguous, high-impact infrastructure projects while remaining close to the code. This is a fully in-person role at our San Francisco office. WHAT YOU'LL DO - Own the technical direction and roadmap for Reducto's ML infrastructure. - Build and maintain our training and inference stack, balancing fast experimentation with high-performance production serving. - Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference. - Design systems for reliable multi-node, multi-GPU training and inference. - Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency. - Develop benchmarks that identify bottlenecks and gu

Similar remote roles

Backend Developer
Solflare · Worldwide · mid
Senior Backend Engineer
Voice AI Space · Worldwide · senior
Full-Stack Engineer, Office of the CEO
Arago · Worldwide · mid
AI Research Engineer - ML Engineering
Helsing · Worldwide · mid
Machine Learning Engineer
Dashlane · Worldwide · mid
AI Full Stack Developer
Sbtglobalinc · Worldwide · mid
Machine Learning Infra Engineer
Reducto · US · mid
Senior Backend Engineer - AI Platform (f/m/d)
Contentful · UK · senior