ironquill.tech/board

$ cat jobs/helix-ai-engineer-training-performance-figure-516ca81f166f.json

Helix AI Engineer, Training Performance

Figure·Worldwide·San Jose, CA·mid
node.jselasticsearch
Apply on greenhouse → Get AI match score →
Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA. Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced AI Training Performance Engineer to take our model training to the next level. This role is focused on improving distributed training frameworks for large scale model training, optimizing GPU kernels, exploring the relative gains of different accelerator types and co-designing our models to maximize utilization of our hardware. Responsibilities Optimize training performance for a 100B+ parameter models across 100k+ GPUs. Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling. Write and optimize custom kernels (Triton/CUDA) Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs Optimize data loading and preprocessing pipelines so I/O never gates the accelerators Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing) Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of

Similar remote roles

Sr Dev-Ops Engineer (Altium)
Renesaselectronics · Worldwide · senior
Sr Dev-Ops Engineer
Renesaselectronics · Worldwide · senior
Senior Software Engineer, Security Solution — Triage & Investigation
Elastic · US · senior
Chief of Staff to the CTO
fundraise up · EU · senior
Senior Full Stack Web Entwickler (w/m/d)
CONTENS · Worldwide · senior
Fullstack Developer
fundraise up · Worldwide · mid
Senior MEAN Developer – Full Remote
gopro consultancy group ltd · US · senior
Senior MEAN Developer - Full Remote
gopro consultancy group ltd · US · senior