$ cat jobs/ml-systems-performance-engineer-mfu-higgsfieldai-728eb9c019c1.json
ML Systems Performance Engineer (MFU)
Why work at Higgsfield AI? Higgsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands. We're building at the absolute frontier of AI-powered video creation and next-generation creative tools. Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like. What you will do • Profile end-to-end training runs and identify bottlenecks across compute, memory, communication, storage, and orchestration. • Define, measure, and improve MFU, tokens/sec/GPU, scaling efficiency, training goodput, and GPU uptime. • Optimize distributed training and model-sharding strategies, including data, tensor, pipeline, context, and expert parallelism. • Improve collective communication through topology-aware placement and compute/communication overlap. • Develop or integrate optimized CUDA and Triton kernels • Optimize data loading, preprocessing, sequence packing, and checkpointing so that I/O does not leave accelerators idle. • Diagnose distributed hangs фтв performance regressions. • Improve fault tolerance for long-running training jobs. What we are looking for • Strong experience running and optimizing multi-GPU or multi-node training. • Experience with PyTorch Distributed or an equivalent training framework. • Understanding of GPU architecture, including memory hierarchy, Tensor Cores • Understanding of collective communication, cluster topology, and distributed-training bottlenecks. • Experience with distributed parallelism technologies such as FSDP, DeepSpeed, Megatron-LM, TorchTitan, or similar. • Ability to debug complex performance and reliability problems across multiple layers of the training stack. Nice to have • CUDA, Triton or GPU-kernel development experience. • Experience with NCCL, MPI, UCX, RDMA, InfiniBand, RoCE,
Similar remote roles
Research Engineer (Agentic Models)
JetBrains · Worldwide · mid
Senior ML Engineer - Offline Team
Voodoo · Worldwide · senior
Software Engineer, DGX Cloud AI Infrastructure
nvidia · US · mid
DESENVOLVEDOR NODE.JS SR
stefanini · LATAM · senior
Technical Architect - Node JS
Necsws · Worldwide · senior
ML Engineer (Data Engine)
Higgsfieldai · Worldwide · mid
Senior/Leading Software Engineer (fintech)
TechBiz Global GmbH · Worldwide · senior
Full Stack Developer
axle informatics · US · mid