ironquill.tech/board

$ cat jobs/senior-principal-ai-engineer-cerence-92de785a58f5.json

Senior Principal AI Engineer

cerence·US·United States·senior
node.jskubernetes
Apply on himalayas → Get AI match score →
A Moving Experience. What You Will Work On Design andoperatedistributed training systems for large neural networks(autoregressive, diffusion,State Space Modelsetc.)across GPU clusters Optimisemulti‑node, multi‑GPU execution tomaximizethroughput andutilization Diagnose&resolve bottlenecks across compute, memory, and network Improve training stability and fault tolerance at scale Partner with research and applied ML teams to productionizelarge‑modeltraining pipelines Core Responsibilities Distributed Training Infrastructure Build andoptimizeGPU cluster orchestration using: Slurm Kubernetes Ray RunAI Ensure efficient scheduling, isolation, and fairness across training workloads Communication & Networking Optimizeand debug distributed communication using: NCCL RDMA InfiniBand NVLink Minimizenetworking bottlenecks that dominateend‑to‑endtraining time Training Frameworks Scale large-model training using: PyTorchDistributed Megatron‑LM DeepSpeed Ownmulti‑nodelaunch configurations, failure recovery, and performance tuning Memory & Performance Optimization Apply advanced memory optimization techniques: Activation checkpointing ZeRO(Stage 1–3) and offload strategies Balance compute, memory, and communication to push model size and batch scale What Success Looks Like GPUutilizationconsistently stays high (>80–90%) Training scales cleanly from single node to dozens or hundreds of GPUs Communication overhead is minimized and predictable Large training jobs run stably for days or weeks without failure New models can be trained faster, larger, and more reliably than before Required Experience & Skills Strongly Required Deephands‑onexperience with distributed systems or ML systems Experience runninglarge‑scaleworkloads on GPU clusters Production experience withPyTorchdistributed training Strong understanding of parallelism strategies (data, tensor, pipeline parallelism) Low‑levelunderstanding of GPU communication and networking Critical Technical Skills GPU orchestration:Slurm, Kub

Similar remote roles

Senior Machine Learning Engineer
BEES · LATAM · senior
Software QA Engineer - Manual (Linux) - Remote
jitterbit · APAC · mid
Technical Support Engineer (GPU Clusters) - US Weekends
together ai · US · mid
Cybersecurity Engineer (SOAR) [JOB ID 20260804]
phoenix cyber · US · mid
Full Stack Developer (React/Node.js)
Elixirr Digital · Worldwide · mid
Senior DevOps / Infrastructure Engineer
Category Labs · US · senior
Senior Fullstack (MERN) Developer
Proxify · EU · senior
Technical Lead - Full Stack
mindplus pvt ltd · Worldwide · senior