ironquill.tech/board

$ cat jobs/sr-software-engineer-ai-reliability-mixmode-c66441c89e2c.json

Sr. Software Engineer-AI Reliability

mixmode·US·United States·senior
Apply on himalayas → Get AI match score →
MixMode is a leading provider of AI-powered cybersecurity solutions at scale, pioneering a patented third-wave, context-aware AI approach that automatically learns and adapts to dynamic environments. The MixMode platform delivers self-supervised, real-time threat detection for known and unknown threats across cloud, hybrid, and on-premises environments. Large organizations with big data workloads – including those in enterprise, critical infrastructure, US Department of War and US Intelligence Community – trust MixMode to defend their most important assets. Backed by PSG and Entrada Ventures, MixMode is headquartered in Santa Barbara, California. Learn more at www.mixmode.ai . Job Summary: We are hiring a Senior Software Engineer, based in the United States, to enhance the reliability, performance, and scalability of our production AI systems. We value clear thinking, incremental improvement, and engineers with real production incident experience. This role focuses on understanding, refining, and strengthening existing distributed services across application, database, and container orchestration layers. You will collaborate with ML researchers to make our systems more robust, maintainable, flexible, and scalable. This is a remote position but requires the ability to travel to the MixMode headquarters in Santa Barbara for in person meetings a few times a year. What you’ll be doing (responsibilities): Own the reliability, performance, and operational health of production AI services Refactor and harden existing systems to improve resilience, clarity, and maintainability Diagnose and resolve issues across distributed services, data pipelines, and storage layers Design and implement monitoring, alerting, and debugging tools for high-availability systems Partner with researchers and engineers to productionize predictive systems at scale Establish best practices for testing, deployment, capacity planning, and incident response Contribute to incident response and postmort