ironquill.tech/board

$ cat jobs/applied-ai-engineer-kernel-performance-etched-372de4dbc12d.json

Applied AI Engineer, Kernel Performance

Etched·Worldwide·San Jose·mid
Apply on ashby → Get AI match score →
About Etched Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history. Job Summary Every model release presents a new opportunity to push the frontier on kernel engineering. Future performance breakthroughs will come from AI systems that can understand model architectures and hardware, run thousands of experiments, learn from compiler and profiler feedback, and discover the most performant implementations faster than the best engineers. You will build that system. Your mandate is to build AI systems that autonomously turn newly released model architectures into correct, production-ready implementations optimized for Etched hardware. These systems should explore broader design spaces, learn from every experiment, and reach peak performance faster than any traditional kernel-development workflows. Etched offers a uniquely tight research loop: proprietary hardware, compiler, runtime, kernels, production workloads, and dedicated in-office compute under one roof. You will teach models using proprietary performance signals, iterate on their proposals, and make every experiment improve both the performance optimization system and the hardware it runs on. Key Responsibilities - Own the system that turns new model architectures into verified, production-ready kernels and model mappings. - Build agents that understand Etched hardware, design experiments, generate implementations, compile and profile them, diagnose bottlenecks, and iterate with our teams, to the limits of model autonomy. - Design evals covering correctness, numerical stability, latency and efficiency. - Turn profiler traces, simulation, hardw