ironquill.tech/board

$ cat jobs/member-of-technical-staff-coding-research-micro1-6ffe1ad064b3.json

Member of Technical Staff, Coding Research

micro1·Worldwide·senior
pythonc++ml
Apply on himalayas → Get AI match score →
Job Title: Member of Technical Staff, Coding Research Job Type: Full-time Location: Remote The Role We are seeking a Member of Technical Staff to help advance the evaluation and development of frontier coding agents. Sitting at the intersection of AI research, software engineering, and model evaluation, you will design the benchmarks, methodologies, and data systems that shape how next-generation coding models are measured and improved. What You'll Do Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, rubrics, and quality standards. Lead end-to-end research initiatives focused on measuring and improving coding model performance across diverse software engineering tasks. Develop high-quality datasets, golden examples, and evaluation protocols that enable reliable assessment of frontier coding systems. Analyze model behavior and failure modes, identifying systematic weaknesses and translating findings into actionable improvements for training and evaluation. Build tooling and infrastructure that support large-scale experimentation, data generation, review workflows, and evaluation pipelines. Establish best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality. Partner closely with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities. Contribute to technical reports, benchmark studies, and client-facing research initiatives that communicate model performance and insights. What We're Looking For Strong software engineering background with expertise in Python, C++, or comparable programming languages. 3+ years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines. Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies. Familiarity with large language models, coding agent

Similar remote roles

Open Source Contributor
micro1 · Worldwide · mid
Machine Learning Engineer
Docusign Women in Sales · Worldwide · mid
Back End Developer
Explore Group · Worldwide · mid
Data Scientist
NICE · UK · mid
AI Engineer
micro1 · Worldwide · mid
Unitary Foundation | Member of Technical Staff, Fault Tolerant Quantum Compilation + Simu…
Unitary Foundation · Worldwide · senior
Senior Machine Learning Engineer
Bjakcareer · Worldwide · senior
Member of Technical Staff, Machine Learning
Bjakcareer · Worldwide · senior