ironquill.tech/board

$ cat jobs/software-engineer-model-evaluation-and-improvement-benchling-7db3ef1af61d.json

Software Engineer, Model Evaluation and Improvement

Benchling·US·San Francisco, CA·mid
llm
Apply on ashby → Get AI match score →
We are rebuilding biotech for the AI era. When a breakthrough is delayed, the world waits. Getting a molecule from discovery to patients, or a crop from lab to field, involves thousands of slow, manual, disconnected steps. AI has the potential to change this, compressing decades of R&D work into years. But that only happens when clean, structured scientific data and AI are built into how science gets done. Benchling is the AI platform for biotech R&D. Scientists use Benchling to design experiments, capture structured data, and run AI agents and models directly in their workflows. Over 200,000 scientists around the world trust Benchling to power their most important work, from academic labs to Sanofi, Moderna, and more than half of the world's top 50 biopharma. We’re building an AI scientist for our customers. We can’t do that if we haven’t built the muscle ourselves. AI fluency is the foundation we build on; it's core to how we work, and we're committed to helping every new hire integrate it into their day-to-day. As part of our interview process, you'll complete a brief AI-focused exercise or discussion so we can understand how you think about and use AI to drive impact in your role. Feel free to reference any tools, platforms, or workflows you use today. ROLE OVERVIEW We’re a team focused on making frontier AI models better at science. LLMs know an extraordinary amount of biology, but there’s still a large gap in reasoning for the real-world problems scientists face every day. We recently published some of our work here https://www.benchling.com/blog/can-llms-work-in-the-wet-lab. You’ll build the datasets, evaluations, and systems that help close that gap. You’ll work with scientists to turn complex scientific work into rigorous tasks that models can learn from and be evaluated against. You’ll partner with leading AI labs to understand where models fail and how to improve them. This is an early and rapidly evolving area. You’ll work at the intersection of software

Similar remote roles

Technical Product Support Specialist, Enterprise
triple whale · US · mid
Software Engineer [New Grads Welcome]
mechanical orchard · Canada · mid
Product Designer
Bland · US · mid
Private Equity & Investment Funds Attorney
micro1 · Worldwide · mid
Prompting and Context Engineer
1mind · Worldwide · mid
Senior Manager, Information Security
Benevity · Canada · senior
Senior Manager, Information Security
Benevity · Worldwide · senior
Don’t see what you’re looking for?
RelationalAI · Worldwide · mid