ironquill.tech/board

$ cat jobs/evaluation-engineer-elicit-fc197041735c.json

Evaluation Engineer

Elicit·Worldwide·Oakland, CA (or remote within US timezones)·mid
Apply on ashby → Get AI match score →
ABOUT ELICIT Elicit is an AI research platform that uses language models to help researchers figure out what's true and make better decisions, starting with common research tasks like literature review. What we're aiming for: 1. Elicit radically increases the amount of good reasoning in the world. - For experts, Elicit pushes the frontier forward. - For non-experts, Elicit makes good reasoning more affordable. People who don't have the tools, expertise, time, or mental energy to make well-reasoned decisions on their own can do so with Elicit. 2. Elicit is a scalable ML system based on human-understandable task decompositions, with supervision of process, not outcomes. This expands our collective understanding of safe AGI architectures. Visit our Twitter https://twitter.com/elicitorg to learn more about how Elicit is helping researchers and making progress on our mission. THE MISSION OF ELICIT EVALS Some orgs build evals to warn us about dangerous capabilities. Some build evals to understand trends and predict where models are heading. Some build evals to hill-climb toward models that users will like more. At Elicit, we're after something different. We want to understand, and hill-climb toward, models that help us make better decisions. This is harder than "what will users like better." Decision support is difficult to evaluate, and users' knee-jerk reactions don't always track with what actually helps them decide. Because it's hard, and because the sales pitch is more complicated, few are doing it well. If we get this right, we have a real shot at pushing AI toward better decision-making, both inside Elicit and beyond. WHY WE'RE HIRING FOR THIS ROLE We need someone to own the technical foundation of our auto-evaluation systems. Our evals are much slower than they need to be, and our interfaces aren't built for the range of people who rely on them: ML engineers iterating on models, product managers monitoring quality, and customers assessing how much to trust a resul