ironquill.tech/board

$ cat jobs/staff-software-engineer-systems-infrastructure-agent-evaluat-linkedin3-c321ab49b415.json

Staff Software Engineer, Systems Infrastructure - Agent Evaluation

Linkedin3·Worldwide·Mountain View, US·senior
llm
Apply on smartrecruiters → Get AI match score →
LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We're also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that's built on trust, care, inclusion, and fun – where everyone can succeed. Join us to transform the way the world works. This role will be based in Mountain View, CA. At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. LinkedIn’s Core AI is building the Evaluation Operating System (EOS), a foundational Agent Evaluation platform that defines how all AI agents and GenAI products at LinkedIn are measured, evaluated, and continuously improved in production. This is a brand-new, industry-defining problem space with no established playbook, focused on evaluating multi-step, non-deterministic, and personalized AI systems where traditional metrics and testing approaches fall short. EOS acts as the central intelligence layer for AI quality, combining large-scale data pipelines, evaluator models (e.g., LLM-as-a-judge, reward models), and real-time production monitoring to understand how AI systems behave, where they fail, and how to improve them. The platform includes capabilities like synthetic data generation, adversarial testing, golden dataset management, recursive Self Improving Agents and live “agent arena” experimentation frameworks (champion/challenger testing) to measure performance across multiple dimensions of quality. This platform also is responsible for tracing infras

Similar remote roles

Senior Staff Engineer - DevOps Engineer
Nagarro1 · Worldwide · senior
Senior Software Engineer, RL Environments
Pareto Ai · US · senior
AI Solutions Engineer
leega consultoria · LATAM · mid
Lead AI Engineer
bridge it · APAC · senior
Principal Data Scientist - Omaha, Ne.
mutual of omaha · US · senior
AI Engineer
huzzle · EU · mid
Litigation Lawyer - Legal AI Reviewer
mercor · UK · mid
Data Scientist (AI Data & LLM Specialist)
eclipse laboratories · Worldwide · mid