ironquill.tech/board

$ cat jobs/research-scientist-engineer-agentic-systems-whitecircle-3e96b0c4afa2.json

Research Scientist/Engineer (Agentic Systems)

Whitecircle·Worldwide·Paris·mid
llmdatadog
Apply on ashby → Get AI match score →
TLDR: We're looking for a research scientist to build autonomous, large-scale environments that push LLM agents (single and multi-agent) to failure, and study how they actually break. About us White Circle https://whitecircle.ai/ is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do. We automatically test, enforce, and continuously improve these policies at scale. - We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others - We process over 100M+ API calls every month - We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model We’re a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built – you’re the one we need. About the team White Circle's fundamental research team works on the science of how AI systems fail: where agents break, why misalignment and unsafe behaviours emerge, and how to catch them before they reach the real world. We build the evals, benchmarks, environments, and tooling that empirically study the most pressing AI safety concerns — some of which become the guardrails shipped in our products, and some of which become public writeups. You will - Build adversarial environments for agents: complex, uncertain settings that sit on the boundary of agent capability and alignment, where failure is informative rather than trivial. - Build realistic multi-agent environments and instrument them so emergent breakdowns are observable — failures that arise from the agents themselves, not ones scripted from the outside. - Run experiments end to end, against external APIs and our own models, orchestrating many agents in parallel. - Catalogue concrete agent failure

Similar remote roles

DevOps Engineer
Jobsforhumanity · Worldwide · mid
AI Engineer
Senovo IT Ltd · Worldwide · mid
AI Engineer
Chabre · Worldwide · mid
Senior Full Stack Engineer - ClickStack
ClickHouse · Worldwide · senior
AI Engineer
Kuro · Worldwide · mid
Backend Engineer - Search (all genders)
ABOUT YOU · Worldwide · mid
Head of Data Labeling
Whitecircle · Worldwide · mid
Data Engineer
Whitecircle · Worldwide · mid