ironquill.tech/board

$ cat jobs/senior-ai-engineer-f-m-d-remote-in-germany-synera-3e1865586c85.json

Senior AI Engineer (f/m/d) - Remote in Germany

synera·EU·Germany·senior
restllm
Apply on himalayas → Get AI match score →
Most AI teams talk about evaluation. At Synera , you'd own it — end-to-end, in production, for two real agentic products used by companies like BMW, Airbus, and NASA. WHAT YOU WILL DO You’ll join the Agentic Ants team as our second native AI engineer. You’ll build and extend our agentic systems alongside the rest of the team — new tools, sub-agents, prompt iterations — and co-own the evaluation framework that gates every change: golden datasets, LLM-as-judge pipelines, regression suites, production-trace mining. This isn’t about adding scores to a notebook — it’s about shipping agents the team can trust, and proving it. You’ll work across both Synera MAS and Synera Assistant, picking up reliability concerns (error handling, retries) as they show up in your work. Partnering with Ruben and the wider engineering team, you’ll also help the team go deeper on customer insights data. 👉 What a week at Synera could look like: Monday: Kick off sprint planning with the Agentic Ants team, review Langfuse traces from the weekend, and flag any new failure modes worth triaging. Tuesday: Work on the golden dataset for the supervisor routing surface — curating examples, versioning the set, and writing evaluators with Ahmed. Wednesday: Join a cross-team sync with QA and product to align on new eval coverage for an upcoming agent feature, then push a CI integration so eval regressions block the next PR. Thursday: Pair with the AI and software engineers on extending our agentic system — a new tool, a routing tweak, or a prompt iteration — then write the eval that gates the change before it ships. Friday: Review a calibrated LLM-as-judge output alongside human labels, refine the rubric, and share findings in the eng review. ⚡ In 6 months: The evaluation framework is live with golden datasets across multiple agent surfaces, CI gates are blocking on regressions, and the team actually trusts the results. You’ve shipped multiple meaningful changes to the agent graphs — new tools, sub-agent

Similar remote roles

AI Engineer
Neurons · Worldwide · mid
Full Stack Developer
Matrix Consulting Group Srl · Worldwide · mid
AI Engineer
Neurons · Worldwide · mid
General QA Engineer in Payment System Team
Namecheap · Worldwide · mid
Full Stack Software Engineer (AI-Native)
R8 TECHNOLOGIES · Worldwide · mid
Senior Developer - Backend
deutsche telekom it solutions · Worldwide · senior
AI Expert Senior Software Engineer – Healthcare Analytics - Remote
oracle corporation · US · senior
Senior Full-Stack Engineer (AI/LLM-focused)
curotec · Worldwide · senior