ironquill.tech/board

$ cat jobs/phd-research-scientist-intern-reinforcement-learning-images-canva-5013f22ec709.json

PhD Research Scientist Intern - Reinforcement Learning, Images

Canva·UK·London, GB·junior
Apply on smartrecruiters → Get AI match score →
Our global HQ is in Sydney, Australia, but our London campus sits in Hoxton Square, right in the middle of Shoreditch. It's a bit of a warren of stairs and rooms — you will get lost at first, and someone will happily give you a tour. It's a space where our UK team comes together to connect, create and collaborate. Fun fact: our London team is one of the places where the AI powering Canva gets built. This role is based in London, and we're looking for someone who calls it home. Our hybrid way of working gives you flexibility — you'll have the option to work from home as well as connecting and collaborating with your team in-person, on campus. We trust teams to choose the balance that empowers them to achieve their goals. Join the team redefining how the world experiences design. Hey, g'day, mabuhay, kia ora, 你好, hallo, vítejte! We’re looking for current PhD students ready to bring their research into the real world and help shape the culture of AI at Canva. Our full-time, 16 week AI Research Internship starts in September. During your internship, you’ll work directly with Canva’s AI team on a live, industry-scale project, turning part of your PhD journey into real world impact. You’ll gain hands on experience with real data, production infrastructure and real deadlines, while learning from and working alongside the researchers and engineerings creating Canva’s next generation of AI-powered experiences. What you'd be doing in this role As Canva scales, change continues to be part of our DNA — but we like to think that's all part of the fun. This gives you a flavour of the work you'd start with, and it will likely evolve over time. At the moment, this role is focused on: Designing and validating a rubric-guided, per-layer VLM judge for RGBA layer decomposition, calibrated against human evaluations. Building VLM-based methods for automatic, human-aligned evaluation of multi-layer designs. Turning VLM-based evaluators into reward functions to train generative models in a