$ cat jobs/phd-research-scientist-intern-reinforcement-learning-for-dif-canva-eebecf72c405.json
PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling
Join the team redefining how the world experiences design. Servus, hey, g'day, mabuhay, kia ora, 你好, hallo, vítejte! Thanks for stopping by. We know job hunting can be a little time consuming and you're probably keen to find out what's on offer, so we'll get straight to the point. Where and how you can work Our flagship campus is in Sydney, Australia but Austria is home to part of our European operations. And you have choice in where and how you work, we trust our Canvanauts to choose the balance that empowers them and their team to achieve their goals. Fun fact, a big part of our Austrian operations is developing the AI product within Canva to help reimagine how artificial intelligence can be used in design. Pretty cool ha! We’re looking for current PhD students ready to bring their research into the real world and help shape the culture of AI at Canva. Our full-time, 16 week AI Research Internship starts in September. During your internship, you’ll work directly with Canva’s AI team on a live, industry-scale project, turning part of your PhD journey into real world impact. You’ll gain hands on experience with real data, production infrastructure and real deadlines, while learning from and working alongside the researchers and engineerings creating Canva’s next generation of AI-powered experiences. What you'd be doing in this role As Canva scales, change continues to be part of our DNA — but we like to think that's all part of the fun. This gives you a flavour of the work you'd start with, and it will likely evolve over time. At the moment, this role is focused on: Designing and validating a rubric-guided, per-layer VLM judge for RGBA layer decomposition, calibrated against human evaluations. Building VLM-based methods for automatic, human-aligned evaluation of multi-layer designs. Turning VLM-based evaluators into reward functions to train generative models in a reinforcement learning setting. Distilling those judges into lightweight reward models that score layer
Similar remote roles
Sr. Principal Product Manager
Twilio · Worldwide · senior
PhD Research Scientist Intern - Edge AI
Canva · Worldwide · junior
Java and Azure Developer : Contract Role
wpp · UK · mid
Senior Program Manager
general dynamics mission systems · US · senior
Site Reliability Engineer (SRE/ DevOps) - Engineering Productivity
Aristanetworks · Worldwide · mid
Senior Data Scientist
dropbox · US · senior
Machine Learning Engineer
Code Compass 🧭 · Worldwide · mid
Machine Learning Engineer
Fruition Group Ireland · Worldwide · mid