$ cat jobs/phd-research-scientist-intern-edge-ai-canva-e281bcde4347.json
PhD Research Scientist Intern - Edge AI
Hey, g'day, mabuhay, kia ora, 你好, hallo, vítejte! Our global HQ is in Sydney, Australia, but our London campus sits in Hoxton Square, right in the middle of Shoreditch. It's a bit of a warren of stairs and rooms — you will get lost at first, and someone will happily give you a tour. It's a space where our UK team comes together to connect, create and collaborate. Fun fact: our London team is one of the places where the AI powering Canva gets built. This role is based in London, and we're looking for someone who calls it home. Our hybrid way of working gives you flexibility — you'll have the option to work from home as well as connecting and collaborating with your team in-person, on campus. We trust teams to choose the balance that empowers them to achieve their goals. At Canva, our mission is to empower the world to design. We're building AI that feels magical and lands real impact for millions of people, helping anyone create with confidence. We're looking for a research intern who is excited by efficient ML and edge deployment to help us bring video-capable vision-language models onto the devices in people's pockets. About the team We're the Video Storytelling team, working on the models and systems behind Canva's video AI experiences. We partner closely with our Edge AI group, who are building Canva's on-device inference capability, to explore what's possible when AI runs directly on users' own hardware. We already own several of the server-side capabilities this work builds on, so you'll be joining a team with a strong command of the data, models, and pipelines behind the problem. About the role This is a 14-week research internship focuses on one clear question: can a video-capable vision-language model be optimised to run efficiently on high-traffic consumer phones, while retaining enough capability to serve a real product use case? The use case is intelligent captioning, where the model's visual understanding of a video drives context-aware, intelligently pl