Other · In person

Applied Science Intern - Clipchamp

Microsoft

Location
Sydney, New South Wales, Australia · Melbourne, Victoria, Australia · Brisbane, Queensland, Australia
Opening date
Opened today
Source
Microsoft official careers
Employer posting

Job description

Captured from employer · Aug 24, 2026

Research, prototype and evaluate computer vision and multimodal models for video understanding. Design and build agentic workflows in which language and multimodal models plan, invoke tools, and act over an editing timeline, then measure their reliability, latency and quality against real product scenarios. Translate ambiguous product goals into well-defined machine learning tasks, and design experiments, baselines and metrics that allow rapid iteration and optimization.

Fine-tune, adapt and benchmark state-of-the-art foundation models (for example, vision-language models, diffusion models and LLM-based agents) on domain data, under the guidance of a senior scientist. Prepare and curate datasets for training and evaluation, reviewing data for quality and technical constraints, and documenting the actions taken to address data quality issues. Implement prototypes of scalable AI components and contribute to code reviews, analysis and technical documentation.

Build an understanding of the broader research area and industry trends, and share your findings with the team through demos, write-ups and presentations. Currently pursuing a Doctorate degree in Computer Science, Applied Science, Statistics, or a related field, with a research focus in computer vision, machine learning or multimodal AI. Must have at least 1 semester/term remaining following the completion of the internship.

Hands-on experience building and training deep learning models in Python with frameworks such as PyTorch, Hugging Face Transformers or Diffusers. Demonstrated experience in computer vision or video understanding, evidenced by research publications, open-source contributions or substantial project work. Experience with agentic AI systems, including tool use and function calling, planning and reasoning, multi-step orchestration, or agent evaluation frameworks.

Familiarity with state-of-the-art architectures and techniques, such as transformers, attention mechanisms, vision-language models, diffusion models, transfer learning and parameter-efficient fine-tuning. Publication record at major AI conferences, including NeurIPS, ICLR, CVPR, ICML, ACL, EMNLP, ECCV/ICCV, etc. Experience designing rigorous evaluation methodology for generative or agentic systems, including human evaluation and automated benchmarks.

Experience taking research prototypes towards production, and fluency in one or more of Python, C# or TypeScript.

Private to you

Application tracker

Keep the exact resume, contact, and follow-up notes with this role.View all applications →

Saved to this browser unless you connect an account; connected records follow your verified account across devices.

Timeline

0 entries