All roles with salary

Principal AI Evaluations Platform Engineer

Microsoftcorporation LondonCHF 183,800–309,700 / yearLead

Key requirements

  • Python
  • Go
  • Rust
  • C++
  • Kubernetes
  • Llm
As Microsoft continues to push the boundaries of AI, we are looking for passionate individuals to work with us on some of the most interesting and challenging AI questions of our time. Our vision is bold and broad: to build systems with true artificial intelligence across agents, applications, services, and infrastructure. It is also inclusive: we aim to make AI accessible to consumers, businesses, and developers so that everyone can realize its benefits. Microsoft AI (MS AI) is seeking an experienced engineer to help build and operate the evaluation platform that supports large-scale model training and development. We’re looking for someone who combines strong systems engineering fundamentals with operational excellence and who can build reliable evaluation infrastructure at scale. This role will work closely with researchers and training teams across our European offices to ensure evaluations are reliable, efficient, and available when teams need them. We seek a versatile engineer who can build solutions that stand the test of time and who brings positive energy, empathy, and kindness to the team while remaining highly effective in a fast-paced environment. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees, we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day, we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. This role is based in London, U.K., or Zurich, Switzerland. Candidates are expected to be local to the applicable office and work in the office four days per week. Responsibilities Serve as a primary engineer for the evaluation platform during European hours, providing dedicated coverage for the evaluation stack supporting large-scale model training. Develop and extend core evaluation platform capabilities, including benchmark configurations, problem sets, graders, and evaluation runners. Improve evaluation throughput and scheduling efficiency. Build and maintain monitoring, alerting, and dashboards to proactively detect evaluation degradation. Partner closely with researchers and training teams across European offices to define evaluation requirements, coordinate launches, and unblock experiments. Contribute to technical design and code reviews across the evaluation platform stack, and maintain high-quality operational documentation and on-call runbooks. Embody our culture and values. Required Qualifications Bachelor's Degree in Computer Science, Engineering, Mathematics, or a related field AND 4+ years of technical engineering experience with coding in languages including, but not limited to, Python, C++, Go, or Rust; OR Master's Degree in Computer Science, Engineering, Mathematics, or a related field AND 2+ years of technical engineering experience with coding in languages including, but not limited to, Python, C++, Go, or Rust; OR equivalent experience. Demonstrated experience operating production distributed systems, including participation in a formal on-call rotation. Strong proficiency in Python, including experience working in a large, shared, typed codebase. Experience debugging distributed system failures across the stack, including scheduling, resource exhaustion, networking, and process-level faults. Willingness and availability to participate in a scheduled on-call rotation covering European business hours. Preferred Qualifications Experience with GPU-based workloads, distributed training, or large-scale inference infrastructure. Experience with distributed compute frameworks and schedulers, such as Ray, Kubernetes, Slurm, or comparable systems. Experience with observability and alerting tooling, such as Datadog, Prometheus, or Grafana. Familiarity with LLM evaluation, benchmarking, or reinforcement learning training infrastruc

See your match score for this role.

Xecodai maps the interview stages and shows what is preventing a 95% match.

Analyse this role