Senior Software Developer, ML Platform & Infrastructure
Wealthsimple Toronto HeadquartersEst. Est. CAD 130,000–180,000 / yearSenior
Estimated range based on role, country and industry — not published by the company.
Key requirements
- Python
- Snowflake
- Aws
- Kubernetes
- Terraform
- Machine Learning
- Llm
- Data Science
Build something people love
Wealthsimple is Canada’s leading financial innovator. The company offers a full suite of simple, sophisticated financial products across managed investing, do-it-yourself trading, cryptocurrency, tax filing, spending and saving. Wealthsimple currently serves more than 4 million Canadians and holds over $155 billion in assets under administration. The company was founded in 2014 by a team of financial experts and technology entrepreneurs, and is headquartered in Toronto, Canada.
We're proud of what we've built — and we're just getting started. Read our Culture Manual and learn more about how we work .
ML Platform & Infrastructure Team
The Machine Learning Infrastructure & Platform team builds the foundational architecture powering AI and GenAI initiatives across Wealthsimple. We sit at the intersection of production MLOps and cutting-edge GenAI enablement.
As our AI footprint expands rapidly, our priority is evolving our robust MLOps foundations into a scalable, high-performance LLM serving and routing platform. We build self-serve systems that allow Data Scientists and Engineers to host open-source LLMs reliably, optimize inference latencies, manage GPU infrastructure, and benchmark model performance safely in production.
The Role
We are looking for an experienced MLOps or ML Platform Engineer who is excited to pivot their deep background in model orchestration, serving, and platform tooling toward solving the unique challenges of LLM inference and infrastructure .
In this role, you will bridge the gap between traditional MLOps (model lifecycles, pipeline orchestration, serving infrastructure) and modern GenAI stack requirements (vLLM, GPU cluster management, intelligent model routing, and automated Evals). You will take end-to-end ownership of setting the technical direction for operating enterprise-grade LLM systems company-wide.
In this role, you will have the opportunity to:
Turn MLOps expertise to GenAI: Transition traditional ML lifecycle and serving patterns into state-of-the-art LLM inference engines and GPU orchestration systems.
Build model-routing architecture: Design low-latency routing frameworks (e.g., LiteLLM integration) to dynamically direct requests across managed cloud providers (AWS Bedrock) and self-hosted open-source models.
Provision & scale GPU infrastructure: Architect and manage high-performance GPU serving environments on Kubernetes using engines like vLLM, Ray, and Triton.
Develop evaluation & benchmarking tooling: Build automated Evals and observability frameworks to empower engineers and data scientists to validate model quality, latency, and drift against production requirements.
Empower self-serve ML across Wealthsimple: Partner with product engineering and data science teams to build framework-agnostic platform tooling that abstracts infrastructure complexity.
Drive cost & performance optimization: Improve price-performance across self-hosted and managed inference by optimizing capacity, utilization, batching, routing, and model selection while meeting quality and reliability objectives
We are looking for people who have:
7+ years of software engineering experience in ML Infrastructure, MLOps, ML Tooling, or Data Platform engineering.
Deep experience in MLOps/ML Platform practices: Proven track record building and operating self-serve ML platforms, model registry workflows, experiment tracking, or production serving infrastructure (Kubeflow, MLflow, Ray, Triton, SageMaker).
Strong platform fundamentals: Advanced proficiency in Python, container orchestration via Kubernetes , infrastructure-as-code ( Terraform ), and cloud provider ecosystem (AWS).
Strong appetite to specialize in LLM serving: A genuine desire to leverage your existing MLOps skillset to tackle LLM-specific challenges (vLLM, model routing, prompt engineering tooling, vector databases, GPU memory optimization, or LLM evaluation frameworks).
Backend performan
See your match score for this role.
Xecodai maps the interview stages and shows what is preventing a 95% match.
