All roles with salary

ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs)

Cosine London OfficeGBP 80,000–110,000 / yearMid

Key requirements

  • Python
  • Go
  • Sql
  • Aws
  • Azure
  • Gcp
  • Kubernetes
  • Docker
  • Machine Learning
  • Llm
Job title: ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs) Location: London; full in-office working as default Start date: ASAP Compensation: £80,000 - £110,000 Base Salary & £80,000 - £110,000 Share options. ___________________________________________________________________________ Help build the software engineers of the future Cosine is building autonomous AI engineers that plan, write and ship code inside real development workflows. Our agents work across complex software systems, and our Lumen models are trained to do more than produce code that looks correct. They are built to understand existing architectures, follow established patterns and produce software that engineers can actually maintain. We develop our agent tooling entirely in-house and post-train open-source models for reliable, enterprise-grade coding performance. Our products are designed for on-premise, VPC and fully air-gapped environments, including security-critical settings where control, privacy and robustness are non-negotiable. In 2024, Cosine achieved a 72% score on OpenAI’s SWE-Lancer benchmark, placing us among the strongest real-world software-engineering AI systems evaluated. We’re now looking for an ML Systems Engineer to help train the next generation of Lumen models. This is a highly hands-on role at the intersection of machine learning, software engineering, data and infrastructure. You’ll build the environments in which models learn to write software, develop the pipelines that generate and curate training data, and run the fine-tuning and reinforcement-learning workloads that shape model behaviour. If you’re excited by the idea that the future quality of coding agents will be determined not just by model architecture, but by the quality of their data, environments and reward functions, this is an opportunity to work directly on that problem. ___________________________________________________________________________ The role You’ll work closely with ML researchers, infrastructure engineers and product teams to decide: What Lumen should learn next. How to generate the right training data. How to design environments that reflect real software-engineering work. How to reward models for producing useful, maintainable code. How to measure whether a new training run genuinely improves the experience for engineers. The systems you build will sit directly inside our model-training loop. Models will write code, use tools, run tests and interact with real repositories. Your work will determine how those interactions are generated, evaluated and fed back into future training. This is not a narrow research role and it is not traditional MLOps. You’ll move between custom PyTorch code, distributed data pipelines, Dockerised services, RL environments, evaluation infrastructure and production-quality software. ___________________________________________________________________________ What you’ll do Build the training systems behind Lumen Contribute to the end-to-end training of software-engineering models. Implement supervised fine-tuning pipelines using curated code and conversation datasets. Build reinforcement-learning loops in which models write code, run tests and use development tools. Develop custom PyTorch dataloaders, training objectives and evaluation workflows. Run and analyse fine-tuning and RL experiments across large, modern open-source models. Create the data that teaches models to engineer Develop synthetic data-generation pipelines for future RL and fine-tuning runs. Design systems for producing, filtering, transforming and sampling large-scale datasets. Work with object storage, dataset sharding and data-quality checks. Investigate which examples, tasks and sampling strategies lead to better model behaviour. Turn model failures into concrete improvements to training data and future experiments. Build reliable RL infrastructure Design, build

See your match score for this role.

Xecodai maps the interview stages and shows what is preventing a 95% match.

Analyse this role