All roles with salary

Member of Technical Staff, Intelligence

Substrate Bio LondonEst. Est. GBP 85,000–120,000 / yearLead

Estimated range based on role, country and industry — not published by the company.

Key requirements

  • Python
  • Sql
  • Aws
  • Kubernetes
  • Docker
LOCATION: King’s Cross, London · PATTERN: Hybrid, with regular time in the lab The opportunity We’re a stealth startup in AI and bio, building infrastructure for data generation. The team is small and elite, the problems are hard, and the foundations are being laid right now. You’d help write them from the first line, with real ownership and a direct line to the founders. This role builds our analysis from the start. Every result leaves an instrument as a trace or a vendor file, and it has to become a value someone can rely on, linked back to the sequence and sample it came from. That means fitting it properly, checking it against controls, catching the plate or reagent lot that drifted, and showing the evidence. Someone has to own that work. That is this role. About us AI for biology has a data problem. The data that matters most doesn’t exist yet, so we’re building the infrastructure to produce it, with quality and traceability built in from day one. We’re in stealth and heads-down on execution. We’ll share more about what we’re building once you’ve spoken with the team. The role You will join the intelligence team and own the analysis of our assay data, and the data and ML infrastructure it runs on, from the point the data has been captured to the result that gets delivered. You will work most closely with the protein characterisation scientists who run the assays, and with the software team, who get data out of the instruments and own the platform it lives on. You build on top of that platform, and you own the quality of the analysis that runs on it. The role is part data engineer and part biostatistician, and all of the data is biological. Much of it is building the analysis and ML pipelines that run on our captured data, and the statistics inside them. The rest is presenting those results clearly, with the statistics that matter, such as uncertainty and QC, shown in a form they can understand. The value is in doing that carefully, in code that behaves the same way on every run. What you will own ● Analysis and ML pipelines that turn captured assay and sequence data into consistent, versioned datasets, linked to the sequence, sample, plate and well each result came from, built with the software team who own data capture. ● Data validation before any analysis runs, catching a mislabelled plate, a swapped sample or a missing well. ● Sequence-level bioinformatics, such as computing the properties of each protein sequence and checking constructs for problems before the DNA is ordered. ● Biostatistics for experimental data: experimental design, replicates and controls, error propagation, outlier handling, and assay acceptance criteria, such as CV across replicates, Z′ for plate assays, and reference standards staying within range. ● Curve fitting, including dose-response (EC50 and IC50) with sound normalisation and fits to SPR, BLI, HPLC and nanoDSF traces, with a measure of confidence on each fitted parameter and automated QC that sends doubtful fits to a person. ● Detection of plate, lot and batch effects across a dataset before it is delivered, and control charting of reference standards so drift shows up over weeks as well as within a run. ● The data visualisation that accompanies each delivered result, showing it with its uncertainty, QC and batch structure, so anyone can understand and check it without asking us. Who you are You have worked with messy experimental data and built things other people relied on. That might have been in a biotech or pharma team, an academic lab, a core facility, or a data-heavy field outside biology where you have since picked up the biology. Nobody arrives with every part of this role. If you are strong in most of it and quick to learn the rest, we want to hear from you. You are happy to get your hands dirty on the data. Much of the job is careful, repeated work on real instrument output, and we want someone who takes pride in doing that well. You write c

See your match score for this role.

Xecodai maps the interview stages and shows what is preventing a 95% match.

Analyse this role