Senior ML Research Engineer, Virtual Cell
SandboxAQ United KingdomGBP 71,400–126,000 / yearSenior
Key requirements
- Python
- Machine Learning
About SandboxAQ
SandboxAQ is a high-growth company delivering AI solutions that address some of the world's greatest challenges. The company’s Large Quantitative Models (LQMs) power advances in life sciences, financial services, navigation, cybersecurity, and other sectors.
We are a global team that is tech-focused and includes experts in AI, chemistry, cybersecurity, physics, mathematics, medicine, engineering, and other specialties. The company emerged from Alphabet Inc. as an independent, growth capital-backed company in 2022, funded by leading investors and supported by a braintrust of industry leaders.
At SandboxAQ, we’ve cultivated an environment that encourages creativity, collaboration, and impact. By investing deeply in our people, we’re building a thriving, global workforce poised to tackle the world's epic challenges. Join us to advance your career in pursuit of an inspiring mission, in a community of like-minded people who value entrepreneurialism, ownership, and transformative impact.
The Opportunity
The AI Sim R&D team builds leading-edge ML and physics-based models ("LQMs") to advance drug discovery. Within this team, AQCell is our virtual cell platform: it takes a cell representation (e.g. basal gene expression) and a perturbation descriptor (e.g. SMILES, dose, and time), predicts the resulting transcriptomic response, and maps that response through pathway activity to functional endpoints such as cell viability, IC50, and toxicity dose-response — helping drug discovery scientists understand the biological repercussions of a compound across the cell, not just whether it binds its target.
As a Machine Learning Engineer on AQCell, you will build and maintain the models and data infrastructure that power this pipeline. You will work across large, heterogeneous transcriptomic and functional-endpoint datasets (e.g. LINCS L1000, GDSC, Tahoe-100M, and DILImap), train and evaluate expression-perturbation and cell-viability prediction models over them, and help harden our evaluation pipeline and baselines so we can trust and improve model performance over time. This role sits at the intersection of machine learning and biology, and offers the opportunity to shape how virtual cell models are trained, validated, and scaled toward real drug discovery decisions.
Key Responsibilities
Model Development: Build, train, and maintain machine learning models for expression-perturbation prediction (e.g. transcriptomic response to a drug perturbation) and downstream functional-endpoint prediction (e.g. cell viability, IC50, toxicity dose-response).
Large-Scale Dataset Management: Acquire, harmonize, and manage large-scale biological datasets (e.g. LINCS L1000, GDSC, Tahoe-100M, DILImap) — including schema harmonization, normalization, and de-duplication across cell, drug, and assay identifiers — and manage model training pipelines over these pooled datasets.
Evaluation & Baselines: Contribute to automating and hardening the end-to-end evaluation pipeline, including implementing robust statistical baselines (e.g. cell- and drug-conditioned mean baselines) to rigorously benchmark model performance.
Research Translation: Translate ideas from the scientific literature (e.g. transformer-based perturbation models, knowledge graph and GNN-based embeddings) into working, well-tested code integrated into our modeling framework.
Cross-Functional Collaboration: Partner with computational biologists, software engineers, and product stakeholders to ensure models are grounded in sound biology and are usable in real drug discovery workflows.
Communication: Clearly document methods, assumptions, and results, and communicate findings to both technical and non-technical stakeholders.
Essential Skills & Experience
Academic Foundation: Bachelor's degree in a scientific or quantitative field (Computer Science, Physics, Mathematics, Biology, Chemistry, or related); an advanced degree (MS or PhD) is preferred.
Applied ML in
See your match score for this role.
Xecodai maps the interview stages and shows what is preventing a 95% match.
