All roles with salary

Staff Research Scientist, STEM

Turing Palo Alto, California, United States; San Francisco, California, United States; Seattle, Washington, United StatesUSD 250,000–400,000 / yearLead

Key requirements

  • Python
  • Machine Learning
  • Agile
About Turing Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com . The Role Turing is seeking exceptional Staff Research Scientists to join our STEM research organization and develop new ways to evaluate, train, and improve frontier AI systems. This is a research-first role focused on problems where the right benchmark, dataset, or methodology often does not yet exist. You will identify important gaps in the literature, propose ambitious new research directions, and take projects from initial hypothesis through experimentation, benchmark construction, and publication. Our research is deliberately focused on frontier STEM evaluation, synthetic data, hallucination and reliability, and agentic science . We are looking for scientists who can recognize important problems early, formulate them precisely, and design rigorous research programs to answer them. What You'll Do Frontier benchmarks and evaluation Identify high-impact gaps in existing benchmark and evaluation literature. Design novel benchmarks in and across STEM fields and on general model functionality. Develop evaluations for emerging model capabilities that are poorly captured by traditional static benchmarks. Design rigorous task-generation, grading, contamination-control, difficulty-calibration, and validation methodologies. Build benchmarks that can become both valuable research contributions and meaningful standards for evaluating frontier models. Synthetic data and post-training Develop methods for generating high-quality synthetic STEM training data. Study how task selection, difficulty, diversity, verification, filtering, and data quality affect downstream performance. Explore methods for generating useful training signal in domains where expert human data is scarce or expensive. Design experiments that determine when synthetic data genuinely improves capabilities rather than simply increasing training volume. Hallucination, reliability, and verification Study hallucination, uncertainty, calibration, and epistemic failure in technical domains. Develop evaluations and methods for improving factual reliability, self-correction, verification, citation, and appropriate abstention. Investigate when models should reason internally, invoke tools, seek external evidence, or recognize that they do not know. Agentic science Research AI systems capable of performing extended scientific and technical work. Develop workflows involving literature search, coding, simulation, tool use, experimentation, verification, and iterative reasoning. Evaluate long-horizon scientific agents and identify the bottlenecks preventing them from reliably performing real research. Explore new approaches to human-AI and multi-agent scientific collaboration. New research directions The areas above are our core focus, not an exhaustive list. Researchers will also have significant latitude to propose new programs in areas such as reasoning, model evaluation, AI-for-science, data generation, and emerging capabilities. What We’re Looking For PhD or equivalent research experi

See your match score for this role.

Xecodai maps the interview stages and shows what is preventing a 95% match.

Analyse this role