Research Engineer, Frontier Data
Turing BrazilEst. Est. BRL 70,000–110,000 / monthMid
Estimated range based on role, country and industry — not published by the company.
Key requirements
- Python
- Java
- Go
- Rust
- C++
- Sql
- Deep Learning
- Agile
About Turing
Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com .
*This is a remote role and can be performed anywhere in Brazil/Colombia*
The Role
We are looking for a Research Engineer to help deliver frontier-quality datasets, RL environments, and evaluations that improve state-of-the-art models for a frontier AI lab client. You will work directly with the client's researchers and engineers, turning inbound requests and post-training goals into concrete technical proposals and data/environment specifications, and then owning the technical execution: building the quality and verification systems that ensure what we deliver meets extremely high standards for correctness, realism, diversity, difficulty, and measurable model lift.
Most of this work is bespoke and custom, built for one client's evolving needs.You'll typically stay attached to a single frontier lab client so you can build real context and a working relationship with their team, with flexibility to support other accounts when needed. Because bespoke work is judged on both turnaround time and quality, both matter equally here.
This role is designed for candidates with experience building and improving deep learning systems, especially where strong results depend on data quality, data curation, denoising, synthetic data generation, and rigorous evaluation. You'll operate in one or more of the following capability areas:
Coding and software engineering agents (repositories, unit tests, debugging, tool use, code reviews, long-horizon workflows)
RL environments and verifier-based training (tasks, rewards/verifiers, trajectories, evaluation harnesses)
Multimodal data and reasoning (text + images + documents + tables/charts; optional audio/video)
STEM reasoning (math, physics, chemistry, bio, engineering – solution verification and error analysis)
Modern embodied AI / VLM-driven agents (vision-language(-action) models, embodied task suites, tool/sensor/action abstractions, long-horizon interaction data)
What You'll Do
1) Own data and environment quality from an AI researcher perspective
Respond to inbound requests from the client and translate ambiguous, evolving research goals into a technical proposal and clear data requirements: target skills, failure modes, difficulty calibration, coverage, and success metrics.
Provide the technical feasibility assessment, assumptions, acceptance criteria, and evaluation requirements needed to scope the contract or statement of work.
Define what “good” looks like by creating detailed rubrics, counterexamples, and boundary cases (what to include vs. exclude).
Perform deep, detail-oriented audits of produced data: spot subtle errors, reward hacking opportunities, leakage, ambiguity, inconsistent assumptions, and distribution shifts.
Drive iterative improvements using evidence: group failures (for example, by semantic similarity) to find the underlying gap, then turn that into concrete instructions for the production team — better few-shot examples, explicit counterexamples, clearer guidance on what's causing rejections.
2) Design and build
See your match score for this role.
Xecodai maps the interview stages and shows what is preventing a 95% match.
