Software Engineer - Cyber & Autonomous Systems Team
Aisi LondonGBP 65,000–75,000 / yearMid
Key requirements
- Python
- Llm
- Security
About the AI Security Institute
The AI Security Institute is the world's largest and best-funded team dedicated to understanding advanced AI risks and translating that knowledge into action. We’re in the heart of the UK government with direct lines to No. 10 (the Prime Minister's office), and we work with frontier developers and governments globally.
We’re here because governments are critical for advanced AI going well, and UK AISI is uniquely positioned to mobilise them. With our resources, unique agility and international influence, this is the best place to shape both AI development and government action.
The deadline for applying to this role is 11th October 2026, end of day, anywhere on Earth.
About the team
AI capabilities in cybersecurity and autonomy are advancing faster than at any point in history. Frontier models can now work through multi-step network intrusions, discover and exploit software vulnerabilities, and carry out long-horizon technical tasks with increasing independence. These are extraordinary tools for scientific and economic progress, but also have the potential for serious harm if misused or deployed without adequate oversight, as seen in the recent cybersecurity incidents.
Our Cyber & Autonomous Systems Team evaluates the capability of both frontier and open-weight AI models in cybersecurity, autonomy and AI R&D, ensuring the UK government and its partners have an accurate view of risks and capabilities. We use realistic cyber ranges and a large CTF suite for our evaluations, run pre-deployment testing of frontier models, and collaborate with our partners across UK government, frontier labs, and NCSC.
Role Description
We are looking for exceptional Software Engineers at all experience levels, from junior through to senior or staff, who want to work at the forefront of frontier AI security
In this role, you’ll work with cutting-edge technologies on research problems with real-world impact, and receive mentorship and coaching from your manager and the technical leads on your team.
Your day-to-day work might include building state-of-the-art evaluations (e.g., cyber ranges ), running pre-deployment testing exercises (e.g., Claude Mythos Preview ), validating novel model behaviours (e.g., models cheating in evaluations ), and answering timely research questions (e.g., the capabilities of open-weight models ). You would be working alongside engineers and researchers who care deeply about the impact of their work, take pride in their craft, and have a high level of autonomy.
If that sounds exciting to you, we'd love for you to apply!
Core Responsibilities
Build evaluation infrastructure : develop systems and code that make cyber and autonomy evaluations possible at scale.
Develop analysis tooling: turn raw results into insights, including LLM-as-a-judge and other analysis workflows.
Lead testing exercises: identify key research questions, run evaluations to collect evidence, analyse results, and communicate findings and recommendations.
Deliver engineering work: own technical projects and shared components from design through implementation, iteration, and maintenance.
Write production-quality code: build scalable, robust, maintainable Python software with strong testing, documentation, and observability practices.
Example projects
Onboard a cyber range: deploy a new cyber range on AISI’s evaluation infrastructure (for example, Proxmox) and verify it meets the internal Evaluation Quality Standard.
Build an image pipeline: create a pipeline that converts infrastructure-as-code into AMIs ready for evaluation.
Develop a more secure sandbox service: partner with Core Technology to improve AISI’s sandbox platform, representing CAST to ensure the service meets its evaluation requirements.
Prepare data for publication: collect, validate, and organise evaluation data to support publication of research findings.
Impact
Your
See your match score for this role.
Xecodai maps the interview stages and shows what is preventing a 95% match.
