Lead Site Reliability Engineer
Zego LondonEst. Est. GBP 75,000–110,000 / yearSenior
Estimated range based on role, country and industry — not published by the company.
Key requirements
- Python
- Aws
- Kubernetes
About Zego 🚀
At Zego, we're on a mission to do the good thing, not the insurance thing.
Insurance hasn't changed much in over a century. The way we live, work and travel has. We're building the real-time, AI-driven infrastructure that powers innovative insurance, so good drivers get cover that works the way they actually drive.
We're not just updating insurance; We're leading the AI evolution in insurance 🤖
For us, AI isn't a line on a roadmap or a buzzword on a slide. It's our operating reality, and it's how we build products that back drivers instead of the old insurance playbook.
We don't do things slowly, and we don't do bureaucracy. We back high-performance builders who want ownership, early responsibility and the chance to do the most career-defining work of their lives. You'll get the space to try things, the tools to move fast, and the room to see your ideas reach millions of drivers.
Do not take our word for it. Read what Zegons say about us on Glassdoor .
If you're ready to build the future of insurance, we're hiring.
Overview of the Role
You will build the Site Reliability Engineering function at Zego, embedding reliability, observability and operational excellence as core engineering concerns - AI is the primary lever for doing that at scale, not a bolt on.
You will be our dedicated SRE, working alongside Systems Engineering and embedded with Product and Engineering across roughly 120 engineers. Teams own their services and their own on-call. You own the framework, the instrumentation and the AI tooling that makes that ownership work.
You will make teams good at running what they build: alert quality over alert volume, runbooks an agent can execute rather than prose that describes, and diagnostics good enough that an engineer or an agent reaches resolution without an SRE in the room.
You will be the primary advocate for reliability with Product and Engineering, making the case with data rather than assertion, so platform health is prioritised alongside delivery commitments.
Key Responsibilities
Define and land the reliability standard: SLIs, SLOs, error budgets and production readiness criteria that teams apply to their own services, with adoption measured rather than assumed.
Encode standards into automation rather than enforcing them by hand. Guardrails in CI, agents that check reliability and observability posture on pull requests, and safe defaults in shared infrastructure so the reliable path is the easy path.
Build SRE capability on Zego's AI platform, extending our MCP servers and agents, and making the estate legible to them through machine readable runbooks, structured telemetry and diagnostics an agent can act on.
Own observability as a practice, including instrumentation design, signal quality, cardinality and cost.
Raise the standard of incident response end-to-end, from detection and triage through to retrospectives that produce change. Put AI to work where it pays off most, stripping toil out of incidents so responders can focus on judgement
Measure toil, publish it, and remove it through tooling that others can run without you.
What you will need to be successful in the Role
We are looking for an engineer who lives SRE and DevOps culture, and treats AI tooling as a default part of how the work gets done, not a side experiment. You will engage and empower teams through decisions grounded in data, and you will be as comfortable building with AI as consuming it: extending MCP servers, writing agents, and automating the operational work that would otherwise fill your week. We treat reliability at Zego as a platform capability, so we care more about what you can make repeatable for others than what you can fix yourself.
What you’ll bring to the Team
Deep SRE experience: SLI and SLO definition, error budget management, and ownership of incident response through to blameless retrospectives.
Fluency with AI as an engineering tool rather than a chat windo
See your match score for this role.
Xecodai maps the interview stages and shows what is preventing a 95% match.
