Site Reliability Engineer
Keep large-scale systems running fast and reliably.
What does a Site Reliability Engineer do?
Site reliability engineers are the people who keep production systems alive — building the automation, monitoring, and runbooks that mean things stay up even when they want to fall over. You write code, respond to incidents, and constantly improve how reliable a system is. This path takes time to get into but pays off significantly. Good fit if you enjoy deep technical problems and want to specialise in something most people never attempt.
Who it fits
People who enjoy systems thinking, stay calm under pressure, and get satisfaction from preventing problems before they happen.
Suits hands-on learners who enjoy reading engineering postmortems and learning from real-world system failures.
Skills you need
Tools you will use
Starter projects
Good projects to practise and add to a portfolio.
- Set up Prometheus and Grafana to monitor a local service
- Write an Ansible playbook to configure a server
- Simulate an incident and write a postmortem
- Build a basic health-check and alerting system
// ROADMAP
Site Reliability Engineer career roadmap
Estimated time: 12–18 months
Badges show what each resource is. Only items marked certification lead to a professional credential — everything else is learning material.
Linux Systems Administration
8–10 weeksSREs live on the command line. Master Linux deeply — systemd, process management, networking tools (ss, ip, tcpdump), file systems, and performance tools (htop, strace, perf).
Programming (Python or Go)
6–8 weeksSREs automate everything. Python for scripting and quick automation; Go for building reliable tools and services. Learn enough to write production-quality automation scripts.
Kubernetes & Container Orchestration
6–8 weeksMost modern SRE work involves Kubernetes. Learn how to debug running pods, understand resource limits, read logs, and interpret events. KodeKloud has excellent hands-on labs.
Observability: Metrics, Logs & Traces
4–6 weeksYou cannot fix what you cannot see. Learn Prometheus for metrics, Grafana for dashboards, and structured logging. Understand what SLOs, SLIs, and error budgets actually mean.
Incident Management & SRE Book
OngoingRead the Google SRE book — it is available free online and is the definitive guide to how reliable systems are operated. Learn incident management, postmortems, and on-call practices.
// START LEARNING
Site Reliability Engineer learning resources
Videos first — they're the fastest way to get moving — then reading, hands-on practice and courses. Only resources badged professional certification award a formal credential.
Git & GitHub Crash Course 2025
Traversy Media · 1 hour
A current, practical introduction to version control and collaborative GitHub workflows.
Learning resource. This does not award a professional certification.
Docker Tutorial for Beginners
TechWorld with Nana · 3 hours
Containers, images, Docker Compose and the workflow used by cloud and platform teams.
Learning resource. This does not award a professional certification.
Linux Journey
LabEx
Short progressive lessons on the command line, permissions, processes, networking and system administration.
Certificate availability depends on the course provider.
Site Reliability Engineering Books
Google’s free canonical books on SRE principles, production systems, monitoring and incident response.
Learning resource. This does not award a professional certification.
Learn Kubernetes Basics
Kubernetes
Official tutorials for deploying, scaling, updating and debugging a containerised application.
Learning resource. This does not award a professional certification.
// PRACTISE
Gain Site Reliability Engineer work experience
Build practical evidence before your first role. Provider terms and eligibility can change.
GitHub · Good first issues
open-source · REMOTE
Find newcomer-labelled issues and build evidence of collaboration in public repositories.
KodeKloud Engineer · DevOps project tasks on practice systems
hands-on-lab · REMOTE
Complete fictional-company tasks on live practice systems across Linux, containers, automation and infrastructure.
AWS Builder Center · Cloud architecture workshops
hands-on-lab · REMOTE
Follow expert-authored workshops to deploy and connect cloud services in guided practical scenarios.
// FIND WORK
Find Site Reliability Engineer jobs
Start with “Site Reliability Engineer”. Platforms without stable public search URLs open with this suggested phrase.
LinkedIn Jobs
Broad professional job search with keyword, location, experience-level and remote filters.
Wellfound
Startup and technology roles with company and compensation context.
Dice
Specialist technology roles across engineering, data, security and infrastructure.
Built In
Technology and startup job discovery, including remote and city-focused listings.
UK visa and international opportunities
Sponsorship depends on the employer, vacancy and current immigration rules. Use the linked official guidance and verify every role before applying.
Going independent
SRE consulting is growing as companies mature their reliability engineering practices. Fractional SRE roles are emerging as a new service model.
SRE and Reliability Consulting
Medium effortHelp companies improve system reliability, implement observability, and build incident management practices.
Examples
- SRE practice setup for growing engineering teams
- Incident management and postmortem process design
- Observability stack implementation (Prometheus, Grafana)
- SLO and error budget design workshops
Getting started
- 1. Get CKA and cloud certifications as baseline credibility
- 2. Write case studies framed as "from X uptime to Y uptime in Z months"
- 3. Target scale-ups that are starting to feel reliability pain
- 4. Position as "fractional SRE" for companies too small to hire full-time
DevOps Tool Development
High effortBuild tooling that solves reliability or observability problems.
Examples
- Incident management SaaS tools
- SLO tracking and alerting products
- Infrastructure cost optimisation tools
Getting started
- 1. Find friction you personally experience in SRE tooling
- 2. Build open-source first, then offer cloud-hosted commercial tier
- 3. Partner with Grafana or Prometheus ecosystem for distribution
Communities & tools
// EARNING POTENTIAL
Earning potential
2025On-call responsibilities are compensated with additional pay or time off. SREs at large-scale companies earn among the highest salaries in infrastructure.
Freelance rates
Day rate
£400–£900
Hourly rate
£50–£115
What affects salary
- On-call rotas are typically compensated — factor into total compensation calculation
- Go programming experience is preferred at many SRE teams and commands a premium
- Observability tool ownership (Prometheus, Grafana, Datadog) adds leverage
- Incident management and postmortem experience valued at senior levels
- Financial services and e-commerce SRE roles pay at the top of these ranges
🌍 US SREs at Google and major tech companies earn $150k–$250k. Google coined the SRE role — their internal compensation sets industry expectations.
Based on: LinkedIn Salary · Glassdoor UK · ITJobsWatch · Google SRE Hiring Data. All figures approximate.
// RESOURCES
Free resources to get started
Recommended starting points — no payment required.
- Book
- Course
- Lab
- Lab
- Reading
Paid courses worth considering
These are not required — free resources above can get you far.
- Professional certification
- Course
Browse all resources
Resource directory.
Free and paid resources across every tech career path — searchable and filterable.
Not sure this fits?
Take the assessment.
Answer a few short questions and get your top matches — with reasons why they fit you.
Wondering if you are ready for this path? Analyse your CV →
Related careers
Compare paths that share skills, tools or ways of working.


