Site Reliability Engineer
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
Site Reliability Engineer (Contract)
Location: India, remoteHours: 09:00–18:00 UK time, Mon–Fri (13:30–22:30 IST summer, 14:30–23:30 IST winter)Experience: 5–6 yearsCloud: Azure and AWS
We run production infrastructure for UK clients during UK office hours, plus out-of-hours support and fixes for US clients. The platform team is two engineers. This is the reliability seat: you own how we know the platform is healthy, how we find out when it isn't, and how fast we recover. Your counterpart owns delivery pipelines, and you cover each other.
We mean SRE in the substantive sense, writing code to remove operational work rather than only operating things. If that isn't how you want to spend your week, the DevOps seat is the better fit.
Technical
Cloud. Strong depth in either Azure or AWS, working knowledge of the other. We don't expect equal depth in both.
AWS: EC2, S3, IAM, Lambda, CloudWatch, VPC, RDS, EKS/ECS
Azure: VMs, Blob Storage, Entra ID/RBAC, Functions, Azure Monitor, VNet, Azure SQL, AKS
Observability. Design and run logging, metrics, tracing and alerting. Prometheus, Grafana, ELK/OpenSearch, Azure Monitor, Datadog or similar. Dashboards that work during an incident, not just in a review.
Software engineering. Python or Go at production standard: tested, reviewed, maintained by others. Scripting alone isn't enough for this seat.
Kubernetes. Production experience debugging workloads, not only deploying them.
Reliability practice. Defining SLIs and SLOs, capacity forecasting, performance investigation across the application and database boundary, cloud cost efficiency.
Foundations. Strong Linux administration and deep troubleshooting. Solid Terraform. Networking: DNS, TCP/IP, HTTPS, load balancers, security groups, TLS. Enough CI/CD knowledge to cover the other seat.
Useful to have: OpenTelemetry and distributed tracing; formal SLO or error-budget practice; FinOps; database performance tuning; GitOps with ArgoCD or Flux; AZ-305 or AWS Solutions Architect.
Procedural
You lead incidents during UK hours and on your cover weeks, then run the post-incident review. Blameless, with root cause identified and follow-up actions tracked to completion.
Alerting is tuned continuously. Every page should be actionable, and noisy alerts are treated as defects. A rota this small can't absorb false pages.
Production readiness review before new services or clients go live.
Runbooks are deliverables. Whoever picks up a page overnight should be able to work through it without calling the other engineer.
Recurring manual work gets identified and removed with code rather than absorbed.
Infrastructure changes go through code review and pipelines, alongside the DevOps engineer.
The working day
Fixed hours in UK time, so the window follows UK clock changes rather than drifting.
Full overlap with your counterpart and with the UK team, so most work is collaborative rather than solo.
Out-of-hours cover for US clients on alternating weeks, included in the scope of the engagement. Response expectations are set by severity, and we track page volume with the intent of keeping it low.
A real objective of this role is making that rota quieter: better alerting and automated remediation so overnight breakage is caught and where possible fixed without a human.
Agile/Scrum cadence with a distributed team: standups, planning and retros in UK hours.
Written English at CEFR C1 or equivalent (IELTS 7.0+, or comparable professional experience). A large share of communication is asynchronous, and you'll regularly hand a live incident to someone who was asleep when it started.
BYOD. You supply your machine; all production access runs through a managed virtual desktop, with nothing sensitive stored locally. Disk encryption, supported OS, screen lock and endpoint protection required, plus reliable broadband and a backup connection for on-call weeks.
Location: India, remoteHours: 09:00–18:00 UK time, Mon–Fri (13:30–22:30 IST summer, 14:30–23:30 IST winter)Experience: 5–6 yearsCloud: Azure and AWS
We run production infrastructure for UK clients during UK office hours, plus out-of-hours support and fixes for US clients. The platform team is two engineers. This is the reliability seat: you own how we know the platform is healthy, how we find out when it isn't, and how fast we recover. Your counterpart owns delivery pipelines, and you cover each other.
We mean SRE in the substantive sense, writing code to remove operational work rather than only operating things. If that isn't how you want to spend your week, the DevOps seat is the better fit.
Technical
Cloud. Strong depth in either Azure or AWS, working knowledge of the other. We don't expect equal depth in both.
AWS: EC2, S3, IAM, Lambda, CloudWatch, VPC, RDS, EKS/ECS
Azure: VMs, Blob Storage, Entra ID/RBAC, Functions, Azure Monitor, VNet, Azure SQL, AKS
Observability. Design and run logging, metrics, tracing and alerting. Prometheus, Grafana, ELK/OpenSearch, Azure Monitor, Datadog or similar. Dashboards that work during an incident, not just in a review.
Software engineering. Python or Go at production standard: tested, reviewed, maintained by others. Scripting alone isn't enough for this seat.
Kubernetes. Production experience debugging workloads, not only deploying them.
Reliability practice. Defining SLIs and SLOs, capacity forecasting, performance investigation across the application and database boundary, cloud cost efficiency.
Foundations. Strong Linux administration and deep troubleshooting. Solid Terraform. Networking: DNS, TCP/IP, HTTPS, load balancers, security groups, TLS. Enough CI/CD knowledge to cover the other seat.
Useful to have: OpenTelemetry and distributed tracing; formal SLO or error-budget practice; FinOps; database performance tuning; GitOps with ArgoCD or Flux; AZ-305 or AWS Solutions Architect.
Procedural
You lead incidents during UK hours and on your cover weeks, then run the post-incident review. Blameless, with root cause identified and follow-up actions tracked to completion.
Alerting is tuned continuously. Every page should be actionable, and noisy alerts are treated as defects. A rota this small can't absorb false pages.
Production readiness review before new services or clients go live.
Runbooks are deliverables. Whoever picks up a page overnight should be able to work through it without calling the other engineer.
Recurring manual work gets identified and removed with code rather than absorbed.
Infrastructure changes go through code review and pipelines, alongside the DevOps engineer.
The working day
Fixed hours in UK time, so the window follows UK clock changes rather than drifting.
Full overlap with your counterpart and with the UK team, so most work is collaborative rather than solo.
Out-of-hours cover for US clients on alternating weeks, included in the scope of the engagement. Response expectations are set by severity, and we track page volume with the intent of keeping it low.
A real objective of this role is making that rota quieter: better alerting and automated remediation so overnight breakage is caught and where possible fixed without a human.
Agile/Scrum cadence with a distributed team: standups, planning and retros in UK hours.
Written English at CEFR C1 or equivalent (IELTS 7.0+, or comparable professional experience). A large share of communication is asynchronous, and you'll regularly hand a live incident to someone who was asleep when it started.
BYOD. You supply your machine; all production access runs through a managed virtual desktop, with nothing sensitive stored locally. Disk encryption, supported OS, screen lock and endpoint protection required, plus reliable broadband and a backup connection for on-call weeks.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.