Job Description
Role Title: Senior Site Reliability Engineer Location: Knutsford or Glasgow - Hybrid (2 days per week onsite) Role Category: Permanent Overview: We're recruiting for an experienced Senior Site Reliability Engineer to drive reliability, scalability and performance across critical banking systems. This role combines hands-on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities: * Build and maintain reliable, scalable and secure infrastructure platforms and solutions. * Apply SRE and software engineering practices to improve reliability, availability and performance. * Monitor systems, manage incidents and lead complex troubleshooting and root cause analysis. * Develop automation using programming and scripting to reduce manual intervention and improve efficiency. * Develop and improve observability, monitoring, instrumentation and performance capabilities. * Use data and reliability metrics to drive continuous improvement and optimisation. * Lead technical discussions, blameless retrospectives and problem-solving activities. * Work with architects, engineers and stakeholders to define requirements and deliver effective solutions. * Provide technical leadership, mentoring and guidance while helping to drive SRE maturity across teams. Required Skills: * 5+ years' experience in SRE, Production Engineering, Platform Engineering or a related discipline. * Strong practical experience with SRE principles and production environments. * Strong programming/scripting skills, such as Python, Go, Java, C# or Bash. * Proven experience in incident management, troubleshooting and root cause analysis. * Strong observability and monitoring experience. * Experience with AWS, Azure or GCP. * Good understanding of operating systems, networking, cloud infrastructure and automation. * Experience with Infrastructure-as-Code. * Strong communication and technical leadership skills. Desirable Skills: * SLOs, SLIs, SLAs and error budgets. * Prometheus, Grafana, Elastic/ELK or OpenTelemetry. * Kubernetes, containers and distributed systems. * Performance and resilience engineering. * Experience driving SRE maturity across engineering teams. * Financial services or banking experience. This is an excellent opportunity for a Senior SRE to combine hands-on engineering with technical leadership and help improve reliability, observability and performance across critical banking systems. GCS is acting as an Employment Agency in relation to this vacancy