$ whoami

Simon Witheridge

Lead Site Reliability Engineer

$ skills

LeadershipPythonGoAWSDockerKubernetes

Projects

Global Scale Kubernetes Platform

Designed and implemented a Kubernetes platform supporting 100+ microservices with high availability and auto-scaling

Key Achievements:

  • Built around a shift-left mentality
  • Implemented Istio in Ambient mode for service mesh capabilities and Gateway API adoption
  • Designed and built cluster templating engine to enable rapid deployment to new regions
  • Built developer self-service platform for deploying and managing microservices
KubernetesDockerIstioEKSAWSTerraformHelm

SLAs, SLOs, SLIs and Error Budget Implementation

Implemented a comprehensive SRE framework for monitoring and maintaining service reliability across multiple microservices

Key Achievements:

  • Designed and implemented a metadata-driven SLO framework for easy configuration and management of SLOs across services
  • Led the education and adoption of SLOs and error budgets across the organization, resulting in improved service reliability and customer satisfaction
  • Utilised a shift-left approach to SLOs, integrating them into the development process and ensuring that reliability is considered from the start of the software lifecycle
DatadogPythonGithubGrafanaPrometheusTerraform

Impact & Achievements

Platform Engineering and Reliability

Global Platform Scale

  • • Designed and managed global Kubernetes (EKS) infrastructure spanning 12+ AWS regions
  • • Engineered self-service Developer Platforms enabling rapid prototyping while eliminating platform team bottlenecks
  • • Architected high-availability distributed systems processing over 200 million requests daily

Observability & Operational Resilience

  • • Designed enterprise-wide Observability frameworks (SLIs/SLOs, monitoring, tracing) across AWS, Azure, and on-premise
  • • Served as Incident Commander & first responder across 24x7 and follow-the-sun operational models
  • • Established blameless post-mortem cultures to systematically lower MTTR and maintain high platform uptime

Infrastructure & DevOps Architecture

Cloud Infrastructure & IaC

  • • Architected AWS Landing Zones and Transit Gateway backbones, reducing cloud account provisioning time from weeks to minutes
  • • Standardized multi-cloud infrastructure delivery using Infrastructure as Code (Terraform) and GitOps
  • • Scaled enterprise Kubernetes (EKS) PaaS platforms for high-profile clients (including 3 Fortune 100 companies and major UK retailers)

Continuous Delivery & DevEx

  • • Streamlined deployment pipelines to cut prototype-to-production lead time from 2 weeks to 1 day
  • • Built self-service tooling and standardized deployment workflows to enhance developer productivity
  • • Implemented automated end-to-end test suites and compliance verification for platform infrastructure

Development & Strategic Leadership

Security & Governance

  • • Engineered Secure-by-Design platform architectures maintaining ISO 27001 and SOC 2 compliance.
  • • Embedded Shift-Left security practices (container vulnerability scanning, IaC linting) into CI/CD pipelines
  • • Spearheaded DevOps and SRE transformations to modernize traditional engineering organizations

Strategic Leadership & Delivery

  • • Led, mentored, and scaled distributed global engineering teams across multiple time zones, fostering technical growth and operational excellence
  • • Successfully directed 4 major enterprise platform build-outs and complex multi-cloud migrations
  • • Championed engineering standards including Conventional Commits, GitOps workflows, and IaC practices

$ contact --info

Let's Connect

$ location --current

Sheffield/Manchester/Remote, UK

$ contact --email

simon@witheridge.dev

$ cat resume.pdf

Download Resume

$ ls ./social-links