$ whoami
Simon Witheridge
Lead Site Reliability Engineer
$ skills
LeadershipPythonGoAWSDockerKubernetes
Projects
Global Scale Kubernetes Platform
Designed and implemented a Kubernetes platform supporting 100+ microservices with high availability and auto-scaling
Key Achievements:
- Built around a shift-left mentality
- Implemented Istio in Ambient mode for service mesh capabilities and Gateway API adoption
- Designed and built cluster templating engine to enable rapid deployment to new regions
- Built developer self-service platform for deploying and managing microservices
KubernetesDockerIstioEKSAWSTerraformHelm
SLAs, SLOs, SLIs and Error Budget Implementation
Implemented a comprehensive SRE framework for monitoring and maintaining service reliability across multiple microservices
Key Achievements:
- Designed and implemented a metadata-driven SLO framework for easy configuration and management of SLOs across services
- Led the education and adoption of SLOs and error budgets across the organization, resulting in improved service reliability and customer satisfaction
- Utilised a shift-left approach to SLOs, integrating them into the development process and ensuring that reliability is considered from the start of the software lifecycle
DatadogPythonGithubGrafanaPrometheusTerraform
Impact & Achievements
Platform Engineering and Reliability
Global Platform Scale
- • Designed and managed global Kubernetes (EKS) infrastructure spanning 12+ AWS regions
- • Engineered self-service Developer Platforms enabling rapid prototyping while eliminating platform team bottlenecks
- • Architected high-availability distributed systems processing over 200 million requests daily
Observability & Operational Resilience
- • Designed enterprise-wide Observability frameworks (SLIs/SLOs, monitoring, tracing) across AWS, Azure, and on-premise
- • Served as Incident Commander & first responder across 24x7 and follow-the-sun operational models
- • Established blameless post-mortem cultures to systematically lower MTTR and maintain high platform uptime
Infrastructure & DevOps Architecture
Cloud Infrastructure & IaC
- • Architected AWS Landing Zones and Transit Gateway backbones, reducing cloud account provisioning time from weeks to minutes
- • Standardized multi-cloud infrastructure delivery using Infrastructure as Code (Terraform) and GitOps
- • Scaled enterprise Kubernetes (EKS) PaaS platforms for high-profile clients (including 3 Fortune 100 companies and major UK retailers)
Continuous Delivery & DevEx
- • Streamlined deployment pipelines to cut prototype-to-production lead time from 2 weeks to 1 day
- • Built self-service tooling and standardized deployment workflows to enhance developer productivity
- • Implemented automated end-to-end test suites and compliance verification for platform infrastructure
Development & Strategic Leadership
Security & Governance
- • Engineered Secure-by-Design platform architectures maintaining ISO 27001 and SOC 2 compliance.
- • Embedded Shift-Left security practices (container vulnerability scanning, IaC linting) into CI/CD pipelines
- • Spearheaded DevOps and SRE transformations to modernize traditional engineering organizations
Strategic Leadership & Delivery
- • Led, mentored, and scaled distributed global engineering teams across multiple time zones, fostering technical growth and operational excellence
- • Successfully directed 4 major enterprise platform build-outs and complex multi-cloud migrations
- • Championed engineering standards including Conventional Commits, GitOps workflows, and IaC practices
$ contact --info
Let's Connect
$ location --current
Sheffield/Manchester/Remote, UK
$ contact --email
simon@witheridge.dev$ cat resume.pdf
Download Resume$ ls ./social-links