Senior site reliability and platform engineer specializing in production reliability, infrastructure automation, observability, and CI/CD across Kubernetes and legacy Linux environments. Builds Python tooling and Infrastructure as Code to modernize operational systems at scale, with recent work spanning centralized logging, configuration-management resilience, ephemeral integration environments, and Prometheus/Grafana observability.