Skip to content

$ cat about.md

I'm a DevOps & Platform Engineer who builds for the worst-case scenario.

I design and automate infrastructure that holds up at 3 AM when a node dies mid-deploy and the dashboard still says green. With 80+ interdependent jobs running at billion-row scale, I focus on making the fast path the safe one, so nobody has to choose. I built an agentic automation platform solo, giving back 1000+ hours a month, and learned that an agent you can't replay is one you can't debug.

Job success

100%

Deploy Failures

↓ 25%

Hours Saved/mo

1000+

Sai Pisey, Senior DevOps Engineer & SRE
$ whoami

$ ./experience.sh

Where I've worked

Senior DevOps Engineer · BrightEdge

Jan 2025 - Present

Own reliability and delivery for production Kubernetes across GCP, AWS and on-prem. Cut deployment failures 25% with blue-green Argo Rollouts and regression gates, migrated 300+ machines from CentOS 8 to Rocky Linux with a staged Ansible workflow, and standardised provisioning on Terraform across 150+ MySQL VMs and 10 GKE clusters. Built the agentic AI platform on LangGraph and Temporal that removed 1000+ manual hours a month. On the AIOps, MLOps and LLMOps side: AIOps self-healing across 175+ hosts that turns correlated Prometheus alerts into automated remediation, MLOps training pipelines on Argo Workflows with GPU node scheduling, and LLMOps observability with LangSmith tracing and per-workflow token and cost attribution.

  • Kubernetes
  • Argo Workflows
  • GitOps
  • Terraform
  • AIOps
  • MLOps
  • LLMOps

DevOps Engineer · Sherlock AI

Jun 2022 - Dec 2024

Hardened Kubernetes workloads with PDBs, HPA, network policies and topology spread for a 40% availability gain. Moved releases onto GitOps CI/CD on ArgoCD for 50% fewer deployment failures and 40% faster deploys, cut container images 80% with multi-stage builds, and put Trivy, OWASP and Sealed Secrets in front of every release.

  • Prometheus
  • Grafana
  • Tempo
  • Thanos
  • ArgoCD
  • AWS

Open source · Upstream / CNCF

2025 - Present

Fixes merged upstream into open-source projects including Argo CD (Flink health assessment), Prometheus node_exporter, prometheus-operator, Flux, MetalLB and Dapr. I contribute back to the tools I run in production, tracing bugs into upstream Go source rather than working around them.

  • Go
  • Argo CD
  • Prometheus
  • Flux
  • MetalLB
  • Dapr

Ready to Build Something Resilient?

Whether it's hardening your platform, untangling a messy CI/CD setup, or just talking infrastructure over a call, let's connect. I'm always up for a good systems conversation.

Book a call

$ ./contact.sh

Get in Touch