Skip to content
$ saipisey ~ init --portfolio

I build systems that don't fall over - especially when they should.

$ ./stats.sh

Results that speak for themselves

↓25%

Deployment failures

1000+

Manual hours removed / month

100%

Service observability coverage

15+

Live failovers, zero data loss

$ ls ./projects

Things I've built

us-east-1aus-east-1bus-east-1cZONE DOWNcheckout-api LOST 3/3fix: topology spread ✓

Most multi-AZ clusters are multi-AZ only on paper. survive-zone reads where pods actually run, takes each zone away in turn and reports which workloads go dark. Every fix it prints is checked against the real Kubernetes scheduler first. Accepted into the Kubernetes krew plugin index.

  • Go
  • kube-scheduler framework
  • kubectl plugin
  • Prometheus exporter
$ kubectl krew install survive-zone
agent session → assai@get podsREADscale web 3SCALEdelete pvc/dbDELETEexec psql DROPEXEC·SQLimpersonatemeasureCELallowkube-apiserverdataservicepodsmeasured on live stateimpactdata 1 PVCsvc db → 0env=prodHELD · data-destructionby a humanapprovedenypolicy.yaml- when: impact.dataDestroyed > 0then: hold- when: size(pdbViolations) > 0auditget podsscale webdelete pvcexec psql

AI agents with kubectl access usually share one admin credential. blastgate runs every agent request as the human behind the session, measures what each write would destroy, and holds the dangerous ones until a person approves them.

  • Go
  • Kubernetes impersonation
  • CEL policy
  • React
alerts / minCPU5xxOOMlatency5xxCPUAIcorrelateINCIDENT #1root causedb poolexhausted41 alerts → 1heal: scale pool ✓one page, not forty

AI Alert Analyzer

Turns an Alertmanager flood into one incident per root cause, then runs policy-guarded fixes (restart, scale, clear crash-looping pods) with a dry-run switch, so the on-call engineer sees the cause, not fifty symptoms.

  • Python
  • Alertmanager
  • Kubernetes
  • AWS EKS
temporal · rollup-mini · run 7f3apreflight12s ✓signed · saibudget$31k > $15kus_en✓uk_en✓de_de✗fr_fr…tier 1tier 2drainre-run ×2publishmanualquota vendor-budget100%LLM advisor · proposescause: de_de quota exhaustedevidence: 14× HTTP 429conf 0.82playbookwait-for-quotapre-approved ✓sandboxed k8s Jobretry tier 2history: 23 events · replayable

Runbooks as YAML, run on Temporal. Steps stop at real sign-off gates instead of a Slack message nobody enforces. When a step fails, an LLM advisor proposes the root cause, but only a pre-approved playbook can ever run.

  • Python
  • Temporal
  • Argo Workflows
  • Claude
scenario home-address · subject syntheticvector storepurged ✓chat historypurged ✓summaryholds ✗"…lives near Lakeview"probesrecall0 hits ✓Where do I live?I don't know. ✓Gyms near me?Try near Lakeview… ✗leak sourcejudge panel✓✗✗2 / 3 reveal the addressreport.htmlLEAKEDjunit: 1 failuresig ed25519 3f9a…c11 plant2 exercise3 erase4 probe5 judge6 signexit 1 · run signed, verifiable offline

Tests whether an AI agent's memory really forgot. It plants a synthetic fact, asks the agent to erase it, probes with indirect questions, and signs the evidence into a report anyone can verify.

  • Python
  • LLM evaluation
  • Ed25519 signing

$ ./skills.sh

My tech stack

Infrastructure

Kubernetes • Docker • Terraform • Helm

CI/CD & GitOps

ArgoCD • Jenkins • Argo Workflows • GitOps

Observability

Prometheus • Grafana • Tempo • Thanos

Cloud

AWS • GCP • Azure

Automation & Scripting

Python • Bash • Ansible

Security

Trivy • OWASP • Sealed Secrets • Kubernetes Security

$ ./services.sh

Services & Offerings

01

Continuous delivery

I build rock-solid pipelines that turn deployments from high-stress events into non-events. Your code moves to production automatically, securely, and with instant zero-downtime rollback capabilities.

02

Site Reliability

I design observability systems that tell you exactly what's failing before your users notice. By setting up clean alerting and auto-healing infrastructure, I make sure you get a full night's sleep.

03

SecOps & Hardening

I integrate security directly into your delivery flow. From least-privilege cloud IAM to automated container vulnerability scanning and secrets management, your defense is built in, not bolted on.

04

Compliance as Code

I codify security and compliance policies into automated checks in the pipeline, so evidence is collected continuously instead of assembled before an audit.

05

FinOps Cost Control

I audit and right-size your cloud footprint to eliminate waste: budget boundaries, autoscaling that matches demand, and spot capacity where the workload tolerates it.

06

AIOps, MLOps & LLMOps

I run AI and ML workloads with the same discipline as production services: self-healing that turns correlated alerts into automated remediation, training pipelines on Argo Workflows with GPU scheduling, and LLM observability with tracing and token-cost attribution.

$ ./feedback.sh

Testimonials

One thing I noticed about Sai is that he doesn't need everything to be clearly defined before getting started. Give him a production problem and some context, and he'll usually work his way through the infrastructure, identify where the actual problem is, and come back with a practical solution. He has a very strong ownership mindset, which is something I value a lot in infrastructure engineers.

Sanket Nighot

CTO, InfraThrone

Sai is the engineer I would bring into a problem when the first few things we've tried haven't worked. He is comfortable going deep into Kubernetes, networking, CI/CD, cloud infrastructure, and monitoring rather than treating each as a separate area. I particularly liked that he would automate the fix afterwards instead of accepting the same operational issue as something we'd have to deal with again.

Ritik Shukla

Lead DevOps Engineer

What stood out to me about Sai was how seriously he took production issues. He didn't rush into making changes just to get something working again. He would look at the logs, metrics, infrastructure and recent changes, piece everything together, and then explain what he had found. That level of ownership gave me confidence when dealing with difficult technical issues.

Sumedh Gadkari

Assistant Vice President, Barclays

Sai was always someone I could rely on when something technical needed to get done properly. What I appreciated most was that he didn't treat DevOps as just servers, pipelines, and deployments. He understood the operational impact on the business and was good at finding ways to remove manual work and make things more dependable.

Saurav Chaudhary

CEO, InfraThrone

I've always found distributed systems problems interesting, and Sai was one of the few DevOps engineers I worked with who was equally comfortable discussing the infrastructure underneath them. We could talk about databases, Kubernetes, scaling, networking, or failure scenarios without having to simplify the conversation. He approaches infrastructure with an engineer's curiosity rather than just following runbooks.

Murtaza Shajapurwala

Co-founder, KiviDB

Sai doesn't treat security as something that gets added after the infrastructure is built. In conversations around deployments and production systems, security was usually part of the design discussion itself. I appreciated that he was willing to think about the operational side of security - secrets, access, image security, and deployment controls - rather than looking at security as a separate checkbox.

Edafe Ukoh

Product Security Engineer

Sai was very good at taking infrastructure problems that could easily become a blocker for the product team and figuring out a practical way forward. He was responsive when things went wrong, but more importantly, he looked for ways to prevent the same issues from becoming recurring problems. That reliability made it easier for the product team to plan releases and work with engineering.

Rahul Wandile

Senior Product Manager

From a developer's perspective, good DevOps is something you notice when you don't have to think about it. Sai made deployments and infrastructure issues much less painful for the development team. When something broke, he was also willing to look at the application side instead of immediately saying it was an infrastructure problem. That made debugging with him much easier.

Kamlesh Chhipa

Software Development Engineer

Ready to Build Something Resilient?

Whether it's hardening your platform, untangling a messy CI/CD setup, or just talking infrastructure over a call, let's connect. I'm always up for a good systems conversation.

Book a call

$ ./contact.sh

Get in Touch

I only use these details to reply to you. Ask and I'll delete them.