Skip to content

ALL SYSTEMS NOMINAL

Deploying infrastructure that doesn't fail at 3 AM.

DevOps & Site Reliability Engineer — AWS · Kubernetes · CI/CD

SCROLL TO EXPAND TELEMETRY

deck://boot

// TELEMETRY.PROFILE

Operator Profile

DevOps and Site Reliability Engineer with one year at Minfy Technologies, operating AWS, Kubernetes, and CI/CD alongside senior engineers. I treat infrastructure like a mission: instrumented, observable, and built to hold under load — not patched together and hoped for.

I drove 25% cloud-cost savings, raised deploy success to 98%, and engineered ZEPLOY, a CLI that cut deploy effort by 60%. I also built a self-healing MLOps system that predicts failures before they cascade. Certified across AWS SAA, Azure AZ-104, and OCI Professional.

0%
AWS spend reduced
0%
Deploy success rate
0.0%
Uptime upheld
0%
Deploy effort saved · ZEPLOY

Education

B.Tech, Computer Science and Engineering

SRM Institute of Science and Technology

Chennai, India · 2021 — 2025

// TELEMETRY.PIPELINE

Work Experience

MONITOR STAGE● ACTIVE

Associate DevOps Engineer / Site Reliability Engineer

Minfy Technologies · Oct 2025 — Present
0.0%
Uptime
-0%
MTTD
-0%
MTTR
0+/wk
Deploys
  • Provisioned multi-tenant AWS environments with senior engineers — Terraform modules for ECS, ECR, S3, and VPC plus ALB/NLB load balancing with TLS termination, hosting 10+ microservices.
  • Operated Amazon EKS clusters with the platform team, tuning HPA and Helm to uphold 99.9% uptime.
  • Deployed Prometheus, Grafana, and ELK Stack dashboards with senior SREs, cutting MTTD 35% and MTTR 25%.
  • Partnered with security teams to embed SonarQube, container scanning, and AWS Secrets Manager into CI/CD, remediating critical vulnerabilities pre-release.
  • Pinpointed EC2 rightsizing and Savings Plans optimizations, driving 25% lower AWS spend.
  • Defined SLOs/SLIs with senior engineers and automated incident runbooks, lifting availability from 99.5% to 99.9%.
  • Standardized GitOps on GitHub Actions with platform leads, enabling 15+ deploys/week at 98% success.
BUILD STAGE

Tech Intern, DevOps Engineer

Minfy Technologies · Apr 2025 — Sep 2025
0%
Deploy success
-0%
Pipeline runtime
-0%
Image size
0%
Effort saved
  • Shipped ZEPLOY, a Go/Python CLI automating ECS Fargate deployments, saving 60% of effort across 5 teams.
  • Revamped Jenkins pipelines under senior review, lifting deploy success from 85% to 98% and trimming runtime 40%.
  • Streamlined Docker images with multi-stage builds and Trivy scanning, shrinking image sizes by 30%.
  • Authored Ansible playbooks for server provisioning, compressing manual setup from 4 hours to 15 minutes.
  • Coached 2 junior interns on Docker and Jenkins fundamentals during onboarding.

// TELEMETRY.PROJECTS

Deployed Systems

SVC-01

ZEPLOY

AWS Deployment CLI

STABLE
GoPythonDockerAmazon ECSAWS FargateGitHub Actions
25→5 min
Deploy time
commits · 52w
  • Unified Fargate deployments into a single-command CLI, reducing deploy time from 25 to 5 minutes.
  • Implemented health-check-validated rollback via GitHub Actions, restoring failed releases automatically.
SVC-02

Vanix

Cattle Management System

STABLE
TerraformHCLIaCCloud Infrastructure
100% IaC
Provisioning
commits · 52w
  • Vanix — a cattle-management system tackling real livestock-tracking and herd-record problems for farm operations.
  • Cloud infrastructure provisioned entirely as code with Terraform (HCL) — reproducible, versioned environments.
SVC-03

Prompt Challenge

PromptWars 2026 · 1st-Place Project

STABLE
TypeScriptGenAIPrompt Engineering
#1
PromptWars 2026
commits · 52w
  • The GenAI prompt-engineering project that took 1st place at PromptWars Hyderabad 2026 (Google for Developers × Hack2skill).
  • Built in TypeScript as an interactive prompt-challenge experience.
SVC-04

Predictive Infra Health

Self-Healing MLOps System

IN PROGRESS
PythonDockerAmazon EKSPrometheusCloudWatchELK
3 sources
Telemetry fused
awaiting deployment
  • Trained an ML model on Prometheus, CloudWatch, and ELK telemetry to predict failures and localize root causes.
  • Built a remediation layer recommending node drains, rolling restarts, and scale-downs before failures cascade.
  • Closed the MLOps loop with containerized inference on EKS, drift monitoring, and continuous retraining.
AWAITING DEPLOYMENT

// TELEMETRY.STACK

Systems & Capabilities

CLOUD5

Cloud & Containers

  • AWS (EC2, ECS, EKS, Fargate, Lambda, S3, IAM)
  • Azure
  • OCI
  • Docker
  • Kubernetes
PIPELINE8

IaC & CI/CD

  • Terraform
  • Ansible
  • CloudFormation
  • Jenkins
  • GitHub Actions
  • GitLab CI/CD
  • ArgoCD
  • GitOps
MLOPS5

MLOps

  • Model Training
  • Model Serving
  • Anomaly Detection
  • Drift Monitoring
  • Continuous Retraining
TELEMETRY6

Monitoring & Observability

  • Prometheus
  • Grafana
  • CloudWatch
  • ELK Stack
  • SLIs/SLOs
  • MTTD/MTTR
RELIABILITY4

Reliability

  • High Availability
  • Incident Management
  • Runbook Documentation
  • Failure Prediction
SECURITY8

Security & Networking

  • SonarQube
  • Trivy
  • RBAC
  • SAST
  • Secrets Management
  • TLS
  • L4/L7 Load Balancing
  • DNS
RUNTIME6

Scripting & OS

  • Python
  • Bash
  • TypeScript
  • SQL
  • Linux administration
  • Windows Server

// TELEMETRY.CREDENTIALS

Certifications & Achievements

01

1st Place — PromptWars Hyderabad 2026

Google for Developers × Hack2skill

GenAI prompt-engineering competition

VERIFY ↗
AWS Certified Solutions Architect – Associate badge

AWS Solutions Architect — Associate

Amazon Web Services

SAA-C03

Oracle Cloud Infrastructure 2025 Certified DevOps Professional badge

OCI 2025 Certified DevOps Professional

Oracle Cloud Infrastructure

Professional · 2025

Microsoft Certified: Azure Administrator Associate certificate

Microsoft Certified: Azure Administrator Associate

Microsoft

AZ-104

ID · C8C967592FA523CC

OCI 2025 Multicloud Architect Professional

Oracle Cloud Infrastructure

Professional · 2025

// TELEMETRY.CHANNEL

Open a Channel

channel://open

Available for DevOps / SRE opportunities. Open a channel and I will get back to you.