# SANTOSH PARSA
**Staff Platform & Infrastructure Engineer | Kubernetes, Cloud (AWS/GCP), Linux & Databases**  
Hyderabad, India | +91 9160396545 | parsa.santosh@gmail.com | [linkedin.com/in/santosh-parsa-1b615554](https://linkedin.com/in/santosh-parsa-1b615554) | [github.com/santosh91parsa](https://github.com/santosh91parsa) | [credly.com/users/santosh-parsa/badges](https://www.credly.com/users/santosh-parsa/badges)

---

## PROFESSIONAL SUMMARY
Staff Platform and Infrastructure Systems Engineer with **13+ years of experience** architecting, managing, and automating mission-critical multi-service infrastructure across enterprise hybrid cloud (**AWS, GCP**) and datacenters. Proven track record across 7+ years at ServiceNow managing foundational infrastructure tiers powering dozens of interconnected platform microservices and distributed workloads under strict **99.99%+ availability SLAs**. Deep technical expertise across **Kubernetes cluster orchestration (EKS/GKE)**, **Infrastructure as Code with Terraform**, **Python and Bash automation**, deep **Linux OS/kernel internals**, and **enterprise networking/DNS (Route53, BIND, CoreDNS, TCP/IP)**. Extensive hands-on experience supporting relational database tiers (**MySQL, PostgreSQL**), web/application middleware (**Apache HTTP Server, Apache Tomcat, Nginx**), software load balancing (**HAProxy, Nginx, ALB/NLB**), and full-stack observability (**Prometheus, Grafana, Splunk**). Experienced Primary Incident Commander driving cross-service root cause analysis, resolving complex cascading failures, and slashing MTTR by **42%** through automated self-healing daemons.

---

## CORE COMPETENCIES & TECHNICAL SKILLS

- **Containerization & Kubernetes:** Kubernetes (EKS, GKE, Multi-Tenant Cluster Management, Helm, Operators, CoreDNS, Ingress Controllers, Service Mesh, CNI/CSI Plugins, HPA, Pod Disruption Budgets), Docker, Container Security, Internal Developer Platforms (IDP), GitOps workflows.
- **Cloud Infrastructure & IaC:** Terraform (Modular Infrastructure Registries, State Management, Drift Detection), Ansible (Fleet Configuration Automation, Dynamic Inventories), AWS (EC2, EKS, VPC, Route53, ALB/NLB, RDS, S3, IAM, KMS), Google Cloud Platform (GCP - GKE, VPC, Cloud SQL, IAM), Multi-Region Active-Active Architectures.
- **Operating Systems, Networking & DNS:** Deep Enterprise Linux Administration (RHEL, Ubuntu, CentOS), Kernel Tuning, Process & Thread Scheduling, Systemd, Cgroups, eBPF Tracing, Performance Profiling (`perf`, `strace`, `tcpdump`, `ss`, `htop`, `iotop`), Network Protocols (TCP/IP, HTTP/2, gRPC, VPC Routing, Subnetting, Firewalls), DNS Architecture (Route53, BIND DNS, CoreDNS, Anycast).
- **Databases & Relational Data Tiers:** MySQL, PostgreSQL, High-Availability Clustering & Replication (Primary-Replica, Failover), Database Connection Pooling (PgBouncer, ProxySQL), Automated Backups & Point-in-Time Recovery (PITR), Schema & Storage Optimization, Query Performance Profiling.
- **Web, Middleware & Software Load Balancing:** Apache HTTP Server, Nginx, Apache Tomcat, Software Load Balancing (**HAProxy, Nginx Load Balancer**, ALB/NLB), Reverse Proxy Architecture, SSL/TLS Offloading & Certificate Management, Connection Pooling & Keep-Alive Tuning, Upstream Health Probes, Sticky Sessions, High-Concurrency Request Routing.
- **Programming & Automation:** Python (Infrastructure Automation, APIs, Scripting Frameworks, Standard Libraries), Advanced Bash / Shell Scripting, FastMCP, Model Context Protocol (MCP), LLM-Augmented Telemetry Pipelines, Automated Incident Triage Engines.
- **Site Reliability Engineering (SRE) & Incident Command:** SLI/SLO Framework Design, Error Budget Governance, Cross-Service Cascading Failure Triage, MTTD/MTTR Optimization, Blameless Postmortems, Chaos Simulation Game Days, Automated Remediation & Self-Healing Daemons, 24/7 Incident Command.
- **Observability & Telemetry:** Prometheus, Grafana, OpenTelemetry (Traces, Metrics, Logs), Splunk, Alertmanager, Multi-Window Multi-Burn-Rate Alerting, Synthetic Cross-Service Health Probes, Availability Dashboards.
- **Version Control, Git Internals & CI/CD:** Advanced Git Architecture & DAG mechanics, 3-Way Merge conflict resolution (Lowest Common Ancestor / LCA analysis, recursive strategies, rebase vs merge governance), branch protection policies, DevSecOps hooks (pre-commit, pre-receive secret scanning), disaster recovery (`git reflog`), GitHub Actions, Jenkins declarative pipelines, JFrog Artifactory, Container Registries, Automated Canary & Blue/Green Deployments, Multi-Service Release Validation.
- **Email & Messaging Infrastructure:** Enterprise Mail Transfer Agents (Postfix, Sendmail, Zimbra, Qmail), SMTP Protocol & Relays, Email Authentication & Security Standards (SPF, DKIM, DMARC Enforcement), Transport Security (STARTTLS), DNS MX Record Routing.

---

## PROFESSIONAL EXPERIENCE

### ServiceNow | 2018 – Present (7+ Years)
**Sr Staff Linux Systems & Infrastructure Platform Engineer** | Mar 2025 – Present  
**Staff Systems Design & Infrastructure Engineer** | Mar 2021 – Mar 2025  
**Senior Linux Systems Administrator** | May 2018 – Mar 2021  
*Hyderabad, India | Technical leader for enterprise multi-service infrastructure management, platform reliability, and cloud operations powering ServiceNow's Global 2000 multi-tenant cloud (99.99%+ uptime).*

#### Multi-Service Infrastructure, Middleware & Database Operations
- **Multi-Service Fleet Operations:** Managed foundational cloud and Linux infrastructure supporting dozens of interconnected enterprise platform services, API gateways, and distributed data pipelines; maintained strict 99.99%+ uptime across thousands of compute nodes.
- **Relational Database Operations (MySQL & PostgreSQL):** Supported enterprise relational database tiers powered by **MySQL and PostgreSQL**, managing automated backup pipelines, primary-replica replication health, connection pooling (PgBouncer, ProxySQL), and failover validation across distributed platform services.
- **Middleware & Software Load Balancing:** Configured and tuned high-throughput **software load balancers (HAProxy, Nginx)** fronting multi-service **Apache and Tomcat** application tiers; optimized connection pools, keep-alive timeouts, buffer limits, and health checks to ensure seamless request distribution and zero-downtime rolling service deployments.
- **Resource Isolation & Capacity Governance:** Enforced strict compute, memory, and I/O isolation across co-located multi-service workloads utilizing Linux cgroups and systemd resource slices, eliminating noisy-neighbor degradation and optimizing multi-tenant cluster density.
- **Zero-Downtime Fleet Maintenance:** Orchestrated automated rolling OS patching, kernel upgrades, and security remediations across multi-service server fleets without customer traffic interruption or service degradation.

#### Cross-Service Reliability Engineering, Incident Command & Observability
- **Cross-Service Incident Commander:** Served as Primary On-Call Incident Commander for high-severity P1/P0 outages, leading rapid cross-functional triage of complex cascading failures across upstream load balancers, API gateways, compute clusters, and backend persistence layers.
- **Kernel Profiling & Latency Elimination:** Spearheaded OS-level latency and saturation profiling using Linux diagnostic tooling (`perf`, `strace`, `eBPF`, network socket analysis), eliminating inter-service communication bottlenecks and achieving a **28% reduction in p99 API latencies**.
- **Multi-Service Observability & Alerting:** Engineered centralized observability stacks ingesting billions of daily telemetry events across multi-service fleets using Prometheus, Grafana, OpenTelemetry, and Splunk; implemented multi-window burn rate alerts linked to SLO consumption, cutting false-positive pager alerts by 65%.
- **Automated Self-Healing Daemons:** Engineered intelligent remediation daemons in Python and Bash that detect hung service threads, socket exhaustion, and deadlocks across service tiers, executing safe auto-recovery workflows that slashed MTTR by **42%**.
- **Blameless Postmortems & Resilience:** Cultivated a blameless postmortem engineering culture, converting repeat failure patterns into automated health probes, regression checks, and multi-service failover runbooks; decreased recurring production incident volume by 35%.
- **Git Branching & Merge Conflict Governance:** Standardized enterprise trunk-based Git workflows and release branching models across 40+ engineering service squads; authored automated merge conflict triage runbooks (resolving complex multi-parent 3-way merge collisions and rebasing bottlenecks), established pre-receive branch protection hooks, and reduced release integration delays by **45%**.

---

### CtrlS Datacenters Ltd | Jul 2012 – May 2018 (6 Years)
**Linux Infrastructure & Systems Engineer** | Hyderabad, India  
*Managing heterogeneous datacenter infrastructure, database servers, web tiers, and networking across Tier-4 datacenters.*

- **Multi-Service Datacenter Infrastructure:** Managed physical and virtualized enterprise Linux server fleets (RHEL/CentOS) supporting diverse client service tiers, overseeing kernel tuning, storage layout (LVM), network bonding, and multi-tier OS diagnostics across 300+ servers.
- **Web, Middleware & Proxy Administration:** Managed and supported production **Apache HTTP Server, Nginx, and Tomcat** environments; configured software load balancing and reverse proxying with **HAProxy**, SSL/TLS certificate management, and application server clustering.
- **Database Administration (MySQL & PostgreSQL):** Installed, secured, and maintained production **MySQL and PostgreSQL** database instances; configured automated daily backup scripts, user permissions, replication checks, and storage optimization.
- **Enterprise Messaging & Mail Infrastructure:** Administered, tuned, and secured high-volume enterprise Mail Transfer Agents (**Postfix, Sendmail, Zimbra, Qmail**); engineered reliable SMTP mail relays, managed queue processing, and enforced strict **SPF, DKIM, and DMARC** authentication standards.
- **Fleet Automation & Toil Reduction:** Engineered automated operational workflows and system maintenance scripts using Python, Bash, and Cron, automating backups, log rotation, security patching, and health checks.
- **Networking Services & Firewalls:** Configured and tuned core networking services and proxies including BIND DNS (MX, A, PTR records), HAProxy, and IPTables firewall policies.
- **Telemetry & Availability Monitoring:** Implemented proactive server telemetry and infrastructure monitoring using Nagios and MRTG, tracking link saturation, CPU/disk metrics, and hardware availability.

---

## KEY RELIABILITY & OPEN-SOURCE PLATFORM PROJECTS

### ForgeOps AI — Safety-First Telemetry & Autonomous SRE Platform | [github.com/santosh91parsa/forgeops](https://github.com/santosh91parsa/forgeops)
- **Autonomous Infrastructure SRE:** Designed and engineered a capability-bounded reliability platform acting as an unprivileged, read-only autonomous SRE agent over computing infrastructure via Model Context Protocol (MCP).
- **Security & Policy Boundary:** Built a deterministic policy engine and security boundary that prevents confused deputy execution and sanitizes diagnostic logs from indirect prompt injection, masking sensitive secrets (JWTs, API keys, DB credentials) prior to telemetry ingestion.
- **Host Telemetry Collectors:** Built cross-platform host telemetry collectors (CPU load, memory breakdown, active socket mappings, process parentage trees) with zero external runtime dependencies.

### Enterprise Mail & Incident Intelligence Pipeline | [github.com/santosh91parsa/DecodeAIwithSantosh](https://github.com/santosh91parsa/DecodeAIwithSantosh)
- **Email Protocol Parsing & Triage:** Engineered an automated incident and ticket triage ingestion engine for Linux environments; parsed SMTP headers and extracted SPF/DKIM/DMARC metadata, routing confidential PII to local on-premise Ollama instances and slashing API overhead by 85%.

### FastMCP Linux Observability Server | [github.com/santosh91parsa/DecodeAIwithSantosh](https://github.com/santosh91parsa/DecodeAIwithSantosh)
- **Zero-Root Telemetry Microservice:** Developed a lightweight Model Context Protocol microservice exposing realtime CPU, RAM, and storage health metrics over Server-Sent Events (SSE) and local stdio, providing zero-root diagnostic access for remote telemetry clients.

---

## CERTIFICATIONS & EDUCATION
- **CKA: Certified Kubernetes Administrator** — The Linux Foundation ([Credly Verified](https://www.credly.com/users/santosh-parsa/badges))
- **Google Cloud Certified: Associate Cloud Engineer (ACE)** — Google Cloud ([Credly Verified](https://www.credly.com/users/santosh-parsa/badges))
- **Google Cloud Certified: Generative AI Leader** — Google Cloud ([Credly Verified](https://www.credly.com/users/santosh-parsa/badges))
- **Microsoft Certified: Azure Administrator Associate** — Microsoft ([Credly Verified](https://www.credly.com/users/santosh-parsa/badges))
- **Microsoft Certified: Azure Fundamentals (AZ-900)** — Microsoft ([Credly Verified](https://www.credly.com/users/santosh-parsa/badges))
- **Bachelor of Technology (B.Tech)** in Computer Science / Engineering
- **Technical Educator, "Decode AI with Santosh":** Author and educator publishing practical deep dives on Linux internals, Kubernetes orchestration, and AI systems reliability for engineers.
