Curriculum Vitae

Kubernetes • Cloud (AWS/GCP) • Linux Internals • Databases • SRE

📥 Download PDF 📄 Word (.docx) 📋 Markdown (.md)

SANTOSH PARSA

Staff Platform & Infrastructure Engineer | Kubernetes, Cloud (AWS/GCP), Linux & Databases

Professional Summary

Staff Platform and Infrastructure Systems Engineer with 13+ years of experience architecting, managing, and automating mission-critical multi-service infrastructure across enterprise hybrid cloud (AWS, GCP) and datacenters. Proven track record across 7+ years at ServiceNow managing foundational infrastructure tiers powering dozens of interconnected platform microservices and distributed workloads under strict 99.99%+ availability SLAs. Deep technical expertise across Kubernetes cluster orchestration (EKS/GKE), Infrastructure as Code with Terraform, Python and Bash automation, deep Linux OS/kernel internals, and enterprise networking/DNS (Route53, BIND, CoreDNS, TCP/IP). Extensive hands-on experience supporting relational database tiers (MySQL, PostgreSQL), web/application middleware (Apache HTTP Server, Apache Tomcat, Nginx), software load balancing (HAProxy, Nginx, ALB/NLB), and full-stack observability (Prometheus, Grafana, Splunk). Experienced Primary Incident Commander driving cross-service root cause analysis, resolving complex cascading failures, and slashing MTTR by 42% through automated self-healing daemons.

Core Competencies & Technical Skills

  • Containerization & Kubernetes: Kubernetes (EKS, GKE, Multi-Tenant Cluster Management, Helm, Operators, CoreDNS, Ingress Controllers, Service Mesh, CNI/CSI Plugins, HPA, Pod Disruption Budgets), Docker, Container Security, Internal Developer Platforms (IDP), GitOps workflows.
  • Cloud Infrastructure & IaC: Terraform (Modular Infrastructure Registries, State Management, Drift Detection), Ansible (Fleet Configuration Automation, Dynamic Inventories), AWS (EC2, EKS, VPC, Route53, ALB/NLB, RDS, S3, IAM, KMS), Google Cloud Platform (GCP - GKE, VPC, Cloud SQL, IAM), Multi-Region Active-Active Architectures.
  • Operating Systems, Networking & DNS: Deep Enterprise Linux Administration (RHEL, Ubuntu, CentOS), Kernel Tuning, Process & Thread Scheduling, Systemd, Cgroups, eBPF Tracing, Performance Profiling (perf, strace, tcpdump, ss, htop, iotop), Network Protocols (TCP/IP, HTTP/2, gRPC, VPC Routing, Subnetting, Firewalls), DNS Architecture (Route53, BIND DNS, CoreDNS, Anycast).
  • Databases & Relational Data Tiers: MySQL, PostgreSQL, High-Availability Clustering & Replication (Primary-Replica, Failover), Database Connection Pooling (PgBouncer, ProxySQL), Automated Backups & Point-in-Time Recovery (PITR), Schema & Storage Optimization, Query Performance Profiling.
  • Web, Middleware & Software Load Balancing: Apache HTTP Server, Nginx, Apache Tomcat, Software Load Balancing (HAProxy, Nginx Load Balancer, ALB/NLB), Reverse Proxy Architecture, SSL/TLS Offloading & Certificate Management, Connection Pooling & Keep-Alive Tuning, Upstream Health Probes, Sticky Sessions, High-Concurrency Request Routing.
  • Programming & Automation: Python (Infrastructure Automation, APIs, Scripting Frameworks, Standard Libraries), Advanced Bash / Shell Scripting, FastMCP, Model Context Protocol (MCP), LLM-Augmented Telemetry Pipelines, Automated Incident Triage Engines.
  • Site Reliability Engineering (SRE) & Incident Command: SLI/SLO Framework Design, Error Budget Governance, Cross-Service Cascading Failure Triage, MTTD/MTTR Optimization, Blameless Postmortems, Chaos Simulation Game Days, Automated Remediation & Self-Healing Daemons, 24/7 Incident Command.
  • Observability & Telemetry: Prometheus, Grafana, OpenTelemetry (Traces, Metrics, Logs), Splunk, Alertmanager, Multi-Window Multi-Burn-Rate Alerting, Synthetic Cross-Service Health Probes, Availability Dashboards.
  • Version Control, Git Internals & CI/CD: Advanced Git Architecture & DAG mechanics, 3-Way Merge conflict resolution (Lowest Common Ancestor / LCA analysis, recursive strategies, rebase vs merge governance), branch protection policies, DevSecOps hooks (pre-commit, pre-receive secret scanning), disaster recovery (git reflog), GitHub Actions, Jenkins declarative pipelines, JFrog Artifactory, Container Registries, Automated Canary & Blue/Green Deployments, Multi-Service Release Validation.
  • Email & Messaging Infrastructure: Enterprise Mail Transfer Agents (Postfix, Sendmail, Zimbra, Qmail), SMTP Protocol & Relays, Email Authentication & Security Standards (SPF, DKIM, DMARC Enforcement), Transport Security (STARTTLS), DNS MX Record Routing.

Professional Experience

ServiceNow 2018 – Present (7+ Years)
Sr Staff Linux Systems & Infrastructure Platform Engineer (Mar 2025 – Present)
Staff Systems Design & Infrastructure Engineer (Mar 2021 – Mar 2025)
Senior Linux Systems Administrator (May 2018 – Mar 2021)

Hyderabad, India | Technical leader for enterprise multi-service infrastructure management, platform reliability, and cloud operations powering ServiceNow's Global 2000 multi-tenant cloud (99.99%+ uptime).

  • Multi-Service Fleet Operations: Managed foundational cloud and Linux infrastructure supporting dozens of interconnected enterprise platform services, API gateways, and distributed data pipelines; maintained strict 99.99%+ uptime across thousands of compute nodes.
  • Relational Database Operations (MySQL & PostgreSQL): Supported enterprise relational database tiers powered by MySQL and PostgreSQL, managing automated backup pipelines, primary-replica replication health, connection pooling (PgBouncer, ProxySQL), and failover validation across distributed platform services.
  • Middleware & Software Load Balancing: Configured and tuned high-throughput software load balancers (HAProxy, Nginx) fronting multi-service Apache and Tomcat application tiers; optimized connection pools, keep-alive timeouts, buffer limits, and health checks to ensure seamless request distribution and zero-downtime rolling service deployments.
  • Resource Isolation & Capacity Governance: Enforced strict compute, memory, and I/O isolation across co-located multi-service workloads utilizing Linux cgroups and systemd resource slices, eliminating noisy-neighbor degradation and optimizing multi-tenant cluster density.
  • Zero-Downtime Fleet Maintenance: Orchestrated automated rolling OS patching, kernel upgrades, and security remediations across multi-service server fleets without customer traffic interruption or service degradation.
  • Cross-Service Incident Commander: Served as Primary On-Call Incident Commander for high-severity P1/P0 outages, leading rapid cross-functional triage of complex cascading failures across upstream load balancers, API gateways, compute clusters, and backend persistence layers.
  • Kernel Profiling & Latency Elimination: Spearheaded OS-level latency and saturation profiling using Linux diagnostic tooling (perf, strace, eBPF, network socket analysis), eliminating inter-service communication bottlenecks and achieving a 28% reduction in p99 API latencies.
  • Multi-Service Observability & Alerting: Engineered centralized observability stacks ingesting billions of daily telemetry events across multi-service fleets using Prometheus, Grafana, OpenTelemetry, and Splunk; implemented multi-window burn rate alerts linked to SLO consumption, cutting false-positive pager alerts by 65%.
  • Automated Self-Healing Daemons: Engineered intelligent remediation daemons in Python and Bash that detect hung service threads, socket exhaustion, and deadlocks across service tiers, executing safe auto-recovery workflows that slashed MTTR by 42%.
  • Blameless Postmortems & Resilience: Cultivated a blameless postmortem engineering culture, converting repeat failure patterns into automated health probes, regression checks, and multi-service failover runbooks; decreased recurring production incident volume by 35%.
  • Git Branching & Merge Conflict Governance: Standardized enterprise trunk-based Git workflows and release branching models across 40+ engineering service squads; authored automated merge conflict triage runbooks (resolving complex multi-parent 3-way merge collisions and rebasing bottlenecks), established pre-receive branch protection hooks, and reduced release integration delays by 45%.
CtrlS Datacenters Ltd Jul 2012 – May 2018 (6 Years)
Linux Infrastructure & Systems Engineer

Hyderabad, India | Managing heterogeneous datacenter infrastructure, database servers, web tiers, and networking across Tier-4 datacenters.

  • Multi-Service Datacenter Infrastructure: Managed physical and virtualized enterprise Linux server fleets (RHEL/CentOS) supporting diverse client service tiers, overseeing kernel tuning, storage layout (LVM), network bonding, and multi-tier OS diagnostics across 300+ servers.
  • Web, Middleware & Proxy Administration: Managed and supported production Apache HTTP Server, Nginx, and Tomcat environments; configured software load balancing and reverse proxying with HAProxy, SSL/TLS certificate management, and application server clustering.
  • Database Administration (MySQL & PostgreSQL): Installed, secured, and maintained production MySQL and PostgreSQL database instances; configured automated daily backup scripts, user permissions, replication checks, and storage optimization.
  • Enterprise Messaging & Mail Infrastructure: Administered, tuned, and secured high-volume enterprise Mail Transfer Agents (Postfix, Sendmail, Zimbra, Qmail); engineered reliable SMTP mail relays, managed queue processing, and enforced strict SPF, DKIM, and DMARC authentication standards.
  • Fleet Automation & Toil Reduction: Engineered automated operational workflows and system maintenance scripts using Python, Bash, and Cron, automating backups, log rotation, security patching, and health checks.
  • Networking Services & Firewalls: Configured and tuned core networking services and proxies including BIND DNS (MX, A, PTR records), HAProxy, and IPTables firewall policies.
  • Telemetry & Availability Monitoring: Implemented proactive server telemetry and infrastructure monitoring using Nagios and MRTG, tracking link saturation, CPU/disk metrics, and hardware availability.

Key Reliability & Open-Source Projects

  • ForgeOps AI (Autonomous SRE & Telemetry Platform): Designed capability-bounded reliability platform acting as an unprivileged, read-only autonomous SRE agent over computing infrastructure via Model Context Protocol (MCP) with prompt-injection neutralization and credential masking. [github.com/santosh91parsa/forgeops]
  • FastMCP Linux Observability Server: Lightweight MCP microservice exposing realtime CPU, RAM, and storage health metrics over Server-Sent Events (SSE) and local stdio, providing zero-root diagnostic access. [github.com/santosh91parsa/DecodeAIwithSantosh]
  • Enterprise Mail & Incident Intelligence Pipeline: Automated incident triage ingestion engine for Linux environments; parsed SMTP headers and extracted SPF/DKIM/DMARC metadata, routing confidential PII to local on-premise Ollama instances.

Certifications & Education

  • CKA: Certified Kubernetes Administrator — The Linux Foundation [Credly Verified]
  • Google Cloud Certified: Associate Cloud Engineer (ACE) — Google Cloud [Credly Verified]
  • Google Cloud Certified: Generative AI Leader — Google Cloud [Credly Verified]
  • Microsoft Certified: Azure Administrator Associate — Microsoft [Credly Verified]
  • Microsoft Certified: Azure Fundamentals (AZ-900) — Microsoft [Credly Verified]
  • Bachelor of Technology (B.Tech) in Computer Science / Engineering
  • Technical Educator, "Decode AI with Santosh": Author and educator publishing practical deep dives on Linux internals, Kubernetes orchestration, and AI systems reliability for engineers.