LAB 01 • NETWORKING & KERNEL
Linux Kernel TCP Backlog & Socket Saturation Diagnostics
The Problem: When high-throughput services experience micro-bursts of traffic, clients observe connection timeouts or connection resets (RST), but CPU and memory utilization remain low. This experiment isolates listen backlog queue drops from TCP handshake accept queue saturation.
1. Inspecting Listen Queue Overflow
# Check active listener socket queues (Send-Q = backlog limit, Recv-Q = current pending connections)
$ ss -lnt '( sport = :8080 )'
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 129 128 0.0.0.0:8080 0.0.0.0:*
# Check system-wide TCP drop counter metrics
$ netstat -s | grep -i "listen"
14820 times the listen queue of a socket overflowed
14820 SYNs to LISTEN sockets dropped
2. Root Cause & Kernel Parameter Tuning
When Recv-Q > Send-Q, the application's listen(fd, backlog) parameter or the kernel net.core.somaxconn cap was exceeded.
# Tune OS socket listen queue maximums in /etc/sysctl.d/99-network.conf
net.core.somaxconn = 4096
net.ipv4.tcp_max_syn_backlog = 8192
net.ipv4.tcp_abort_on_overflow = 0
# Apply dynamically
$ sudo sysctl -p /etc/sysctl.d/99-network.conf
LAB 02 • SRE & OBSERVABILITY
Multi-Window Multi-Burn-Rate SLO Alerting in Prometheus
The Problem: Traditional static threshold alerts (e.g., "alert if error rate > 1%") trigger massive alert fatigue during low-traffic periods and fail to alert quickly during rapid outages. This lab implements Google SRE multi-window multi-burn-rate alerting.
1. Calculating Burn Rate (99.9% Target / 0.1% Error Budget)
• 14.4x burn rate consumes 2% of budget in 1 hour → Page immediately (Critical)
• 6x burn rate consumes 5% of budget in 6 hours → Ticket/Warning
# Prometheus Alert Rule: Multi-Window Alerting for 99.9% Availability SLO
- alert: ApiHighErrorBudgetBurnRate
expr: (
(rate(http_requests_total{status=~"5.."}[1h]) / rate(http_requests_total[1h])) > (14.4 * 0.001)
and
(rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m])) > (14.4 * 0.001)
)
for: 2m
labels:
severity: page
annotations:
summary: "API 1-hour error budget burn rate is critical (> 14.4x)"
LAB 03 • SYSTEMS DIAGNOSTICS
Production JVM Thread Contention & CPU Spikes on Linux
The Problem: A Java microservice spikes to 100% CPU on a production Linux node. Walkthrough isolating the exact lightweight thread (LWP) and correlating it with thread stack traces without taking the service down.
# Step 1: Find the high CPU Thread ID (Lightweight Process / LWP)
$ top -H -p 24810
PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
24892 appuser 20 0 16.4g 8.2g 32400 R 98.4 25.6 5:24.12 java
# Step 2: Convert decimal Thread ID to hexadecimal
$ printf '%x\n' 24892
613c
# Step 3: Grep thread dump for the nid (Native Thread ID)
$ jstack 24810 | grep -A 30 'nid=0x613c'
"http-nio-8080-exec-42" #142 daemon prio=5 os_prio=0 cpu=324120.45ms elapsed=340.12s tid=0x00007f3b nid=0x613c runnable [0x00007f3a9e32f000]
java.lang.Thread.State: RUNNABLE
at com.enterprise.auth.TokenValidator.verifySignature(TokenValidator.java:184)
at com.enterprise.security.JwtFilter.doFilter(JwtFilter.java:62)
LAB 04 • SYSTEMS ARCHITECTURE & INTERACTIVE SIMULATOR
How Git Works Under the Hood: State Machine & Conflict Engine
The Problem: Engineers often treat Git as black-box magic, leading to merge panic, accidental data loss, and corrupted branch topologies. This hands-on lab exposes the internal content-addressable DAG, simulates 4-zone state transitions in real time, and walks through a live 3-way conflict resolution engine.
4-Zone State Simulator
Interactive Conflict Editor
Object DAG & Blobs
10 Systems Topics
Knowledge Check
⚡ Launch Interactive Git Lab →
LAB 05 • DISTRIBUTED SYSTEMS & DETERMINISTIC AI ORCHESTRATION
DAG-Ops: Deterministic DAG vs. Monolithic LLM Pipeline Simulator
The Problem: Monolithic single-prompt AI workflows suffer from severe quadratic attention degradation, cascade poisoning, and catastrophic 100% pipeline crashes on transient errors (like HTTP 504 timeouts). This lab proves why Directed Acyclic Graph (DAG) workflow engines and topological scheduling deliver 42% faster execution, 48% lower token costs, localized fault boundaries, and deterministic security firewalls.
Dual-Pipeline Visualizer
Topological Sort (Kahn's)
Parallel Fan-Out
Isolated Blast Radius
Regex PII Firewall
Zero Dependencies
⚡ Launch Interactive DAG Simulator →